The Eval Index / Agent Eval / #163
QuesmaOrg/BinaryAudit
by QuesmaOrg · Agent Eval · updated 1mo ago
An open-source benchmark for evaluating AI agents' ability to find backdoors hidden in compiled binaries.
46
momentum
99
stars
6
forks
#163
rank
aibenchmarkbinary-analysiscybersecurityllm-evalreverse-engineering
View on GitHub →