The Eval Index / Agent Eval / #163

QuesmaOrg/BinaryAudit

by QuesmaOrg · Agent Eval · updated 1mo ago

An open-source benchmark for evaluating AI agents' ability to find backdoors hidden in compiled binaries.

46
momentum
99
stars
6
forks
#163
rank
aibenchmarkbinary-analysiscybersecurityllm-evalreverse-engineering
View on GitHub →