The Eval Index / Agent Eval / #135

AutoTrustAI/PaperGuru-Benchmark

by AutoTrustAI · Agent Eval · updated 3mo ago

Lifecycle-Aware Memory for long-horizon LLM agents — 66.05% on PaperBench, 94.66% on SurveyBench, 10 peer-reviewed acceptances at FSE/ICML/TOSEM/AEI/ICoGB

53
momentum
1,324
stars
197
forks
#135
rank
View on GitHub →