The Eval Index / RAG Eval / #62
chunxiaoxx/nautilus-compass
by chunxiaoxx · RAG Eval · updated today
Open memory & reliability layer for AI agents + independent judging/verification (NACRE judge, caliber-bench meta-benchmark, preregistered criteria) — by Nautilus Platform
69
momentum
1,256
stars
36
forks
#62
rank
a2aagent-memoryagentic-aibenchmarkclaude-codecross-agentevaluationllmllm-as-judgelocal-firstlong-term-memorymcp
View on GitHub →