The Eval Index / RAG Eval / #62

chunxiaoxx/nautilus-compass

by chunxiaoxx · RAG Eval · updated today

Open memory & reliability layer for AI agents + independent judging/verification (NACRE judge, caliber-bench meta-benchmark, preregistered criteria) — by Nautilus Platform

69
momentum
1,256
stars
36
forks
#62
rank
a2aagent-memoryagentic-aibenchmarkclaude-codecross-agentevaluationllmllm-as-judgelocal-firstlong-term-memorymcp
View on GitHub →