The Eval Index / Agent Eval / #83
suyoumo/ClawProBench
by suyoumo · Agent Eval · updated 18d ago
ClawProBench is a live-first benchmark harness for evaluating LLM agents in the OpenClaw runtime with deterministic grading and repeated-trial reliability.
64
momentum
823
stars
54
forks
#83
rank
agentbenchmarkevaluationharnessleaderboardllmopenclaw
View on GitHub →