The Eval Index / Agent Eval / #117

pinchbench/skill

by pinchbench · Agent Eval · updated 2mo ago

PinchBench is a benchmarking system for evaluating LLM models as OpenClaw coding agents. Made with 🦀 by the humans at https://kilo.ai

57
momentum
1,343
stars
158
forks
#117
rank
View on GitHub →