The Eval Index / Benchmarks / #202

yance521/super_pe_evaluation

by yance521 · Benchmarks · updated 2mo ago

A TDD-style prompt evaluation skill — build reusable eval datasets, run prompt versions against locked test cases, score outputs with evidence, and compare results without overwriting history. Local-first, platform-agnostic, append-only.

36
momentum
17
stars
0
forks
#202
rank
View on GitHub →