The Eval Index / Benchmarks / #110

yance521/super_pe_evaluation

by yance521 · Benchmarks · updated 11d ago

A TDD-style prompt evaluation skill — build reusable eval datasets, run prompt versions against locked test cases, score outputs with evidence, and compare results without overwriting history. Local-first, platform-agnostic, append-only.

58
momentum
21
stars
0
forks
#110
rank
View on GitHub →