The Eval Index / Benchmarks / #116

openai/evals

by openai · Benchmarks · updated 5mo ago

Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.

57
momentum
19,442
stars
3,084
forks
#116
rank
View on GitHub →