The Eval Index / Benchmarks / #182

huggingface/evaluation-guidebook

by huggingface · Benchmarks · updated 9mo ago

Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard and designing lighteval!

40
momentum
2,145
stars
126
forks
#182
rank
evaluationevaluation-metricsguidebooklarge-language-modelsllmmachine-learningtutorial
View on GitHub →