The Eval Index / Reasoning / #124

ai-twinkle/Eval

by ai-twinkle · Reasoning · updated 1d ago

High-performance LLM evaluation framework with parallel API calls — up to 17× faster than sequential tools. Supports box, math, and logit-based evaluation.

57
momentum
110
stars
19
forks
#124
rank
evalevaluationllm
View on GitHub →