The Eval Index / Benchmarks / #99
sunxin-ai/dsh-design-qa
by sunxin-ai · Benchmarks · updated 5d ago
Design-fidelity QA for DeepSeek Harness: lend any text-only model an eye, then judge whether the implementation matches the mock. Ships the benchmark behind that judgement — four fixtures, 23 injected defects, and every raw model transcript. Retires itself when DeepSeek ships vision.
61
momentum
43
stars
3
forks
#99
rank
benchmarkdeepseek-harnessdesign-qadesign-reviewdshdsh-pluginllm-evalmultimodalopenai-compatibleqwen-vlvisionvisual-regression
View on GitHub →