You are looking at the frozen release 2026-09-18. See the latest boards
rate on a rubric benchmark · release 2026-09-18
score → sort
Place the state on an ordered rubric. The model returns a probability for every level, ready to sort by.
“How helpful is the response to the prompt?”
ranking
System One LLM baseline
No model clearly beats guessing yet: the top rows can't be told apart from answering with the label base rates.
| model | 95% range of the Decision Score | wins | ||||||
|---|---|---|---|---|---|---|---|---|
| 1 | Jev | 9.2 | 41.3% | 19.7 | $0.036 | 670 ms | 3 | |
accuracy41.3%95% range 35.9–46.9% calibration gap19.7 ptslower is better · 0.3% of answers at confidence 1.00 automation at 95%—never reaches 95% accuracy speed478 msmedian · p95 670 ms stability2.3%of answers change on an identical rerun valid output100.0%0.0% refused · 1,500 of 1,500 answered setupai-sdk experimental_evaluate ai@7.0.106 · native probabilities via Vercel AI Gateway · reasoning — · temperature — · 300 items × 5 runs · 2026-09-18 · served as typesafe-ai/jevhelpsteer2 run log | ||||||||
| =2 | GLM-5.3 | 7.8 | 43.0% | 12.9 | $0.78 | 3.6 s | 2 | |
accuracy43.0%95% range 38.6–47.9% calibration gap12.9 ptslower is better · 0.0% of answers at confidence 1.00 automation at 95%—never reaches 95% accuracy speed2.2 smedian · p95 3.6 s stability31.3%of answers change on an identical rerun valid output100.0%0.0% refused · 1,500 of 1,500 answered setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Novita (fp8) · reasoning low (mandatory) · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as z-ai/glm-5.3helpsteer2 run log | ||||||||
| =2 | Mercury 2.5 | −5.5 | 41.7% | 31.5 | $0.034 | 1.2 s | 2 | |
accuracy41.7%95% range 37.3–46.2% calibration gap31.5 ptslower is better · 0.1% of answers at confidence 1.00 automation at 95%—never reaches 95% accuracy speed618 msmedian · p95 1.2 s stability38.7%of answers change on an identical rerun valid output100.0%0.0% refused · 1,500 of 1,500 answered setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Inception · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as inception/mercury-2.5helpsteer2 run log | ||||||||
| 4 | Mistral Medium 3.5 | −13.7 | 43.9% | 33.6 | $1.3 | 1.2 s | 1 | |
accuracy43.9%95% range 38.5–49.3% calibration gap33.6 ptslower is better · 2.1% of answers at confidence 1.00 automation at 95%—never reaches 95% accuracy speed706 msmedian · p95 1.2 s stability3.7%of answers change on an identical rerun valid output100.0%0.0% refused · 1,500 of 1,500 answered setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Mistral · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as mistralai/mistral-medium-3-5helpsteer2 run log | ||||||||
| =5 | Gemini 3.8 Flash | 4.6 | 42.4% | 22.6 | $1.3 | 4.0 s | 0 | |
accuracy42.4%95% range 36.9–47.5% calibration gap22.6 ptslower is better · 0.1% of answers at confidence 1.00 automation at 95%—never reaches 95% accuracy speed1.9 smedian · p95 4.0 s stability11.7%of answers change on an identical rerun valid output100.0%0.0% refused · 1,500 of 1,500 answered setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Google AI Studio · reasoning low (mandatory) · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as google/gemini-3.8-flashhelpsteer2 run log | ||||||||
| — | Label prior | 0.0 | 41.7% | 0.0 | — | — | — | |
accuracy41.7%95% range 36.0–47.0% calibration gap0.0 ptslower is better · 0.0% of answers at confidence 1.00 automation at 95%—no probabilities speed—median · p95 — stability0.0%of answers change on an identical rerun valid output100.0%0.0% refused · 1,500 of 1,500 answered setupconstant probabilities · reasoning — · temperature — · 300 items × 5 runs | ||||||||
| =5 | Qwen3.8 Flash | −1.4 | 36.1% | 25.8 | $0.13 | 2.6 s | 0 | |
accuracy36.1%95% range 31.5–40.7% calibration gap25.8 ptslower is better · 0.0% of answers at confidence 1.00 automation at 95%—never reaches 95% accuracy speed1.4 smedian · p95 2.6 s stability26.0%of answers change on an identical rerun valid output100.0%0.0% refused · 1,500 of 1,500 answered setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Alibaba · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as qwen/qwen3.8-flashhelpsteer2 run log | ||||||||
| =5 | DeepSeek V4.1 Flash | −19.0 | 34.7% | 33.5 | $0.13 | 1.2 s | 0 | |
accuracy34.7%95% range 30.5–38.7% calibration gap33.5 ptslower is better · 1.5% of answers at confidence 1.00 automation at 95%—never reaches 95% accuracy speed912 msmedian · p95 1.2 s stability34.3%of answers change on an identical rerun valid output100.0%0.0% refused · 1,500 of 1,500 answered setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · DeepSeek · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as deepseek/deepseek-v4.1-flashhelpsteer2 run log | ||||||||
Score: 100 = perfect, 0 = no better than guessing. The bar is the 95% range. Select a model to see everything behind its row.
Wins: a model wins a column when it is in the top two for that column (score, accuracy, calibration gap, cost or speed). The table is sorted by most wins, then by score. =3 means tied on wins.
score vs price
score vs speed
automation at 95% accuracy
Share of questions each model can answer alone and stay 95% right. It answers only above its confidence cut-off; the rest go to a person.
No model reaches 95% accuracy at any confidence on this task. Keep a person in the loop.
release 2026-09-18 · suite 0.1.0 · board.json · permalink · explore every question · how we score