Skip to content

rate on a rubric benchmark · release 2026-09-18

scoresort

Place the state on an ordered rubric. The model returns a probability for every level, ready to sort by.

the test

How helpful is the response to the prompt?

HelpSteer2 helpfulness · 5 levels · 300 questions × 5 runs · human labels

ranking

System One LLM baseline

No model clearly beats guessing yet: the top rows can't be told apart from answering with the label base rates.

Ordered by wins (top-2 finishes across score, accuracy, calibration gap, cost and speed), then Decision Score. Select a model for details; select a column header to sort.
model95% range of the Decision Scorewins
1
Jev
System One TypeSafe AI
9.241.3%19.7$0.036670 ms3
accuracy41.3%95% range 35.946.9%
calibration gap19.7 ptslower is better · 0.3% of answers at confidence 1.00
10said 27%, right 47% (34 answers)said 36%, right 31% (118 answers)said 45%, right 31% (261 answers)said 54%, right 31% (339 answers)said 65%, right 41% (324 answers)said 75%, right 53% (222 answers)said 83%, right 61% (157 answers)said 94%, right 76% (45 answers)
automation at 95%never reaches 95% accuracy
speed478 msmedian · p95 670 ms
stability2.3%of answers change on an identical rerun
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupai-sdk experimental_evaluate ai@7.0.106 · native probabilities via Vercel AI Gateway · reasoning · temperature · 300 items × 5 runs · 2026-09-18 · served as typesafe-ai/jevmodel pagehelpsteer2 run log
=2
GLM-5.3
LLM Z.ai · reasoning low
7.843.0%12.9$0.783.6 s2
accuracy43.0%95% range 38.647.9%
calibration gap12.9 ptslower is better · 0.0% of answers at confidence 1.00
10said 35%, right 25% (4 answers)said 44%, right 24% (279 answers)said 52%, right 40% (722 answers)said 61%, right 53% (280 answers)said 71%, right 62% (130 answers)said 83%, right 67% (63 answers)said 92%, right 77% (22 answers)
automation at 95%never reaches 95% accuracy
speed2.2 smedian · p95 3.6 s
stability31.3%of answers change on an identical rerun
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Novita (fp8) · reasoning low (mandatory) · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as z-ai/glm-5.3model pagehelpsteer2 run log
=2
Mercury 2.5
LLM Inception · no reasoning
−5.541.7%31.5$0.0341.2 s2
accuracy41.7%95% range 37.346.2%
calibration gap31.5 ptslower is better · 0.1% of answers at confidence 1.00
10said 31%, right 0% (4 answers)said 43%, right 27% (109 answers)said 52%, right 31% (152 answers)said 62%, right 35% (368 answers)said 72%, right 31% (193 answers)said 84%, right 48% (287 answers)said 94%, right 58% (387 answers)
automation at 95%never reaches 95% accuracy
speed618 msmedian · p95 1.2 s
stability38.7%of answers change on an identical rerun
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Inception · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as inception/mercury-2.5model pagehelpsteer2 run log
4
Mistral Medium 3.5
LLM Mistral AI · no reasoning
−13.743.9%33.6$1.31.2 s1
accuracy43.9%95% range 38.549.3%
calibration gap33.6 ptslower is better · 2.1% of answers at confidence 1.00
10said 40%, right 0% (5 answers)said 60%, right 42% (90 answers)said 70%, right 35% (578 answers)said 80%, right 46% (474 answers)said 91%, right 56% (353 answers)
automation at 95%never reaches 95% accuracy
speed706 msmedian · p95 1.2 s
stability3.7%of answers change on an identical rerun
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Mistral · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as mistralai/mistral-medium-3-5model pagehelpsteer2 run log
=5
Gemini 3.8 Flash
LLM Google · reasoning low
4.642.4%22.6$1.34.0 s0
accuracy42.4%95% range 36.947.5%
calibration gap22.6 ptslower is better · 0.1% of answers at confidence 1.00
10said 44%, right 33% (106 answers)said 53%, right 24% (579 answers)said 64%, right 39% (269 answers)said 71%, right 60% (153 answers)said 84%, right 70% (252 answers)said 93%, right 63% (141 answers)
automation at 95%never reaches 95% accuracy
speed1.9 smedian · p95 4.0 s
stability11.7%of answers change on an identical rerun
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Google AI Studio · reasoning low (mandatory) · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as google/gemini-3.8-flashmodel pagehelpsteer2 run log
Label prior
Baseline
guessing
0.041.7%0.0
accuracy41.7%95% range 36.047.0%
calibration gap0.0 ptslower is better · 0.0% of answers at confidence 1.00
automation at 95%no probabilities
speedmedian · p95
stability0.0%of answers change on an identical rerun
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupconstant probabilities · reasoning · temperature · 300 items × 5 runsmodel page
=5
Qwen3.8 Flash
LLM Alibaba · no reasoning
−1.436.1%25.8$0.132.6 s0
accuracy36.1%95% range 31.540.7%
calibration gap25.8 ptslower is better · 0.0% of answers at confidence 1.00
10said 44%, right 23% (232 answers)said 50%, right 27% (353 answers)said 62%, right 34% (594 answers)said 73%, right 48% (46 answers)said 85%, right 62% (151 answers)said 95%, right 60% (124 answers)
automation at 95%never reaches 95% accuracy
speed1.4 smedian · p95 2.6 s
stability26.0%of answers change on an identical rerun
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Alibaba · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as qwen/qwen3.8-flashmodel pagehelpsteer2 run log
=5
DeepSeek V4.1 Flash
LLM DeepSeek · no reasoning
−19.034.7%33.5$0.131.2 s0
accuracy34.7%95% range 30.538.7%
calibration gap33.5 ptslower is better · 1.5% of answers at confidence 1.00
10said 34%, right 9% (11 answers)said 44%, right 16% (102 answers)said 54%, right 22% (351 answers)said 62%, right 28% (346 answers)said 73%, right 34% (267 answers)said 85%, right 54% (261 answers)said 95%, right 60% (162 answers)
automation at 95%never reaches 95% accuracy
speed912 msmedian · p95 1.2 s
stability34.3%of answers change on an identical rerun
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · DeepSeek · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as deepseek/deepseek-v4.1-flashmodel pagehelpsteer2 run log

Score: 100 = perfect, 0 = no better than guessing. The bar is the 95% range. Select a model to see everything behind its row.

Wins: a model wins a column when it is in the top two for that column (score, accuracy, calibration gap, cost or speed). The table is sorted by most wins, then by score. =3 means tied on wins.

score vs price

Up and to the left is better. The line joins the best value at each price.
−25−20−15−10−5051015guessing$0.05$0.1$0.2$0.5$1$ per 1k decisions, log scale →Jev: Decision Score 9.2 · $0.036 per 1k decisionsJevGLM-5.3: Decision Score 7.8 · $0.78 per 1k decisionsGLM-5.3Gemini 3.8 Flash: Decision Score 4.6 · $1.3 per 1k decisionsGemini 3.8 FlashQwen3.8 Flash: Decision Score −1.4 · $0.13 per 1k decisionsQwen3.8 FlashMercury 2.5: Decision Score −5.5 · $0.034 per 1k decisionsMercury 2.5Mistral Medium 3.5: Decision Score −13.7 · $1.3 per 1k decisionsMistral Medium 3.5DeepSeek V4.1 Flash: Decision Score −19.0 · $0.13 per 1k decisionsDeepSeek V4.1 Flash

score vs speed

Up and to the left is better. Measured end to end, one question per request.
−25−20−15−10−5051015guessing500 ms1 s2 s5 sp95 latency, log scale →Jev: Decision Score 9.2 · p95 670 msJevGLM-5.3: Decision Score 7.8 · p95 3.6 sGLM-5.3Gemini 3.8 Flash: Decision Score 4.6 · p95 4.0 sGemini 3.8 FlashQwen3.8 Flash: Decision Score −1.4 · p95 2.6 sQwen3.8 FlashMercury 2.5: Decision Score −5.5 · p95 1.2 sMercury 2.5Mistral Medium 3.5: Decision Score −13.7 · p95 1.2 sMistral Medium 3.5DeepSeek V4.1 Flash: Decision Score −19.0 · p95 1.2 sDeepSeek V4.1 Flash

automation at 95% accuracy

Share of questions each model can answer alone and stay 95% right. It answers only above its confidence cut-off; the rest go to a person.

No model reaches 95% accuracy at any confidence on this task. Keep a person in the loop.

release 2026-09-18 · suite 0.1.0 · board.json · permalink · explore every question · how we score