Skip to content

pick one of N benchmark · release 2026-09-18

choicematch

Pick one option from a list. The model returns a probability for every option, like a match over its cases.

the test

Which intent does this banking customer's message express?

Banking77 · 77 options · 300 questions × 5 runs · human labels

ranking

System One LLM baseline top 2 for that metric (score, cost, speed)

Ordered by wins (top-2 finishes across score, accuracy, calibration gap, cost and speed), then Decision Score. Select a model for details; select a column header to sort.
model95% range of the Decision Scorewins
1
Jev
System One TypeSafe AI
67.879.7%9.8$0.043693 ms4
accuracy79.7%95% range 75.583.8%
calibration gap9.8 ptslower is better · 41.5% of answers at confidence 1.00
10said 29%, right 0% (1 answers)said 35%, right 20% (25 answers)said 44%, right 23% (40 answers)said 54%, right 51% (69 answers)said 64%, right 54% (69 answers)said 74%, right 63% (95 answers)said 85%, right 60% (131 answers)said 98%, right 91% (1070 answers)
automation at 95%54%of decisions, acting at confidence ≥ 0.98 (right 95.4%)shared gate ≥ 0.96: acts on 59%, right 94.3%
speed467 msmedian · p95 693 ms
stability2.7%of answers change on an identical rerun · 10.3% change when options are reordered
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupai-sdk experimental_evaluate ai@7.0.106 · native probabilities via Vercel AI Gateway · reasoning · temperature · 300 items × 5 runs · 2026-09-18 · served as typesafe-ai/jevmodel pagebanking77 run log
2
Gemini 3.8 Flash
LLM Google · reasoning low
74.184.6%3.2$1.43.9 s3
accuracy84.6%95% range 80.888.2%
calibration gap3.2 ptslower is better · 0.0% of answers at confidence 1.00
10said 35%, right 17% (6 answers)said 45%, right 47% (19 answers)said 55%, right 62% (66 answers)said 65%, right 63% (68 answers)said 74%, right 73% (82 answers)said 86%, right 80% (588 answers)said 95%, right 96% (671 answers)
automation at 95%45%of decisions, acting at confidence ≥ 0.89 (right 96.1%)shared gate ≥ 0.96: acts on 12%, right 99.4%
speed1.6 smedian · p95 3.9 s
stability3.0%of answers change on an identical rerun · 8.7% change when options are reordered
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Google AI Studio · reasoning low (mandatory) · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as google/gemini-3.8-flashmodel pagebanking77 run log
3
DeepSeek V4.1 Flash
LLM DeepSeek · no reasoning
63.975.9%2.8$0.141.3 s2
accuracy75.9%95% range 71.980.1%
calibration gap2.8 ptslower is better · 0.6% of answers at confidence 1.00
10said 18%, right 0% (2 answers)said 26%, right 0% (3 answers)said 35%, right 22% (50 answers)said 44%, right 44% (109 answers)said 55%, right 50% (155 answers)said 63%, right 66% (183 answers)said 73%, right 69% (216 answers)said 84%, right 88% (219 answers)said 95%, right 96% (561 answers)
automation at 95%39%of decisions, acting at confidence ≥ 0.86 (right 95.8%)shared gate ≥ 0.96: acts on 15%, right 99.1%
speed999 msmedian · p95 1.3 s
stability10.7%of answers change on an identical rerun · 30.2% change when options are reordered
valid output99.9%0.0% refused · 1,498 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · DeepSeek · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as deepseek/deepseek-v4.1-flashmodel pagebanking77 run log
4
Mercury 2.5
LLM Inception · no reasoning
54.070.3%10.4$0.0361.5 s1
accuracy70.3%95% range 66.174.7%
calibration gap10.4 ptslower is better · 0.0% of answers at confidence 1.00
10said 25%, right 0% (6 answers)said 34%, right 6% (18 answers)said 44%, right 28% (82 answers)said 52%, right 23% (31 answers)said 63%, right 49% (144 answers)said 73%, right 56% (179 answers)said 85%, right 73% (451 answers)said 94%, right 90% (584 answers)
automation at 95%never reaches 95% accuracyshared gate ≥ 0.96: acts on 1%, right 100.0%
speed638 msmedian · p95 1.5 s
stability13.8%of answers change on an identical rerun · 36.1% change when options are reordered
valid output99.7%0.0% refused · 1,495 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Inception · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as inception/mercury-2.5model pagebanking77 run log
=5
GLM-5.3
LLM Z.ai · reasoning low
66.878.9%6.2$0.723.2 s0
accuracy78.9%95% range 74.982.7%
calibration gap6.2 ptslower is better · 0.0% of answers at confidence 1.00
10said 25%, right 33% (3 answers)said 33%, right 39% (38 answers)said 44%, right 43% (120 answers)said 54%, right 57% (183 answers)said 62%, right 70% (173 answers)said 73%, right 82% (283 answers)said 84%, right 91% (338 answers)said 92%, right 97% (362 answers)
automation at 95%38%of decisions, acting at confidence ≥ 0.83 (right 95.1%)shared gate ≥ 0.96: acts on 0%, right 100.0%
speed2.1 smedian · p95 3.2 s
stability9.7%of answers change on an identical rerun · 23.7% change when options are reordered
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Novita (fp8) · reasoning low (mandatory) · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as z-ai/glm-5.3model pagebanking77 run log
=5
Qwen3.8 Flash
LLM Alibaba · no reasoning
62.476.8%9.5$0.123.1 s0
accuracy76.8%95% range 72.680.9%
calibration gap9.5 ptslower is better · 0.0% of answers at confidence 1.00
10said 35%, right 43% (14 answers)said 45%, right 36% (76 answers)said 50%, right 33% (3 answers)said 64%, right 50% (123 answers)said 75%, right 50% (4 answers)said 85%, right 66% (435 answers)said 95%, right 91% (843 answers)
automation at 95%never reaches 95% accuracyshared gate ≥ 0.96: acts on 2%, right 100.0%
speed1.7 smedian · p95 3.1 s
stability9.0%of answers change on an identical rerun · 22.1% change when options are reordered
valid output99.9%0.0% refused · 1,498 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Alibaba · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as qwen/qwen3.8-flashmodel pagebanking77 run log
=5
Mistral Medium 3.5
LLM Mistral AI · no reasoning
59.574.5%8.1$1.51.4 s0
accuracy74.5%95% range 70.178.9%
calibration gap8.1 ptslower is better · 0.0% of answers at confidence 1.00
10said 40%, right 43% (14 answers)said 50%, right 33% (113 answers)said 60%, right 59% (472 answers)said 70%, right 81% (320 answers)said 80%, right 92% (417 answers)said 92%, right 100% (157 answers)
automation at 95%11%of decisions, acting at confidence ≥ 0.81 (right 99.4%)shared gate ≥ 0.96: acts on 0%
speed936 msmedian · p95 1.4 s
stability2.3%of answers change on an identical rerun · 22.2% change when options are reordered
valid output99.5%0.0% refused · 1,493 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Mistral · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as mistralai/mistral-medium-3-5model pagebanking77 run log
Label prior
Baseline
← 0 = guessing, off the scale
0.01.3%0.0
accuracy1.3%95% range 0.32.7%
calibration gap0.0 ptslower is better · 0.0% of answers at confidence 1.00
automation at 95%no probabilitiesshared gate ≥ 0.96: acts on 0%
speedmedian · p95
stability0.0%of answers change on an identical rerun · 0.0% change when options are reordered
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupconstant probabilities · reasoning · temperature · 300 items × 5 runsmodel page

Score: 100 = perfect, 0 = no better than guessing. The bar is the 95% range. Select a model to see everything behind its row.

Wins: a model wins a column when it is in the top two for that column (score, accuracy, calibration gap, cost or speed). The table is sorted by most wins, then by score. =3 means tied on wins.

score vs price

Up and to the left is better. The line joins the best value at each price.
50556065707580$0.05$0.1$0.2$0.5$1$2$ per 1k decisions, log scale →Gemini 3.8 Flash: Decision Score 74.1 · $1.4 per 1k decisionsGemini 3.8 FlashJev: Decision Score 67.8 · $0.043 per 1k decisionsJevGLM-5.3: Decision Score 66.8 · $0.72 per 1k decisionsGLM-5.3DeepSeek V4.1 Flash: Decision Score 63.9 · $0.14 per 1k decisionsDeepSeek V4.1 FlashQwen3.8 Flash: Decision Score 62.4 · $0.12 per 1k decisionsQwen3.8 FlashMistral Medium 3.5: Decision Score 59.5 · $1.5 per 1k decisionsMistral Medium 3.5Mercury 2.5: Decision Score 54.0 · $0.036 per 1k decisionsMercury 2.5

score vs speed

Up and to the left is better. Measured end to end, one question per request.
50556065707580500 ms1 s2 s5 sp95 latency, log scale →Gemini 3.8 Flash: Decision Score 74.1 · p95 3.9 sGemini 3.8 FlashJev: Decision Score 67.8 · p95 693 msJevGLM-5.3: Decision Score 66.8 · p95 3.2 sGLM-5.3DeepSeek V4.1 Flash: Decision Score 63.9 · p95 1.3 sDeepSeek V4.1 FlashQwen3.8 Flash: Decision Score 62.4 · p95 3.1 sQwen3.8 FlashMistral Medium 3.5: Decision Score 59.5 · p95 1.4 sMistral Medium 3.5Mercury 2.5: Decision Score 54.0 · p95 1.5 sMercury 2.5

automation at 95% accuracy

Share of questions each model can answer alone and stay 95% right. It answers only above its confidence cut-off; the rest go to a person.

Jev54%acts at confidence ≥ 0.98
Gemini 3.8 Flash45%acts at confidence ≥ 0.89
DeepSeek V4.1 Flash39%acts at confidence ≥ 0.86
GLM-5.338%acts at confidence ≥ 0.83
Mistral Medium 3.511%acts at confidence ≥ 0.81
Qwen3.8 Flashnever reaches 95%
Mercury 2.5never reaches 95%

release 2026-09-18 · suite 0.1.0 · board.json · permalink · explore every question · how we score