pick one of N benchmark · release 2026-09-18
choice → match
Pick one option from a list. The model returns a probability for every option, like a match over its cases.
“Which intent does this banking customer's message express?”
ranking
System One LLM baseline top 2 for that metric (score, cost, speed)
| model | 95% range of the Decision Score | wins | ||||||
|---|---|---|---|---|---|---|---|---|
| 1 | Jev | 67.8 | 79.7% | 9.8 | $0.043 | 693 ms | 4 | |
accuracy79.7%95% range 75.5–83.8% calibration gap9.8 ptslower is better · 41.5% of answers at confidence 1.00 automation at 95%54%of decisions, acting at confidence ≥ 0.98 (right 95.4%)shared gate ≥ 0.96: acts on 59%, right 94.3% speed467 msmedian · p95 693 ms stability2.7%of answers change on an identical rerun · 10.3% change when options are reordered valid output100.0%0.0% refused · 1,500 of 1,500 answered setupai-sdk experimental_evaluate ai@7.0.106 · native probabilities via Vercel AI Gateway · reasoning — · temperature — · 300 items × 5 runs · 2026-09-18 · served as typesafe-ai/jevmodel pagebanking77 run log | ||||||||
| 2 | Gemini 3.8 Flash | 74.1 | 84.6% | 3.2 | $1.4 | 3.9 s | 3 | |
accuracy84.6%95% range 80.8–88.2% calibration gap3.2 ptslower is better · 0.0% of answers at confidence 1.00 automation at 95%45%of decisions, acting at confidence ≥ 0.89 (right 96.1%)shared gate ≥ 0.96: acts on 12%, right 99.4% speed1.6 smedian · p95 3.9 s stability3.0%of answers change on an identical rerun · 8.7% change when options are reordered valid output100.0%0.0% refused · 1,500 of 1,500 answered setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Google AI Studio · reasoning low (mandatory) · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as google/gemini-3.8-flashmodel pagebanking77 run log | ||||||||
| 3 | DeepSeek V4.1 Flash | 63.9 | 75.9% | 2.8 | $0.14 | 1.3 s | 2 | |
accuracy75.9%95% range 71.9–80.1% calibration gap2.8 ptslower is better · 0.6% of answers at confidence 1.00 automation at 95%39%of decisions, acting at confidence ≥ 0.86 (right 95.8%)shared gate ≥ 0.96: acts on 15%, right 99.1% speed999 msmedian · p95 1.3 s stability10.7%of answers change on an identical rerun · 30.2% change when options are reordered valid output99.9%0.0% refused · 1,498 of 1,500 answered setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · DeepSeek · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as deepseek/deepseek-v4.1-flashmodel pagebanking77 run log | ||||||||
| 4 | Mercury 2.5 | 54.0 | 70.3% | 10.4 | $0.036 | 1.5 s | 1 | |
accuracy70.3%95% range 66.1–74.7% calibration gap10.4 ptslower is better · 0.0% of answers at confidence 1.00 automation at 95%—never reaches 95% accuracyshared gate ≥ 0.96: acts on 1%, right 100.0% speed638 msmedian · p95 1.5 s stability13.8%of answers change on an identical rerun · 36.1% change when options are reordered valid output99.7%0.0% refused · 1,495 of 1,500 answered setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Inception · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as inception/mercury-2.5model pagebanking77 run log | ||||||||
| =5 | GLM-5.3 | 66.8 | 78.9% | 6.2 | $0.72 | 3.2 s | 0 | |
accuracy78.9%95% range 74.9–82.7% calibration gap6.2 ptslower is better · 0.0% of answers at confidence 1.00 automation at 95%38%of decisions, acting at confidence ≥ 0.83 (right 95.1%)shared gate ≥ 0.96: acts on 0%, right 100.0% speed2.1 smedian · p95 3.2 s stability9.7%of answers change on an identical rerun · 23.7% change when options are reordered valid output100.0%0.0% refused · 1,500 of 1,500 answered setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Novita (fp8) · reasoning low (mandatory) · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as z-ai/glm-5.3model pagebanking77 run log | ||||||||
| =5 | Qwen3.8 Flash | 62.4 | 76.8% | 9.5 | $0.12 | 3.1 s | 0 | |
accuracy76.8%95% range 72.6–80.9% calibration gap9.5 ptslower is better · 0.0% of answers at confidence 1.00 automation at 95%—never reaches 95% accuracyshared gate ≥ 0.96: acts on 2%, right 100.0% speed1.7 smedian · p95 3.1 s stability9.0%of answers change on an identical rerun · 22.1% change when options are reordered valid output99.9%0.0% refused · 1,498 of 1,500 answered setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Alibaba · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as qwen/qwen3.8-flashmodel pagebanking77 run log | ||||||||
| =5 | Mistral Medium 3.5 | 59.5 | 74.5% | 8.1 | $1.5 | 1.4 s | 0 | |
accuracy74.5%95% range 70.1–78.9% calibration gap8.1 ptslower is better · 0.0% of answers at confidence 1.00 automation at 95%11%of decisions, acting at confidence ≥ 0.81 (right 99.4%)shared gate ≥ 0.96: acts on 0% speed936 msmedian · p95 1.4 s stability2.3%of answers change on an identical rerun · 22.2% change when options are reordered valid output99.5%0.0% refused · 1,493 of 1,500 answered setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Mistral · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as mistralai/mistral-medium-3-5model pagebanking77 run log | ||||||||
| — | Label prior | 0.0 | 1.3% | 0.0 | — | — | — | |
accuracy1.3%95% range 0.3–2.7% calibration gap0.0 ptslower is better · 0.0% of answers at confidence 1.00 automation at 95%—no probabilitiesshared gate ≥ 0.96: acts on 0% speed—median · p95 — stability0.0%of answers change on an identical rerun · 0.0% change when options are reordered valid output100.0%0.0% refused · 1,500 of 1,500 answered setupconstant probabilities · reasoning — · temperature — · 300 items × 5 runsmodel page | ||||||||
Score: 100 = perfect, 0 = no better than guessing. The bar is the 95% range. Select a model to see everything behind its row.
Wins: a model wins a column when it is in the top two for that column (score, accuracy, calibration gap, cost or speed). The table is sorted by most wins, then by score. =3 means tied on wins.
score vs price
score vs speed
automation at 95% accuracy
Share of questions each model can answer alone and stay 95% right. It answers only above its confidence cut-off; the rest go to a person.
release 2026-09-18 · suite 0.1.0 · board.json · permalink · explore every question · how we score