You are looking at the frozen release 2026-09-18. See the latest boards
yes or no benchmark · release 2026-09-18
noul → if
A yes/no question. The model returns P(yes), and your code acts on it with an if.
“Given the context passages from a biomedical abstract, is the answer to the research question yes?”
ranking
System One LLM baseline top 2 for that metric (score, cost, speed)
| model | 95% range of the Decision Score | wins | ||||||
|---|---|---|---|---|---|---|---|---|
| 1 | Jev | 69.0 | 91.3% | 5.0 | $0.029 | 653 ms | 4 | |
accuracy91.3%95% range 87.9–94.3% calibration gap5.0 ptslower is better · 0.0% of answers at confidence 1.00 automation at 95%86%of decisions, acting at confidence ≥ 0.75 (right 95.0%)shared gate ≥ 0.91: acts on 49%, right 98.6% speed438 msmedian · p95 653 ms stability0.0%of answers change on an identical rerun valid output100.0%0.0% refused · 1,500 of 1,500 answered setupai-sdk experimental_evaluate ai@7.0.106 · native probabilities via Vercel AI Gateway · reasoning — · temperature — · 300 items × 5 runs · 2026-09-18 · served as typesafe-ai/jevpubmedqa run log | ||||||||
| 2 | Gemini 3.8 Flash | 73.0 | 92.5% | 2.0 | $0.8 | 5.5 s | 3 | |
accuracy92.5%95% range 89.3–95.3% calibration gap2.0 ptslower is better · 0.8% of answers at confidence 1.00 automation at 95%79%of decisions, acting at confidence ≥ 0.86 (right 96.6%)shared gate ≥ 0.91: acts on 73%, right 97.1% speed1.9 smedian · p95 5.5 s stability1.7%of answers change on an identical rerun valid output100.0%0.0% refused · 1,500 of 1,500 answered setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Google AI Studio · reasoning low (mandatory) · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as google/gemini-3.8-flashpubmedqa run log | ||||||||
| =3 | GLM-5.3 | 60.6 | 88.7% | 2.0 | $0.51 | 2.9 s | 1 | |
accuracy88.7%95% range 85.3–91.6% calibration gap2.0 ptslower is better · 0.3% of answers at confidence 1.00 automation at 95%45%of decisions, acting at confidence ≥ 0.91 (right 96.7%)shared gate ≥ 0.91: acts on 45%, right 96.7% speed1.8 smedian · p95 2.9 s stability5.3%of answers change on an identical rerun valid output100.0%0.0% refused · 1,500 of 1,500 answered setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Novita (fp8) · reasoning low (mandatory) · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as z-ai/glm-5.3pubmedqa run log | ||||||||
| =3 | Mistral Medium 3.5 | 58.0 | 88.8% | 5.2 | $0.87 | 917 ms | 1 | |
accuracy88.8%95% range 85.3–92.2% calibration gap5.2 ptslower is better · 14.2% of answers at confidence 1.00 automation at 95%65%of decisions, acting at confidence ≥ 0.91 (right 95.5%)shared gate ≥ 0.91: acts on 65%, right 95.5% speed533 msmedian · p95 917 ms stability1.0%of answers change on an identical rerun valid output100.0%0.0% refused · 1,500 of 1,500 answered setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Mistral · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as mistralai/mistral-medium-3-5pubmedqa run log | ||||||||
| =3 | Mercury 2.5 | 55.7 | 87.1% | 4.7 | $0.023 | 1.0 s | 1 | |
accuracy87.1%95% range 84.0–90.1% calibration gap4.7 ptslower is better · 1.1% of answers at confidence 1.00 automation at 95%19%of decisions, acting at confidence ≥ 0.99 (right 95.8%)shared gate ≥ 0.91: acts on 72%, right 92.5% speed584 msmedian · p95 1.0 s stability10.0%of answers change on an identical rerun valid output100.0%0.0% refused · 1,500 of 1,500 answered setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Inception · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as inception/mercury-2.5pubmedqa run log | ||||||||
| =6 | Qwen3.8 Flash | 62.4 | 89.7% | 2.5 | $0.086 | 3.5 s | 0 | |
accuracy89.7%95% range 86.3–92.7% calibration gap2.5 ptslower is better · 0.7% of answers at confidence 1.00 automation at 95%—never reaches 95% accuracyshared gate ≥ 0.91: acts on 72%, right 93.9% speed1.1 smedian · p95 3.5 s stability4.7%of answers change on an identical rerun valid output100.0%0.0% refused · 1,500 of 1,500 answered setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Alibaba · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as qwen/qwen3.8-flashpubmedqa run log | ||||||||
| =6 | DeepSeek V4.1 Flash | 47.5 | 83.7% | 6.2 | $0.078 | 1.1 s | 0 | |
accuracy83.7%95% range 79.9–87.3% calibration gap6.2 ptslower is better · 19.5% of answers at confidence 1.00 automation at 95%—never reaches 95% accuracyshared gate ≥ 0.91: acts on 54%, right 93.0% speed838 msmedian · p95 1.1 s stability6.7%of answers change on an identical rerun valid output100.0%0.0% refused · 1,500 of 1,500 answered setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · DeepSeek · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as deepseek/deepseek-v4.1-flashpubmedqa run log | ||||||||
| — | Label prior | 0.0 | 62.0% | 0.0 | — | — | — | |
accuracy62.0%95% range 56.7–67.3% calibration gap0.0 ptslower is better · 0.0% of answers at confidence 1.00 automation at 95%—no probabilitiesshared gate ≥ 0.91: acts on 0% speed—median · p95 — stability0.0%of answers change on an identical rerun valid output100.0%0.0% refused · 1,500 of 1,500 answered setupconstant probabilities · reasoning — · temperature — · 300 items × 5 runs | ||||||||
Score: 100 = perfect, 0 = no better than guessing. The bar is the 95% range. Select a model to see everything behind its row.
Wins: a model wins a column when it is in the top two for that column (score, accuracy, calibration gap, cost or speed). The table is sorted by most wins, then by score. =3 means tied on wins.
score vs price
score vs speed
automation at 95% accuracy
Share of questions each model can answer alone and stay 95% right. It answers only above its confidence cut-off; the rest go to a person.
release 2026-09-18 · suite 0.1.0 · board.json · permalink · explore every question · how we score