Skip to content

You are looking at the frozen release 2026-09-18. See the latest boards

yes or no benchmark · release 2026-09-18

noulif

A yes/no question. The model returns P(yes), and your code acts on it with an if.

the test

Given the context passages from a biomedical abstract, is the answer to the research question yes?

PubMedQA · 2 options · 300 questions × 5 runs · human labels

ranking

System One LLM baseline top 2 for that metric (score, cost, speed)

Ordered by wins (top-2 finishes across score, accuracy, calibration gap, cost and speed), then Decision Score. Select a model for details; select a column header to sort.
model95% range of the Decision Scorewins
1
Jev
System One TypeSafe AI
69.091.3%5.0$0.029653 ms4
accuracy91.3%95% range 87.994.3%
calibration gap5.0 ptslower is better · 0.0% of answers at confidence 1.00
10said 55%, right 58% (48 answers)said 65%, right 73% (109 answers)said 75%, right 82% (170 answers)said 86%, right 91% (371 answers)said 94%, right 98% (802 answers)
automation at 95%86%of decisions, acting at confidence ≥ 0.75 (right 95.0%)shared gate ≥ 0.91: acts on 49%, right 98.6%
speed438 msmedian · p95 653 ms
stability0.0%of answers change on an identical rerun
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupai-sdk experimental_evaluate ai@7.0.106 · native probabilities via Vercel AI Gateway · reasoning · temperature · 300 items × 5 runs · 2026-09-18 · served as typesafe-ai/jevpubmedqa run log
2
Gemini 3.8 Flash
LLM Google · reasoning low
73.092.5%2.0$0.85.5 s3
accuracy92.5%95% range 89.395.3%
calibration gap2.0 ptslower is better · 0.8% of answers at confidence 1.00
10said 55%, right 50% (2 answers)said 63%, right 100% (3 answers)said 73%, right 57% (7 answers)said 85%, right 78% (306 answers)said 96%, right 97% (1182 answers)
automation at 95%79%of decisions, acting at confidence ≥ 0.86 (right 96.6%)shared gate ≥ 0.91: acts on 73%, right 97.1%
speed1.9 smedian · p95 5.5 s
stability1.7%of answers change on an identical rerun
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Google AI Studio · reasoning low (mandatory) · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as google/gemini-3.8-flashpubmedqa run log
=3
GLM-5.3
LLM Z.ai · reasoning low
60.688.7%2.0$0.512.9 s1
accuracy88.7%95% range 85.391.6%
calibration gap2.0 ptslower is better · 0.3% of answers at confidence 1.00
10said 55%, right 67% (3 answers)said 61%, right 71% (34 answers)said 72%, right 70% (99 answers)said 84%, right 80% (372 answers)said 94%, right 95% (992 answers)
automation at 95%45%of decisions, acting at confidence ≥ 0.91 (right 96.7%)shared gate ≥ 0.91: acts on 45%, right 96.7%
speed1.8 smedian · p95 2.9 s
stability5.3%of answers change on an identical rerun
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Novita (fp8) · reasoning low (mandatory) · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as z-ai/glm-5.3pubmedqa run log
=3
Mistral Medium 3.5
LLM Mistral AI · no reasoning
58.088.8%5.2$0.87917 ms1
accuracy88.8%95% range 85.392.2%
calibration gap5.2 ptslower is better · 14.2% of answers at confidence 1.00
10said 61%, right 90% (41 answers)said 70%, right 73% (49 answers)said 80%, right 77% (83 answers)said 95%, right 90% (1327 answers)
automation at 95%65%of decisions, acting at confidence ≥ 0.91 (right 95.5%)shared gate ≥ 0.91: acts on 65%, right 95.5%
speed533 msmedian · p95 917 ms
stability1.0%of answers change on an identical rerun
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Mistral · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as mistralai/mistral-medium-3-5pubmedqa run log
=3
Mercury 2.5
LLM Inception · no reasoning
55.787.1%4.7$0.0231.0 s1
accuracy87.1%95% range 84.090.1%
calibration gap4.7 ptslower is better · 1.1% of answers at confidence 1.00
10said 54%, right 20% (10 answers)said 63%, right 63% (68 answers)said 73%, right 69% (67 answers)said 84%, right 77% (187 answers)said 96%, right 92% (1168 answers)
automation at 95%19%of decisions, acting at confidence ≥ 0.99 (right 95.8%)shared gate ≥ 0.91: acts on 72%, right 92.5%
speed584 msmedian · p95 1.0 s
stability10.0%of answers change on an identical rerun
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Inception · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as inception/mercury-2.5pubmedqa run log
=6
Qwen3.8 Flash
LLM Alibaba · no reasoning
62.489.7%2.5$0.0863.5 s0
accuracy89.7%95% range 86.392.7%
calibration gap2.5 ptslower is better · 0.7% of answers at confidence 1.00
10said 52%, right 0% (1 answers)said 65%, right 70% (30 answers)said 73%, right 80% (10 answers)said 85%, right 79% (332 answers)said 95%, right 94% (1127 answers)
automation at 95%never reaches 95% accuracyshared gate ≥ 0.91: acts on 72%, right 93.9%
speed1.1 smedian · p95 3.5 s
stability4.7%of answers change on an identical rerun
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Alibaba · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as qwen/qwen3.8-flashpubmedqa run log
=6
DeepSeek V4.1 Flash
LLM DeepSeek · no reasoning
47.583.7%6.2$0.0781.1 s0
accuracy83.7%95% range 79.987.3%
calibration gap6.2 ptslower is better · 19.5% of answers at confidence 1.00
10said 52%, right 14% (7 answers)said 62%, right 41% (51 answers)said 71%, right 66% (145 answers)said 84%, right 72% (304 answers)said 96%, right 92% (993 answers)
automation at 95%never reaches 95% accuracyshared gate ≥ 0.91: acts on 54%, right 93.0%
speed838 msmedian · p95 1.1 s
stability6.7%of answers change on an identical rerun
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · DeepSeek · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as deepseek/deepseek-v4.1-flashpubmedqa run log
Label prior
Baseline
← 0 = guessing, off the scale
0.062.0%0.0
accuracy62.0%95% range 56.767.3%
calibration gap0.0 ptslower is better · 0.0% of answers at confidence 1.00
automation at 95%no probabilitiesshared gate ≥ 0.91: acts on 0%
speedmedian · p95
stability0.0%of answers change on an identical rerun
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupconstant probabilities · reasoning · temperature · 300 items × 5 runs

Score: 100 = perfect, 0 = no better than guessing. The bar is the 95% range. Select a model to see everything behind its row.

Wins: a model wins a column when it is in the top two for that column (score, accuracy, calibration gap, cost or speed). The table is sorted by most wins, then by score. =3 means tied on wins.

score vs price

Up and to the left is better. The line joins the best value at each price.
45505560657075$0.02$0.05$0.1$0.2$0.5$1$ per 1k decisions, log scale →Gemini 3.8 Flash: Decision Score 73.0 · $0.8 per 1k decisionsGemini 3.8 FlashJev: Decision Score 69.0 · $0.029 per 1k decisionsJevQwen3.8 Flash: Decision Score 62.4 · $0.086 per 1k decisionsQwen3.8 FlashGLM-5.3: Decision Score 60.6 · $0.51 per 1k decisionsGLM-5.3Mistral Medium 3.5: Decision Score 58.0 · $0.87 per 1k decisionsMistral Medium 3.5Mercury 2.5: Decision Score 55.7 · $0.023 per 1k decisionsMercury 2.5DeepSeek V4.1 Flash: Decision Score 47.5 · $0.078 per 1k decisionsDeepSeek V4.1 Flash

score vs speed

Up and to the left is better. Measured end to end, one question per request.
45505560657075500 ms1 s2 s5 sp95 latency, log scale →Gemini 3.8 Flash: Decision Score 73.0 · p95 5.5 sGemini 3.8 FlashJev: Decision Score 69.0 · p95 653 msJevQwen3.8 Flash: Decision Score 62.4 · p95 3.5 sQwen3.8 FlashGLM-5.3: Decision Score 60.6 · p95 2.9 sGLM-5.3Mistral Medium 3.5: Decision Score 58.0 · p95 917 msMistral Medium 3.5Mercury 2.5: Decision Score 55.7 · p95 1.0 sMercury 2.5DeepSeek V4.1 Flash: Decision Score 47.5 · p95 1.1 sDeepSeek V4.1 Flash

automation at 95% accuracy

Share of questions each model can answer alone and stay 95% right. It answers only above its confidence cut-off; the rest go to a person.

Jev86%acts at confidence ≥ 0.75
Gemini 3.8 Flash79%acts at confidence ≥ 0.86
Mistral Medium 3.565%acts at confidence ≥ 0.91
GLM-5.345%acts at confidence ≥ 0.91
Mercury 2.519%acts at confidence ≥ 0.99
Qwen3.8 Flashnever reaches 95%
DeepSeek V4.1 Flashnever reaches 95%

release 2026-09-18 · suite 0.1.0 · board.json · permalink · explore every question · how we score