Skip to content

LLM · Mistral AI

Mistral Medium 3.5

Ranks =3 of 7 on noul, =5 of 7 on choice, 4 of 7 on score. Verbalized probabilities through the Jevals adapter, no reasoning, via OpenRouter · Mistral.

compare with Jev →

noulif

58.0Decision Score
rank =3 of 8

accuracy
88.8%
calibration gap
5.2
hand-off at 95%
65%
$ per 1k
$0.87
speed p95
917 ms
10said 61%, right 90% (41 answers)said 70%, right 73% (49 answers)said 80%, right 77% (83 answers)said 95%, right 90% (1327 answers)Said vs right. Squares on the diagonal mean honest confidence; bigger squares hold more answers.
choicematch

59.5Decision Score
rank 5 of 8

accuracy
74.5%
calibration gap
8.1
hand-off at 95%
11%
$ per 1k
$1.5
speed p95
1.4 s
10said 40%, right 43% (14 answers)said 50%, right 33% (113 answers)said 60%, right 59% (472 answers)said 70%, right 81% (320 answers)said 80%, right 92% (417 answers)said 92%, right 100% (157 answers)Said vs right. Squares on the diagonal mean honest confidence; bigger squares hold more answers.
scoresort

−13.7Decision Score
rank 5 of 8

accuracy
43.9%
calibration gap
33.6
hand-off at 95%
$ per 1k
$1.3
speed p95
1.2 s
10said 40%, right 0% (5 answers)said 60%, right 42% (90 answers)said 70%, right 35% (578 answers)said 80%, right 46% (474 answers)said 91%, right 56% (353 answers)Said vs right. Squares on the diagonal mean honest confidence; bigger squares hold more answers.

everything behind the numbers

noul
accuracy88.8%95% range 85.392.2%
calibration gap5.2 ptslower is better · 14.2% of answers at confidence 1.00
10said 61%, right 90% (41 answers)said 70%, right 73% (49 answers)said 80%, right 77% (83 answers)said 95%, right 90% (1327 answers)
automation at 95%65%of decisions, acting at confidence ≥ 0.91 (right 95.5%)shared gate ≥ 0.91: acts on 65%, right 95.5%
speed533 msmedian · p95 917 ms
stability1.0%of answers change on an identical rerun
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Mistral · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as mistralai/mistral-medium-3-5pubmedqa run log
choice
accuracy74.5%95% range 70.178.9%
calibration gap8.1 ptslower is better · 0.0% of answers at confidence 1.00
10said 40%, right 43% (14 answers)said 50%, right 33% (113 answers)said 60%, right 59% (472 answers)said 70%, right 81% (320 answers)said 80%, right 92% (417 answers)said 92%, right 100% (157 answers)
automation at 95%11%of decisions, acting at confidence ≥ 0.81 (right 99.4%)shared gate ≥ 0.96: acts on 0%
speed936 msmedian · p95 1.4 s
stability2.3%of answers change on an identical rerun · 22.2% change when options are reordered
valid output99.5%0.0% refused · 1,493 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Mistral · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as mistralai/mistral-medium-3-5banking77 run log
score
accuracy43.9%95% range 38.549.3%
calibration gap33.6 ptslower is better · 2.1% of answers at confidence 1.00
10said 40%, right 0% (5 answers)said 60%, right 42% (90 answers)said 70%, right 35% (578 answers)said 80%, right 46% (474 answers)said 91%, right 56% (353 answers)
automation at 95%never reaches 95% accuracy
speed706 msmedian · p95 1.2 s
stability3.7%of answers change on an identical rerun
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Mistral · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as mistralai/mistral-medium-3-5helpsteer2 run log
model id
mistralai/mistral-medium-3-5
served as
mistralai/mistral-medium-3-5
adapter
jevals-llm 0383a0e3e592 · the prompt
host
OpenRouter · Mistral
price as of
2026-09-18
harness
4da1b6d6bc57
release
2026-09-18

see every question Mistral Medium 3.5 answered →