LLM · Mistral AI
Mistral Medium 3.5
Ranks =3 of 7 on noul, =5 of 7 on choice, 4 of 7 on score. Verbalized probabilities through the Jevals adapter, no reasoning, via OpenRouter · Mistral.
58.0Decision Score
rank =3 of 8
- accuracy
- 88.8%
- calibration gap
- 5.2
- hand-off at 95%
- 65%
- $ per 1k
- $0.87
- speed p95
- 917 ms
Said vs right. Squares on the diagonal mean honest confidence; bigger squares hold more answers.
59.5Decision Score
rank 5 of 8
- accuracy
- 74.5%
- calibration gap
- 8.1
- hand-off at 95%
- 11%
- $ per 1k
- $1.5
- speed p95
- 1.4 s
Said vs right. Squares on the diagonal mean honest confidence; bigger squares hold more answers.
−13.7Decision Score
rank 5 of 8
- accuracy
- 43.9%
- calibration gap
- 33.6
- hand-off at 95%
- —
- $ per 1k
- $1.3
- speed p95
- 1.2 s
Said vs right. Squares on the diagonal mean honest confidence; bigger squares hold more answers.
everything behind the numbers
noul
accuracy88.8%95% range 85.3–92.2%
calibration gap5.2 ptslower is better · 14.2% of answers at confidence 1.00
automation at 95%65%of decisions, acting at confidence ≥ 0.91 (right 95.5%)shared gate ≥ 0.91: acts on 65%, right 95.5%
speed533 msmedian · p95 917 ms
stability1.0%of answers change on an identical rerun
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Mistral · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as mistralai/mistral-medium-3-5pubmedqa run log
choice
accuracy74.5%95% range 70.1–78.9%
calibration gap8.1 ptslower is better · 0.0% of answers at confidence 1.00
automation at 95%11%of decisions, acting at confidence ≥ 0.81 (right 99.4%)shared gate ≥ 0.96: acts on 0%
speed936 msmedian · p95 1.4 s
stability2.3%of answers change on an identical rerun · 22.2% change when options are reordered
valid output99.5%0.0% refused · 1,493 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Mistral · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as mistralai/mistral-medium-3-5banking77 run log
score
accuracy43.9%95% range 38.5–49.3%
calibration gap33.6 ptslower is better · 2.1% of answers at confidence 1.00
automation at 95%—never reaches 95% accuracy
speed706 msmedian · p95 1.2 s
stability3.7%of answers change on an identical rerun
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Mistral · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as mistralai/mistral-medium-3-5helpsteer2 run log
- model id
- mistralai/mistral-medium-3-5
- served as
- mistralai/mistral-medium-3-5
- adapter
- jevals-llm 0383a0e3e592 · the prompt
- host
- OpenRouter · Mistral
- price as of
- 2026-09-18
- harness
- 4da1b6d6bc57
- release
- 2026-09-18