LLM · DeepSeek
DeepSeek V4.1 Flash
Ranks =6 of 7 on noul, 3 of 7 on choice, =5 of 7 on score. Verbalized probabilities through the Jevals adapter, no reasoning, via OpenRouter · DeepSeek.
47.5Decision Score
rank 7 of 8
- accuracy
- 83.7%
- calibration gap
- 6.2
- hand-off at 95%
- —
- $ per 1k
- $0.078
- speed p95
- 1.1 s
Said vs right. Squares on the diagonal mean honest confidence; bigger squares hold more answers.
63.9Decision Score
rank =4 of 8
- accuracy
- 75.9%
- calibration gap
- 2.8
- hand-off at 95%
- 39%
- $ per 1k
- $0.14
- speed p95
- 1.3 s
Said vs right. Squares on the diagonal mean honest confidence; bigger squares hold more answers.
−19.0Decision Score
rank 7 of 8
- accuracy
- 34.7%
- calibration gap
- 33.5
- hand-off at 95%
- —
- $ per 1k
- $0.13
- speed p95
- 1.2 s
Said vs right. Squares on the diagonal mean honest confidence; bigger squares hold more answers.
everything behind the numbers
noul
accuracy83.7%95% range 79.9–87.3%
calibration gap6.2 ptslower is better · 19.5% of answers at confidence 1.00
automation at 95%—never reaches 95% accuracyshared gate ≥ 0.91: acts on 54%, right 93.0%
speed838 msmedian · p95 1.1 s
stability6.7%of answers change on an identical rerun
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · DeepSeek · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as deepseek/deepseek-v4.1-flashpubmedqa run log
choice
accuracy75.9%95% range 71.9–80.1%
calibration gap2.8 ptslower is better · 0.6% of answers at confidence 1.00
automation at 95%39%of decisions, acting at confidence ≥ 0.86 (right 95.8%)shared gate ≥ 0.96: acts on 15%, right 99.1%
speed999 msmedian · p95 1.3 s
stability10.7%of answers change on an identical rerun · 30.2% change when options are reordered
valid output99.9%0.0% refused · 1,498 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · DeepSeek · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as deepseek/deepseek-v4.1-flashbanking77 run log
score
accuracy34.7%95% range 30.5–38.7%
calibration gap33.5 ptslower is better · 1.5% of answers at confidence 1.00
automation at 95%—never reaches 95% accuracy
speed912 msmedian · p95 1.2 s
stability34.3%of answers change on an identical rerun
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · DeepSeek · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as deepseek/deepseek-v4.1-flashhelpsteer2 run log
- model id
- deepseek/deepseek-v4.1-flash
- served as
- deepseek/deepseek-v4.1-flash
- adapter
- jevals-llm 0383a0e3e592 · the prompt
- host
- OpenRouter · DeepSeek
- price as of
- 2026-09-18
- harness
- 4da1b6d6bc57
- release
- 2026-09-18