LLM · Google
Gemini 3.8 Flash
Ranks 2 of 7 on noul, 2 of 7 on choice, =5 of 7 on score. Verbalized probabilities through the Jevals adapter, reasoning low, via OpenRouter · Google AI Studio.
73.0Decision Score
rank =1 of 8
- accuracy
- 92.5%
- calibration gap
- 2.0
- hand-off at 95%
- 79%
- $ per 1k
- $0.8
- speed p95
- 5.5 s
Said vs right. Squares on the diagonal mean honest confidence; bigger squares hold more answers.
74.1Decision Score
rank 1 of 8
- accuracy
- 84.6%
- calibration gap
- 3.2
- hand-off at 95%
- 45%
- $ per 1k
- $1.4
- speed p95
- 3.9 s
Said vs right. Squares on the diagonal mean honest confidence; bigger squares hold more answers.
4.6Decision Score
rank =1 of 8
- accuracy
- 42.4%
- calibration gap
- 22.6
- hand-off at 95%
- —
- $ per 1k
- $1.3
- speed p95
- 4.0 s
Said vs right. Squares on the diagonal mean honest confidence; bigger squares hold more answers.
everything behind the numbers
noul
accuracy92.5%95% range 89.3–95.3%
calibration gap2.0 ptslower is better · 0.8% of answers at confidence 1.00
automation at 95%79%of decisions, acting at confidence ≥ 0.86 (right 96.6%)shared gate ≥ 0.91: acts on 73%, right 97.1%
speed1.9 smedian · p95 5.5 s
stability1.7%of answers change on an identical rerun
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Google AI Studio · reasoning low (mandatory) · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as google/gemini-3.8-flashpubmedqa run log
choice
accuracy84.6%95% range 80.8–88.2%
calibration gap3.2 ptslower is better · 0.0% of answers at confidence 1.00
automation at 95%45%of decisions, acting at confidence ≥ 0.89 (right 96.1%)shared gate ≥ 0.96: acts on 12%, right 99.4%
speed1.6 smedian · p95 3.9 s
stability3.0%of answers change on an identical rerun · 8.7% change when options are reordered
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Google AI Studio · reasoning low (mandatory) · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as google/gemini-3.8-flashbanking77 run log
score
accuracy42.4%95% range 36.9–47.5%
calibration gap22.6 ptslower is better · 0.1% of answers at confidence 1.00
automation at 95%—never reaches 95% accuracy
speed1.9 smedian · p95 4.0 s
stability11.7%of answers change on an identical rerun
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Google AI Studio · reasoning low (mandatory) · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as google/gemini-3.8-flashhelpsteer2 run log
- model id
- google/gemini-3.8-flash
- served as
- google/gemini-3.8-flash
- adapter
- jevals-llm 0383a0e3e592 · the prompt
- host
- OpenRouter · Google AI Studio
- price as of
- 2026-09-18
- harness
- 4da1b6d6bc57
- release
- 2026-09-18