LLM · Alibaba
Qwen3.8 Flash
Ranks =6 of 7 on noul, =5 of 7 on choice, =5 of 7 on score. Verbalized probabilities through the Jevals adapter, no reasoning, via OpenRouter · Alibaba.
62.4Decision Score
rank 2 of 8
- accuracy
- 89.7%
- calibration gap
- 2.5
- hand-off at 95%
- —
- $ per 1k
- $0.086
- speed p95
- 3.5 s
Said vs right. Squares on the diagonal mean honest confidence; bigger squares hold more answers.
62.4Decision Score
rank =4 of 8
- accuracy
- 76.8%
- calibration gap
- 9.5
- hand-off at 95%
- —
- $ per 1k
- $0.12
- speed p95
- 3.1 s
Said vs right. Squares on the diagonal mean honest confidence; bigger squares hold more answers.
−1.4Decision Score
rank =3 of 8
- accuracy
- 36.1%
- calibration gap
- 25.8
- hand-off at 95%
- —
- $ per 1k
- $0.13
- speed p95
- 2.6 s
Said vs right. Squares on the diagonal mean honest confidence; bigger squares hold more answers.
everything behind the numbers
noul
accuracy89.7%95% range 86.3–92.7%
calibration gap2.5 ptslower is better · 0.7% of answers at confidence 1.00
automation at 95%—never reaches 95% accuracyshared gate ≥ 0.91: acts on 72%, right 93.9%
speed1.1 smedian · p95 3.5 s
stability4.7%of answers change on an identical rerun
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Alibaba · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as qwen/qwen3.8-flashpubmedqa run log
choice
accuracy76.8%95% range 72.6–80.9%
calibration gap9.5 ptslower is better · 0.0% of answers at confidence 1.00
automation at 95%—never reaches 95% accuracyshared gate ≥ 0.96: acts on 2%, right 100.0%
speed1.7 smedian · p95 3.1 s
stability9.0%of answers change on an identical rerun · 22.1% change when options are reordered
valid output99.9%0.0% refused · 1,498 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Alibaba · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as qwen/qwen3.8-flashbanking77 run log
score
accuracy36.1%95% range 31.5–40.7%
calibration gap25.8 ptslower is better · 0.0% of answers at confidence 1.00
automation at 95%—never reaches 95% accuracy
speed1.4 smedian · p95 2.6 s
stability26.0%of answers change on an identical rerun
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Alibaba · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as qwen/qwen3.8-flashhelpsteer2 run log
- model id
- qwen/qwen3.8-flash
- served as
- qwen/qwen3.8-flash
- adapter
- jevals-llm 0383a0e3e592 · the prompt
- host
- OpenRouter · Alibaba
- price as of
- 2026-09-18
- harness
- 4da1b6d6bc57, 0401f0112fa8
- release
- 2026-09-18