LLM · Inception
Mercury 2.5
Ranks =3 of 7 on noul, 4 of 7 on choice, =2 of 7 on score. Verbalized probabilities through the Jevals adapter, no reasoning, via OpenRouter · Inception.
55.7Decision Score
rank =3 of 8
- accuracy
- 87.1%
- calibration gap
- 4.7
- hand-off at 95%
- 19%
- $ per 1k
- $0.023
- speed p95
- 1.0 s
Said vs right. Squares on the diagonal mean honest confidence; bigger squares hold more answers.
54.0Decision Score
rank 7 of 8
- accuracy
- 70.3%
- calibration gap
- 10.4
- hand-off at 95%
- —
- $ per 1k
- $0.036
- speed p95
- 1.5 s
Said vs right. Squares on the diagonal mean honest confidence; bigger squares hold more answers.
−5.5Decision Score
rank =3 of 8
- accuracy
- 41.7%
- calibration gap
- 31.5
- hand-off at 95%
- —
- $ per 1k
- $0.034
- speed p95
- 1.2 s
Said vs right. Squares on the diagonal mean honest confidence; bigger squares hold more answers.
everything behind the numbers
noul
accuracy87.1%95% range 84.0–90.1%
calibration gap4.7 ptslower is better · 1.1% of answers at confidence 1.00
automation at 95%19%of decisions, acting at confidence ≥ 0.99 (right 95.8%)shared gate ≥ 0.91: acts on 72%, right 92.5%
speed584 msmedian · p95 1.0 s
stability10.0%of answers change on an identical rerun
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Inception · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as inception/mercury-2.5pubmedqa run log
choice
accuracy70.3%95% range 66.1–74.7%
calibration gap10.4 ptslower is better · 0.0% of answers at confidence 1.00
automation at 95%—never reaches 95% accuracyshared gate ≥ 0.96: acts on 1%, right 100.0%
speed638 msmedian · p95 1.5 s
stability13.8%of answers change on an identical rerun · 36.1% change when options are reordered
valid output99.7%0.0% refused · 1,495 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Inception · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as inception/mercury-2.5banking77 run log
score
accuracy41.7%95% range 37.3–46.2%
calibration gap31.5 ptslower is better · 0.1% of answers at confidence 1.00
automation at 95%—never reaches 95% accuracy
speed618 msmedian · p95 1.2 s
stability38.7%of answers change on an identical rerun
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Inception · reasoning none · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as inception/mercury-2.5helpsteer2 run log
- model id
- inception/mercury-2.5
- served as
- inception/mercury-2.5
- adapter
- jevals-llm 0383a0e3e592 · the prompt
- host
- OpenRouter · Inception
- price as of
- 2026-09-18
- harness
- 4da1b6d6bc57
- release
- 2026-09-18