LLM · Z.ai
GLM-5.3
Ranks =3 of 7 on noul, =5 of 7 on choice, =2 of 7 on score. Verbalized probabilities through the Jevals adapter, reasoning low, via OpenRouter · Novita (fp8).
60.6Decision Score
rank =3 of 8
- accuracy
- 88.7%
- calibration gap
- 2.0
- hand-off at 95%
- 45%
- $ per 1k
- $0.51
- speed p95
- 2.9 s
Said vs right. Squares on the diagonal mean honest confidence; bigger squares hold more answers.
66.8Decision Score
rank =2 of 8
- accuracy
- 78.9%
- calibration gap
- 6.2
- hand-off at 95%
- 38%
- $ per 1k
- $0.72
- speed p95
- 3.2 s
Said vs right. Squares on the diagonal mean honest confidence; bigger squares hold more answers.
7.8Decision Score
rank =1 of 8
- accuracy
- 43.0%
- calibration gap
- 12.9
- hand-off at 95%
- —
- $ per 1k
- $0.78
- speed p95
- 3.6 s
Said vs right. Squares on the diagonal mean honest confidence; bigger squares hold more answers.
everything behind the numbers
noul
accuracy88.7%95% range 85.3–91.6%
calibration gap2.0 ptslower is better · 0.3% of answers at confidence 1.00
automation at 95%45%of decisions, acting at confidence ≥ 0.91 (right 96.7%)shared gate ≥ 0.91: acts on 45%, right 96.7%
speed1.8 smedian · p95 2.9 s
stability5.3%of answers change on an identical rerun
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Novita (fp8) · reasoning low (mandatory) · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as z-ai/glm-5.3pubmedqa run log
choice
accuracy78.9%95% range 74.9–82.7%
calibration gap6.2 ptslower is better · 0.0% of answers at confidence 1.00
automation at 95%38%of decisions, acting at confidence ≥ 0.83 (right 95.1%)shared gate ≥ 0.96: acts on 0%, right 100.0%
speed2.1 smedian · p95 3.2 s
stability9.7%of answers change on an identical rerun · 23.7% change when options are reordered
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Novita (fp8) · reasoning low (mandatory) · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as z-ai/glm-5.3banking77 run log
score
accuracy43.0%95% range 38.6–47.9%
calibration gap12.9 ptslower is better · 0.0% of answers at confidence 1.00
automation at 95%—never reaches 95% accuracy
speed2.2 smedian · p95 3.6 s
stability31.3%of answers change on an identical rerun
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupjevals-llm 0383a0e3e592 · verbalized probabilities via OpenRouter · Novita (fp8) · reasoning low (mandatory) · temperature provider default · 300 items × 5 runs · 2026-09-18 · served as z-ai/glm-5.3helpsteer2 run log
- model id
- z-ai/glm-5.3
- served as
- z-ai/glm-5.3
- adapter
- jevals-llm 0383a0e3e592 · the prompt
- host
- OpenRouter · Novita (fp8)
- price as of
- 2026-09-18
- harness
- 4da1b6d6bc57
- release
- 2026-09-18