# Jev vs GLM-5.3

Jev vs GLM-5.3, scored against human labels. Decision Score: noul 69.0 vs 60.6, choice 67.8 vs 66.8 and score 9.2 vs 7.8. Release 2026-09-18.

## noul (PubMedQA)

|  | Jev | GLM-5.3 |
| --- | --- | --- |
| Decision Score (95% range) | 69.0 (59.8 to 76.5) | 60.6 (50.0 to 69.8) |
| Rank | tied for rank 1 of 8 | tied for rank 3 of 8 |
| Accuracy | 91.3% | 88.7% |
| Calibration gap (ECE, pts) | 5.0 | 2.0 |
| Hand-off at 95% | 86% | 45% |
| $ per 1k decisions | $0.029 | $0.51 |
| Speed p95 | 653 ms | 2.9 s |

Question by question (right = most of 5 runs right), out of 300: both right on 263, only Jev on 11, only GLM-5.3 on 6, neither on 20. Same usual pick on 94%.

## choice (Banking77)

|  | Jev | GLM-5.3 |
| --- | --- | --- |
| Decision Score (95% range) | 67.8 (61.3 to 74.5) | 66.8 (62.0 to 71.7) |
| Rank | tied for rank 2 of 8 | tied for rank 2 of 8 |
| Accuracy | 79.7% | 78.9% |
| Calibration gap (ECE, pts) | 9.8 | 6.2 |
| Hand-off at 95% | 54% | 38% |
| $ per 1k decisions | $0.043 | $0.72 |
| Speed p95 | 693 ms | 3.2 s |

Question by question (right = most of 5 runs right), out of 300: both right on 231, only Jev on 7, only GLM-5.3 on 14, neither on 48. Same usual pick on 91%.

## score (HelpSteer2 helpfulness)

|  | Jev | GLM-5.3 |
| --- | --- | --- |
| Decision Score (95% range) | 9.2 (−4.4 to 21.0) | 7.8 (−4.6 to 18.6) |
| Rank | tied for rank 1 of 8 | tied for rank 1 of 8 |
| Accuracy | 41.3% | 43.0% |
| Calibration gap (ECE, pts) | 19.7 | 12.9 |
| Hand-off at 95% | — | — |
| $ per 1k decisions | $0.036 | $0.78 |
| Speed p95 | 670 ms | 3.6 s |

Question by question (right = most of 5 runs right), out of 300: both right on 79, only Jev on 45, only GLM-5.3 on 53, neither on 123. Same usual pick on 54%.

Decision Score: 100 = perfect, 0 = guessing the label base rates, below 0 = worse than that. Ranks: 1 + the number of rows significantly better (paired item bootstrap, 95%); rows that cannot be told apart share a rank. Hand-off at 95%: the share of decisions a model can take alone, at its own confidence threshold, while staying at least 95% right.

Source: https://jevals.com/compare/jev-vs-glm-5.3/ · release 2026-09-18, suite 0.1.0 · data (CC-BY-4.0): https://jevals.com/data/releases/2026-09-18/board.json · cite as: Jevals (jevals.com), release 2026-09-18, suite 0.1.0. CC-BY-4.0.
