compare · release 2026-09-18
Jev vs GLM-5.3
Jev vs GLM-5.3, scored against human labels. Decision Score: noul 69.0 vs 60.6, choice 67.8 vs 66.8 and score 9.2 vs 7.8. Release 2026-09-18.
noul → if · PubMedQA
| Jev | GLM-5.3 | |
|---|---|---|
| Decision Score | 69.095% range 59.8 to 76.5 | 60.695% range 50.0 to 69.8 |
| Rank | tied for rank 1 of 8 | tied for rank 3 of 8 |
| Accuracy | 91.3% | 88.7% |
| Calibration gap | 5.0 pts | 2.0 pts |
| Hand-off at 95% | 86% | 45% |
| $ per 1k decisions | $0.029 | $0.51 |
| Speed p95 | 653 ms | 2.9 s |
Question by question (right = most of 5 runs right), out of 300: both right on 263, only Jev on 11, only GLM-5.3 on 6, neither on 20. Same usual pick on 94%. See where they differ
choice → match · Banking77
| Jev | GLM-5.3 | |
|---|---|---|
| Decision Score | 67.895% range 61.3 to 74.5 | 66.895% range 62.0 to 71.7 |
| Rank | tied for rank 2 of 8 | tied for rank 2 of 8 |
| Accuracy | 79.7% | 78.9% |
| Calibration gap | 9.8 pts | 6.2 pts |
| Hand-off at 95% | 54% | 38% |
| $ per 1k decisions | $0.043 | $0.72 |
| Speed p95 | 693 ms | 3.2 s |
Question by question (right = most of 5 runs right), out of 300: both right on 231, only Jev on 7, only GLM-5.3 on 14, neither on 48. Same usual pick on 91%. See where they differ
score → sort · HelpSteer2 helpfulness
| Jev | GLM-5.3 | |
|---|---|---|
| Decision Score | 9.295% range −4.4 to 21.0 | 7.895% range −4.6 to 18.6 |
| Rank | tied for rank 1 of 8 | tied for rank 1 of 8 |
| Accuracy | 41.3% | 43.0% |
| Calibration gap | 19.7 pts | 12.9 pts |
| Hand-off at 95% | — | — |
| $ per 1k decisions | $0.036 | $0.78 |
| Speed p95 | 670 ms | 3.6 s |
Question by question (right = most of 5 runs right), out of 300: both right on 79, only Jev on 45, only GLM-5.3 on 53, neither on 123. Same usual pick on 54%. See where they differ
Jev · GLM-5.3 · all models · how we score · board.json