# Jev vs Mercury 2.5

Jev vs Mercury 2.5, scored against human labels. Decision Score: noul 69.0 vs 55.7, choice 67.8 vs 54.0 and score 9.2 vs −5.5. Release 2026-09-18.

## noul (PubMedQA)

|  | Jev | Mercury 2.5 |
| --- | --- | --- |
| Decision Score (95% range) | 69.0 (59.8 to 76.5) | 55.7 (44.8 to 65.4) |
| Rank | tied for rank 1 of 8 | tied for rank 3 of 8 |
| Accuracy | 91.3% | 87.1% |
| Calibration gap (ECE, pts) | 5.0 | 4.7 |
| Hand-off at 95% | 86% | 19% |
| $ per 1k decisions | $0.029 | $0.023 |
| Speed p95 | 653 ms | 1.0 s |

Question by question (right = most of 5 runs right), out of 300: both right on 257, only Jev on 17, only Mercury 2.5 on 7, neither on 19. Same usual pick on 92%.

## choice (Banking77)

|  | Jev | Mercury 2.5 |
| --- | --- | --- |
| Decision Score (95% range) | 67.8 (61.3 to 74.5) | 54.0 (47.8 to 60.6) |
| Rank | tied for rank 2 of 8 | rank 7 of 8 |
| Accuracy | 79.7% | 70.3% |
| Calibration gap (ECE, pts) | 9.8 | 10.4 |
| Hand-off at 95% | 54% | — |
| $ per 1k decisions | $0.043 | $0.036 |
| Speed p95 | 693 ms | 1.5 s |

Question by question (right = most of 5 runs right), out of 300: both right on 209, only Jev on 29, only Mercury 2.5 on 7, neither on 55. Same usual pick on 83%.

## score (HelpSteer2 helpfulness)

|  | Jev | Mercury 2.5 |
| --- | --- | --- |
| Decision Score (95% range) | 9.2 (−4.4 to 21.0) | −5.5 (−16.2 to 4.0) |
| Rank | tied for rank 1 of 8 | tied for rank 3 of 8 |
| Accuracy | 41.3% | 41.7% |
| Calibration gap (ECE, pts) | 19.7 | 31.5 |
| Hand-off at 95% | — | — |
| $ per 1k decisions | $0.036 | $0.034 |
| Speed p95 | 670 ms | 1.2 s |

Question by question (right = most of 5 runs right), out of 300: both right on 67, only Jev on 57, only Mercury 2.5 on 61, neither on 115. Same usual pick on 48%.

Decision Score: 100 = perfect, 0 = guessing the label base rates, below 0 = worse than that. Ranks: 1 + the number of rows significantly better (paired item bootstrap, 95%); rows that cannot be told apart share a rank. Hand-off at 95%: the share of decisions a model can take alone, at its own confidence threshold, while staying at least 95% right.

Source: https://jevals.com/compare/jev-vs-mercury-2.5/ · release 2026-09-18, suite 0.1.0 · data (CC-BY-4.0): https://jevals.com/data/releases/2026-09-18/board.json · cite as: Jevals (jevals.com), release 2026-09-18, suite 0.1.0. CC-BY-4.0.
