# Jev vs Mistral Medium 3.5

Jev vs Mistral Medium 3.5, scored against human labels. Decision Score: noul 69.0 vs 58.0, choice 67.8 vs 59.5 and score 9.2 vs −13.7. Release 2026-09-18.

## noul (PubMedQA)

|  | Jev | Mistral Medium 3.5 |
| --- | --- | --- |
| Decision Score (95% range) | 69.0 (59.8 to 76.5) | 58.0 (45.3 to 69.8) |
| Rank | tied for rank 1 of 8 | tied for rank 3 of 8 |
| Accuracy | 91.3% | 88.8% |
| Calibration gap (ECE, pts) | 5.0 | 5.2 |
| Hand-off at 95% | 86% | 65% |
| $ per 1k decisions | $0.029 | $0.87 |
| Speed p95 | 653 ms | 917 ms |

Question by question (right = most of 5 runs right), out of 300: both right on 256, only Jev on 18, only Mistral Medium 3.5 on 10, neither on 16. Same usual pick on 91%.

## choice (Banking77)

|  | Jev | Mistral Medium 3.5 |
| --- | --- | --- |
| Decision Score (95% range) | 67.8 (61.3 to 74.5) | 59.5 (54.3 to 64.8) |
| Rank | tied for rank 2 of 8 | rank 5 of 8 |
| Accuracy | 79.7% | 74.5% |
| Calibration gap (ECE, pts) | 9.8 | 8.1 |
| Hand-off at 95% | 54% | 11% |
| $ per 1k decisions | $0.043 | $1.5 |
| Speed p95 | 693 ms | 1.4 s |

Question by question (right = most of 5 runs right), out of 300: both right on 212, only Jev on 26, only Mistral Medium 3.5 on 12, neither on 50. Same usual pick on 83%.

## score (HelpSteer2 helpfulness)

|  | Jev | Mistral Medium 3.5 |
| --- | --- | --- |
| Decision Score (95% range) | 9.2 (−4.4 to 21.0) | −13.7 (−28.3 to −0.1) |
| Rank | tied for rank 1 of 8 | rank 5 of 8 |
| Accuracy | 41.3% | 43.9% |
| Calibration gap (ECE, pts) | 19.7 | 33.6 |
| Hand-off at 95% | — | — |
| $ per 1k decisions | $0.036 | $1.3 |
| Speed p95 | 670 ms | 1.2 s |

Question by question (right = most of 5 runs right), out of 300: both right on 63, only Jev on 61, only Mistral Medium 3.5 on 69, neither on 107. Same usual pick on 42%.

Decision Score: 100 = perfect, 0 = guessing the label base rates, below 0 = worse than that. Ranks: 1 + the number of rows significantly better (paired item bootstrap, 95%); rows that cannot be told apart share a rank. Hand-off at 95%: the share of decisions a model can take alone, at its own confidence threshold, while staying at least 95% right.

Source: https://jevals.com/compare/jev-vs-mistral-medium-3.5/ · release 2026-09-18, suite 0.1.0 · data (CC-BY-4.0): https://jevals.com/data/releases/2026-09-18/board.json · cite as: Jevals (jevals.com), release 2026-09-18, suite 0.1.0. CC-BY-4.0.
