# Jev vs DeepSeek V4.1 Flash

Jev vs DeepSeek V4.1 Flash, scored against human labels. Decision Score: noul 69.0 vs 47.5, choice 67.8 vs 63.9 and score 9.2 vs −19.0. Release 2026-09-18.

## noul (PubMedQA)

|  | Jev | DeepSeek V4.1 Flash |
| --- | --- | --- |
| Decision Score (95% range) | 69.0 (59.8 to 76.5) | 47.5 (34.7 to 59.1) |
| Rank | tied for rank 1 of 8 | rank 7 of 8 |
| Accuracy | 91.3% | 83.7% |
| Calibration gap (ECE, pts) | 5.0 | 6.2 |
| Hand-off at 95% | 86% | — |
| $ per 1k decisions | $0.029 | $0.078 |
| Speed p95 | 653 ms | 1.1 s |

Question by question (right = most of 5 runs right), out of 300: both right on 249, only Jev on 25, only DeepSeek V4.1 Flash on 4, neither on 22. Same usual pick on 90%.

## choice (Banking77)

|  | Jev | DeepSeek V4.1 Flash |
| --- | --- | --- |
| Decision Score (95% range) | 67.8 (61.3 to 74.5) | 63.9 (58.5 to 69.6) |
| Rank | tied for rank 2 of 8 | tied for rank 4 of 8 |
| Accuracy | 79.7% | 75.9% |
| Calibration gap (ECE, pts) | 9.8 | 2.8 |
| Hand-off at 95% | 54% | 39% |
| $ per 1k decisions | $0.043 | $0.14 |
| Speed p95 | 693 ms | 1.3 s |

Question by question (right = most of 5 runs right), out of 300: both right on 221, only Jev on 17, only DeepSeek V4.1 Flash on 14, neither on 48. Same usual pick on 88%.

## score (HelpSteer2 helpfulness)

|  | Jev | DeepSeek V4.1 Flash |
| --- | --- | --- |
| Decision Score (95% range) | 9.2 (−4.4 to 21.0) | −19.0 (−36.0 to −3.3) |
| Rank | tied for rank 1 of 8 | rank 7 of 8 |
| Accuracy | 41.3% | 34.7% |
| Calibration gap (ECE, pts) | 19.7 | 33.5 |
| Hand-off at 95% | — | — |
| $ per 1k decisions | $0.036 | $0.13 |
| Speed p95 | 670 ms | 1.2 s |

Question by question (right = most of 5 runs right), out of 300: both right on 74, only Jev on 50, only DeepSeek V4.1 Flash on 26, neither on 150. Same usual pick on 61%.

Decision Score: 100 = perfect, 0 = guessing the label base rates, below 0 = worse than that. Ranks: 1 + the number of rows significantly better (paired item bootstrap, 95%); rows that cannot be told apart share a rank. Hand-off at 95%: the share of decisions a model can take alone, at its own confidence threshold, while staying at least 95% right.

Source: https://jevals.com/compare/jev-vs-deepseek-v4.1-flash/ · release 2026-09-18, suite 0.1.0 · data (CC-BY-4.0): https://jevals.com/data/releases/2026-09-18/board.json · cite as: Jevals (jevals.com), release 2026-09-18, suite 0.1.0. CC-BY-4.0.
