# Jev vs Qwen3.8 Flash

Jev vs Qwen3.8 Flash, scored against human labels. Decision Score: noul 69.0 vs 62.4, choice 67.8 vs 62.4 and score 9.2 vs −1.4. Release 2026-09-18.

## noul (PubMedQA)

|  | Jev | Qwen3.8 Flash |
| --- | --- | --- |
| Decision Score (95% range) | 69.0 (59.8 to 76.5) | 62.4 (50.9 to 72.5) |
| Rank | tied for rank 1 of 8 | rank 2 of 8 |
| Accuracy | 91.3% | 89.7% |
| Calibration gap (ECE, pts) | 5.0 | 2.5 |
| Hand-off at 95% | 86% | — |
| $ per 1k decisions | $0.029 | $0.086 |
| Speed p95 | 653 ms | 3.5 s |

Question by question (right = most of 5 runs right), out of 300: both right on 260, only Jev on 14, only Qwen3.8 Flash on 8, neither on 18. Same usual pick on 93%.

## choice (Banking77)

|  | Jev | Qwen3.8 Flash |
| --- | --- | --- |
| Decision Score (95% range) | 67.8 (61.3 to 74.5) | 62.4 (55.9 to 68.9) |
| Rank | tied for rank 2 of 8 | tied for rank 4 of 8 |
| Accuracy | 79.7% | 76.8% |
| Calibration gap (ECE, pts) | 9.8 | 9.5 |
| Hand-off at 95% | 54% | — |
| $ per 1k decisions | $0.043 | $0.12 |
| Speed p95 | 693 ms | 3.1 s |

Question by question (right = most of 5 runs right), out of 300: both right on 218, only Jev on 20, only Qwen3.8 Flash on 11, neither on 51. Same usual pick on 88%.

## score (HelpSteer2 helpfulness)

|  | Jev | Qwen3.8 Flash |
| --- | --- | --- |
| Decision Score (95% range) | 9.2 (−4.4 to 21.0) | −1.4 (−15.4 to 11.5) |
| Rank | tied for rank 1 of 8 | tied for rank 3 of 8 |
| Accuracy | 41.3% | 36.1% |
| Calibration gap (ECE, pts) | 19.7 | 25.8 |
| Hand-off at 95% | — | — |
| $ per 1k decisions | $0.036 | $0.13 |
| Speed p95 | 670 ms | 2.6 s |

Question by question (right = most of 5 runs right), out of 300: both right on 86, only Jev on 38, only Qwen3.8 Flash on 23, neither on 153. Same usual pick on 68%.

Decision Score: 100 = perfect, 0 = guessing the label base rates, below 0 = worse than that. Ranks: 1 + the number of rows significantly better (paired item bootstrap, 95%); rows that cannot be told apart share a rank. Hand-off at 95%: the share of decisions a model can take alone, at its own confidence threshold, while staying at least 95% right.

Source: https://jevals.com/compare/jev-vs-qwen3.8-flash/ · release 2026-09-18, suite 0.1.0 · data (CC-BY-4.0): https://jevals.com/data/releases/2026-09-18/board.json · cite as: Jevals (jevals.com), release 2026-09-18, suite 0.1.0. CC-BY-4.0.
