# noul (yes or no) benchmark: Jev vs 6 LLMs

Gemini 3.8 Flash (73.0) and Jev (69.0) tie for first in Decision Score on PubMedQA. Jev costs $0.029 per 1,000 decisions, p95 653 ms. Release 2026-09-18.

A yes/no question. The model returns P(yes), and your code acts on it with an if. The test: “Given the context passages from a biomedical abstract, is the answer to the research question yes?” (PubMedQA, 2 options, 300 questions × 5 runs, human labels, https://huggingface.co/datasets/qiaojin/PubMedQA).

| Rank | Model | Type | Decision Score (95% range) | Accuracy | Calibration gap (ECE, pts) | Hand-off at 95% | $ per 1k decisions | Speed p95 |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| =1 | [Gemini 3.8 Flash](https://jevals.com/models/gemini-3.8-flash/) | LLM, Google, reasoning low | 73.0 (62.1 to 82.2) | 92.5% | 2.0 | 79% | $0.8 | 5.5 s |
| =1 | [Jev](https://jevals.com/models/jev/) | System One, TypeSafe AI | 69.0 (59.8 to 76.5) | 91.3% | 5.0 | 86% | $0.029 | 653 ms |
| 2 | [Qwen3.8 Flash](https://jevals.com/models/qwen3.8-flash/) | LLM, Alibaba, no reasoning | 62.4 (50.9 to 72.5) | 89.7% | 2.5 | — | $0.086 | 3.5 s |
| =3 | [GLM-5.3](https://jevals.com/models/glm-5.3/) | LLM, Z.ai, reasoning low | 60.6 (50.0 to 69.8) | 88.7% | 2.0 | 45% | $0.51 | 2.9 s |
| =3 | [Mistral Medium 3.5](https://jevals.com/models/mistral-medium-3.5/) | LLM, Mistral AI, no reasoning | 58.0 (45.3 to 69.8) | 88.8% | 5.2 | 65% | $0.87 | 917 ms |
| =3 | [Mercury 2.5](https://jevals.com/models/mercury-2.5/) | LLM, Inception, no reasoning | 55.7 (44.8 to 65.4) | 87.1% | 4.7 | 19% | $0.023 | 1.0 s |
| 7 | [DeepSeek V4.1 Flash](https://jevals.com/models/deepseek-v4.1-flash/) | LLM, DeepSeek, no reasoning | 47.5 (34.7 to 59.1) | 83.7% | 6.2 | — | $0.078 | 1.1 s |
| 8 | [Label prior](https://jevals.com/models/label-prior/) | baseline | 0.0 (0.0 to 0.0) | 62.0% | 0.0 | — | — | — |

Decision Score: 100 = perfect, 0 = guessing the label base rates, below 0 = worse than that. Ranks: 1 + the number of rows significantly better (paired item bootstrap, 95%); rows that cannot be told apart share a rank. Hand-off at 95%: the share of decisions a model can take alone, at its own confidence threshold, while staying at least 95% right. "=" before a rank marks a tie. Shared confidence gate for this board: 0.91.

Source: https://jevals.com/noul/ · release 2026-09-18, suite 0.1.0 · data (CC-BY-4.0): https://jevals.com/data/releases/2026-09-18/board.json · cite as: Jevals (jevals.com), release 2026-09-18, suite 0.1.0. CC-BY-4.0.
