# DeepSeek V4.1 Flash (DeepSeek)

LLM, DeepSeek, no reasoning. Verbalized probabilities through the Jevals adapter, via OpenRouter · DeepSeek.

| Board | Decision Score (95% range) | Rank | Accuracy | Calibration gap (ECE, pts) | Hand-off at 95% | $ per 1k decisions | Speed p95 |
| --- | --- | --- | --- | --- | --- | --- | --- |
| [noul (PubMedQA)](https://jevals.com/noul/) | 47.5 (34.7 to 59.1) | rank 7 of 8 | 83.7% | 6.2 | — | $0.078 | 1.1 s |
| [choice (Banking77)](https://jevals.com/choice/) | 63.9 (58.5 to 69.6) | tied for rank 4 of 8 | 75.9% | 2.8 | 39% | $0.14 | 1.3 s |
| [score (HelpSteer2 helpfulness)](https://jevals.com/score/) | −19.0 (−36.0 to −3.3) | rank 7 of 8 | 34.7% | 33.5 | — | $0.13 | 1.2 s |

Decision Score: 100 = perfect, 0 = guessing the label base rates, below 0 = worse than that. Ranks: 1 + the number of rows significantly better (paired item bootstrap, 95%); rows that cannot be told apart share a rank. Hand-off at 95%: the share of decisions a model can take alone, at its own confidence threshold, while staying at least 95% right.

Setup: model id deepseek/deepseek-v4.1-flash; served as deepseek/deepseek-v4.1-flash; adapter jevals-llm 0383a0e3e592; reasoning none; temperature provider default; price as of 2026-09-18.

Compare: [Jev vs DeepSeek V4.1 Flash](https://jevals.com/compare/jev-vs-deepseek-v4.1-flash/)

Source: https://jevals.com/models/deepseek-v4.1-flash/ · release 2026-09-18, suite 0.1.0 · data (CC-BY-4.0): https://jevals.com/data/releases/2026-09-18/board.json · cite as: Jevals (jevals.com), release 2026-09-18, suite 0.1.0. CC-BY-4.0.
