# score (rate on a rubric) benchmark: Jev vs 6 LLMs

No model clearly beats guessing on HelpSteer2 helpfulness. Jev (9.2, tied for rank 1 of 8): $0.036 per 1,000 decisions, p95 670 ms. Release 2026-09-18.

Place the state on an ordered rubric. The model returns a probability for every level, ready to sort by. The test: “How helpful is the response to the prompt?” (HelpSteer2 helpfulness, 5 levels, 300 questions × 5 runs, human labels, https://huggingface.co/datasets/nvidia/HelpSteer2).

| Rank | Model | Type | Decision Score (95% range) | Accuracy | Calibration gap (ECE, pts) | Hand-off at 95% | $ per 1k decisions | Speed p95 |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| =1 | [Jev](https://jevals.com/models/jev/) | System One, TypeSafe AI | 9.2 (−4.4 to 21.0) | 41.3% | 19.7 | — | $0.036 | 670 ms |
| =1 | [GLM-5.3](https://jevals.com/models/glm-5.3/) | LLM, Z.ai, reasoning low | 7.8 (−4.6 to 18.6) | 43.0% | 12.9 | — | $0.78 | 3.6 s |
| =1 | [Gemini 3.8 Flash](https://jevals.com/models/gemini-3.8-flash/) | LLM, Google, reasoning low | 4.6 (−12.2 to 19.0) | 42.4% | 22.6 | — | $1.3 | 4.0 s |
| =1 | [Label prior](https://jevals.com/models/label-prior/) | baseline | 0.0 (0.0 to 0.0) | 41.7% | 0.0 | — | — | — |
| =3 | [Qwen3.8 Flash](https://jevals.com/models/qwen3.8-flash/) | LLM, Alibaba, no reasoning | −1.4 (−15.4 to 11.5) | 36.1% | 25.8 | — | $0.13 | 2.6 s |
| =3 | [Mercury 2.5](https://jevals.com/models/mercury-2.5/) | LLM, Inception, no reasoning | −5.5 (−16.2 to 4.0) | 41.7% | 31.5 | — | $0.034 | 1.2 s |
| 5 | [Mistral Medium 3.5](https://jevals.com/models/mistral-medium-3.5/) | LLM, Mistral AI, no reasoning | −13.7 (−28.3 to −0.1) | 43.9% | 33.6 | — | $1.3 | 1.2 s |
| 7 | [DeepSeek V4.1 Flash](https://jevals.com/models/deepseek-v4.1-flash/) | LLM, DeepSeek, no reasoning | −19.0 (−36.0 to −3.3) | 34.7% | 33.5 | — | $0.13 | 1.2 s |

Decision Score: 100 = perfect, 0 = guessing the label base rates, below 0 = worse than that. Ranks: 1 + the number of rows significantly better (paired item bootstrap, 95%); rows that cannot be told apart share a rank. Hand-off at 95%: the share of decisions a model can take alone, at its own confidence threshold, while staying at least 95% right. "=" before a rank marks a tie. Shared confidence gate for this board: none.

Source: https://jevals.com/score/ · release 2026-09-18, suite 0.1.0 · data (CC-BY-4.0): https://jevals.com/data/releases/2026-09-18/board.json · cite as: Jevals (jevals.com), release 2026-09-18, suite 0.1.0. CC-BY-4.0.
