Skip to content

System One · TypeSafe AI

Jev

Ranks 1 of 7 on noul, 1 of 7 on choice, 1 of 7 on score. Native probabilities, via Vercel AI Gateway.

Jev is TypeSafe AI's hosted decision model, launched on 2026-09-15 and the first System One model: it writes no text. You send a state (any text or JSON) with typed questions and get probabilities back in one parallel pass: noul answers yes or no with P(yes), choice picks one of up to 255 options with a probability for each, and score places the state on an ordered rubric of 2 to 10 levels. It is called at api.typesafe.ai/v1/systemone or through Vercel AI Gateway as typesafe-ai/jev, and costs $0.042 per million input tokens, with output free. The numbers below are Jevals' own measurements against human labels, next to LLMs asked the same questions (how we score).

compare: vs Gemini 3.8 Flash · vs GLM-5.3 · vs DeepSeek V4.1 Flash · vs Qwen3.8 Flash · vs Mistral Medium 3.5 · vs Mercury 2.5

noulif

69.0Decision Score
rank =1 of 8

accuracy
91.3%
calibration gap
5.0
hand-off at 95%
86%
$ per 1k
$0.029
speed p95
653 ms
10said 55%, right 58% (48 answers)said 65%, right 73% (109 answers)said 75%, right 82% (170 answers)said 86%, right 91% (371 answers)said 94%, right 98% (802 answers)Said vs right. Squares on the diagonal mean honest confidence; bigger squares hold more answers.
choicematch

67.8Decision Score
rank =2 of 8

accuracy
79.7%
calibration gap
9.8
hand-off at 95%
54%
$ per 1k
$0.043
speed p95
693 ms
10said 29%, right 0% (1 answers)said 35%, right 20% (25 answers)said 44%, right 23% (40 answers)said 54%, right 51% (69 answers)said 64%, right 54% (69 answers)said 74%, right 63% (95 answers)said 85%, right 60% (131 answers)said 98%, right 91% (1070 answers)Said vs right. Squares on the diagonal mean honest confidence; bigger squares hold more answers.
scoresort

9.2Decision Score
rank =1 of 8

accuracy
41.3%
calibration gap
19.7
hand-off at 95%
$ per 1k
$0.036
speed p95
670 ms
10said 27%, right 47% (34 answers)said 36%, right 31% (118 answers)said 45%, right 31% (261 answers)said 54%, right 31% (339 answers)said 65%, right 41% (324 answers)said 75%, right 53% (222 answers)said 83%, right 61% (157 answers)said 94%, right 76% (45 answers)Said vs right. Squares on the diagonal mean honest confidence; bigger squares hold more answers.

everything behind the numbers

noul
accuracy91.3%95% range 87.994.3%
calibration gap5.0 ptslower is better · 0.0% of answers at confidence 1.00
10said 55%, right 58% (48 answers)said 65%, right 73% (109 answers)said 75%, right 82% (170 answers)said 86%, right 91% (371 answers)said 94%, right 98% (802 answers)
automation at 95%86%of decisions, acting at confidence ≥ 0.75 (right 95.0%)shared gate ≥ 0.91: acts on 49%, right 98.6%
speed438 msmedian · p95 653 ms
stability0.0%of answers change on an identical rerun
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupai-sdk experimental_evaluate ai@7.0.106 · native probabilities via Vercel AI Gateway · reasoning · temperature · 300 items × 5 runs · 2026-09-18 · served as typesafe-ai/jevpubmedqa run log
choice
accuracy79.7%95% range 75.583.8%
calibration gap9.8 ptslower is better · 41.5% of answers at confidence 1.00
10said 29%, right 0% (1 answers)said 35%, right 20% (25 answers)said 44%, right 23% (40 answers)said 54%, right 51% (69 answers)said 64%, right 54% (69 answers)said 74%, right 63% (95 answers)said 85%, right 60% (131 answers)said 98%, right 91% (1070 answers)
automation at 95%54%of decisions, acting at confidence ≥ 0.98 (right 95.4%)shared gate ≥ 0.96: acts on 59%, right 94.3%
speed467 msmedian · p95 693 ms
stability2.7%of answers change on an identical rerun · 10.3% change when options are reordered
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupai-sdk experimental_evaluate ai@7.0.106 · native probabilities via Vercel AI Gateway · reasoning · temperature · 300 items × 5 runs · 2026-09-18 · served as typesafe-ai/jevbanking77 run log
score
accuracy41.3%95% range 35.946.9%
calibration gap19.7 ptslower is better · 0.3% of answers at confidence 1.00
10said 27%, right 47% (34 answers)said 36%, right 31% (118 answers)said 45%, right 31% (261 answers)said 54%, right 31% (339 answers)said 65%, right 41% (324 answers)said 75%, right 53% (222 answers)said 83%, right 61% (157 answers)said 94%, right 76% (45 answers)
automation at 95%never reaches 95% accuracy
speed478 msmedian · p95 670 ms
stability2.3%of answers change on an identical rerun
valid output100.0%0.0% refused · 1,500 of 1,500 answered
setupai-sdk experimental_evaluate ai@7.0.106 · native probabilities via Vercel AI Gateway · reasoning · temperature · 300 items × 5 runs · 2026-09-18 · served as typesafe-ai/jevhelpsteer2 run log
model id
typesafe-ai/jev
served as
typesafe-ai/jev
adapter
ai-sdk experimental_evaluate ai@7.0.106
host
Vercel AI Gateway
price as of
2026-09-18
harness
6d93f79a2310
release
2026-09-18

see every question Jev answered →