# Changelog

Every release, scoring change and removed row, newest first.

## 2026-09-18 · release 2026-09-18 · suite 0.1.0 · grader 1

- First release. Tasks: Banking77 (choice, 77 intents), HelpSteer2 helpfulness (score, 5 levels), PubMedQA (yes/no); 300 items × 5 repeats each. [Findings](https://jevals.com/notes/2026-09-18/).
- Listed: Jev (Vercel AI Gateway); Gemini 3.8 Flash, GLM-5.3, DeepSeek V4.1 Flash, Qwen3.8 Flash, Mistral Medium 3.5 and Mercury 2.5 (OpenRouter, adapter v1); label prior.
- Confidence gate frozen for suite 0.1.0: choice 0.96, yes/no 0.91, score none (no level reaches 5% error).
- Started, then stopped and not listed, because this release lists only each lab's newest models within its budget: GPT-5.6 Luna and Claude Haiku 4.5 (OpenAI's and Anthropic's newest models are frontier models, not yet run), Gemini 3.5 Flash-Lite (replaced by Gemini 3.8 Flash), DeepSeek V4 Flash (replaced by V4.1 Flash) and Mistral Small 2603 (replaced by Mistral Medium 3.5). Their logs are in `data/runs`. Completed Banking77 runs, graded for disclosure only: GPT-5.6 Luna 73.6 (82.9% accuracy), Gemini 3.5 Flash-Lite 62.5, Claude Haiku 4.5 58.2, Mistral Small 2603 51.6.
- Measured spend: $15.07 for the listed systems, $3.98 for the stopped runs.

Source: https://jevals.com/changelog/
