Skip to content

Changelog

Every release, scoring change and removed row, newest first.

2026-09-18 · release 2026-09-18 · suite 0.1.0 · grader 1

  • First release. Tasks: Banking77 (choice, 77 intents), HelpSteer2 helpfulness (score, 5 levels), PubMedQA (yes/no); 300 items × 5 repeats each. Findings.
  • Listed: Jev (Vercel AI Gateway); Gemini 3.8 Flash, GLM-5.3, DeepSeek V4.1 Flash, Qwen3.8 Flash, Mistral Medium 3.5 and Mercury 2.5 (OpenRouter, adapter v1); label prior.
  • Confidence gate frozen for suite 0.1.0: choice 0.96, yes/no 0.91, score none (no level reaches 5% error).
  • Started, then stopped and not listed, because this release lists only each lab's newest models within its budget: GPT-5.6 Luna and Claude Haiku 4.5 (OpenAI's and Anthropic's newest models are frontier models, not yet run), Gemini 3.5 Flash-Lite (replaced by Gemini 3.8 Flash), DeepSeek V4 Flash (replaced by V4.1 Flash) and Mistral Small 2603 (replaced by Mistral Medium 3.5). Their logs are in data/runs. Completed Banking77 runs, graded for disclosure only: GPT-5.6 Luna 73.6 (82.9% accuracy), Gemini 3.5 Flash-Lite 62.5, Claude Haiku 4.5 58.2, Mistral Small 2603 51.6.
  • Measured spend: $15.07 for the listed systems, $3.98 for the stopped runs.