Starlight Model Arena
Which model should handle which task? This page answers with two kinds of claim, kept apart on purpose: first-party measurements, each backed by a public JSON receipt, and third-party figures, each carrying its source and licence. No number appears without one or the other, and the two are never merged into a ranking.
Measured — harness receipts
Every round the harness has run, with the raw JSON each claim traces to. 1 round on the books — an honest, growing record, not a leaderboard.
Round 5 - Sonnet 5 Arrives2026-07-013 tasks
Same task prompt dispatched to each model as a parallel subagent via the Claude Code Agent tool's per-spawn model override, no tool access. Reasoning task checked against a ground truth computed independently before dispatch. Coding task: each model's raw code submission saved verbatim and executed with Node.js against 4 fixed assertions (not self-reported by the model). A third task (structured JSON output under hard constraints) was blocked by the harness's own auto-mode safety classifier before any model saw it - recorded as a methodology finding, not a contestant result.
Contestants: Claude Sonnet 5 · Claude Opus 4.8 · Claude Haiku 4.5
- reasoning-divisor-countsonnet-5: passopus-4-8: passhaiku-4-5: pass
- coding-merge-intervalssonnet-5: passopus-4-8: passhaiku-4-5: pass
- constraint-json-schemasonnet-5: blockedopus-4-8: blockedhaiku-4-5: blocked
Task record
Per-task statuses exactly as recorded — including blocked, which means the harness could not verify that task for anyone that day. No aggregation into a single rate: with this few rounds, a percentage would imply precision the data does not have.
reasoning-divisor-count
n=3 · 2026-07-01- sonnet-5pass
- opus-4-8pass
- haiku-4-5pass
coding-merge-intervals
n=3 · 2026-07-01- sonnet-5pass
- opus-4-8pass
- haiku-4-5pass
constraint-json-schema
n=3 · 2026-07-01- sonnet-5blocked
- opus-4-8blocked
- haiku-4-5blocked
Price landscape — third-party figures
Every model the external snapshot prices, on licence-cleared data only. Emerald points have been through the harness at least once. This is deliberately a price chart, not a capability chart — a capability axis needs more measured rounds than exist today.
Cross-reference
The same figures as a table. The receipt column records whether first-party evidence exists for a model — it is a fact, not a score.
| Model | Input $/M [1] | Output $/M [1] | Context | Harness receipt |
|---|---|---|---|---|
| Claude Haiku 4.5 | $1 | $5 | 200K | 1 round |
| Claude Opus 4.8 | $5 | $25 | 1000K | 1 round |
| Claude Sonnet 5 | $2 | $10 | 1000K | 1 round |
| Claude Fable 5 | $10 | $50 | 1000K | none yet |
| Claude Opus 4.5 | $5 | $25 | 200K | none yet |
| Claude Opus 4.6 | $5 | $25 | 200K | none yet |
| Claude Opus 5 | $5 | $25 | 1000K | none yet |
| Claude Sonnet 4.5 | $3 | $15 | 200K | none yet |
| Claude Sonnet 4.6 | $3 | $15 | 200K | none yet |
| Gemini 3.5 Flash | $1.5 | $9 | 1000K | none yet |
| Gemini 3.7 Flash | $0.75 | $3.75 | 1049K | none yet |
| GPT-5.2 Pro | $21 | $168 | 196K | none yet |
| GPT-5.5 | $5 | $30 | 1000K | none yet |
| Grok 4.3 | $1.25 | $2.5 | 1000K | none yet |
| Grok 4.6 | $2 | $6 | 500K | none yet |
| Kimi K2.6 | $0.95 | $4 | 262K | none yet |
| Microsoft Phi-4 (open-weight family) | $0.125 | $0.5 | 16K | none yet |
| Mistral Large 3 | $0.5 | $1.5 | 256K | none yet |
| Qwen3.7-Max | $2.5 | $7.5 | 1000K | none yet |
Data sources & licences
- [1]models-devMITactiveretrieved 2026-08-31
- [2]openrouter-modelsOpenRouter ToS — redistribution not yet verifiedheld — licence unverifiedNo figures from this source appear on this page.
- [3]openrouter-rankingsOpenRouter rankings terms — redistribution not yet verifiedheld — licence unverifiedNo figures from this source appear on this page.
Named but never ingested — their terms do not permit republishing figures here:
Vendor-published benchmarks — Claude Sonnet 5
Released 2026-06-30These are Anthropic's own numbers at launch, not measured by this harness — cited for context only.
Sonnet 5 edges the flagship Opus 4.8 on knowledge work while running at roughly 40% of the list price.
Pricing: $2 / $10 per million input/output tokens; Anthropic cancelled the September increase (checked 2026-09-07)
August 2026 Wave — Public-Report Orientation
Compiled: 2026-08-17Vendor reports point to a strong agentic/speed wave. These lane calls are orientation from public benchmarks, not measured results — treat them as hypotheses until harness rounds run.
The Proving Ground Methodology
How the harness executes evaluations to eliminate cherry-picking bias.
Standardized Task Envelopes
Every model receives identical constraints, schemas, and stop conditions without conversational preambles.
Multi-Model Blind Evaluation
Outputs are judged anonymously by non-contender frontier models using rigid rubric scoring.
Maker ≠ Checker Cross-Verification
Generated solutions must be verified by a model from an independent provider before pass certification.
Durable Proof Receipts
Every execution logs exact token metrics, latencies, AST verification outputs, and JSON receipts.
Read complete verification rules in tools/arena/README.md
Tactical Routing Guidelines
Working hypotheses from the measured rounds and the vendor reports above — labelled per source, revised as rounds accumulate.
Swarm / Queen orchestration + real-time tools
Grok 4.6 (Hermes)Agentic post-training + reliability for loops, memory, and external data.
Fast volume agentic + coding speed
Gemini 3.7 FlashLarge lifts in coding/agentic benchmarks + leading output speed at intro price.
Cost-sensitive or local multi-agent systems
Qwen3.8-27B or DeepSeek-V4-ProOpen, efficient, strong agentic results on consumer or budget infra.
Deep judgment + brand-craft
Claude Opus/Fable-classRetained leadership on complex situational judgment and high-craft output.
Verify / checker (mandatory CROSS-MODEL-GATE)
Rotate provider different from maker (Gemini 3.7 or Qwen)Independent verification catches harness-specific artifacts.
Heavy multi-step work, any model
Enforce contracts structurally (schemas, evals, human gates)Output discipline degrades under load across all frontier models.
Caveats & Safety Guardrails
- n small per task — directional signals. Re-test in your own harness.
- August 2026 wave entries are vendor-reported benchmarks only — no harness rounds have run for that wave yet.
- Pricing introductory for Gemini 3.7 Flash; weights for GLM-5.3 delayed.
- Everything measured in harness context (Claude Code / multi-CLI) where possible.
Frequently Asked Questions
How do the new August 2026 models change routing for multi-agent systems?
Grok 4.6 for orchestration. Gemini 3.7 Flash for speed/volume. Open models (Qwen, DeepSeek) for cost and local. Claude remains for deepest judgment and craft. Always CROSS-MODEL-GATE verify.