Skip to content
FrankX.AI
Open Receipts
Last harness measurement: 2026-07-01
External snapshot: 2026-08-31

Starlight Model Arena

Which model should handle which task? This page answers with two kinds of claim, kept apart on purpose: first-party measurements, each backed by a public JSON receipt, and third-party figures, each carrying its source and licence. No number appears without one or the other, and the two are never merged into a ranking.

Measured — harness receipts

Every round the harness has run, with the raw JSON each claim traces to. 1 round on the books — an honest, growing record, not a leaderboard.

Machine-readable manifest
Round 5 - Sonnet 5 Arrives2026-07-013 tasks

Same task prompt dispatched to each model as a parallel subagent via the Claude Code Agent tool's per-spawn model override, no tool access. Reasoning task checked against a ground truth computed independently before dispatch. Coding task: each model's raw code submission saved verbatim and executed with Node.js against 4 fixed assertions (not self-reported by the model). A third task (structured JSON output under hard constraints) was blocked by the harness's own auto-mode safety classifier before any model saw it - recorded as a methodology finding, not a contestant result.

Contestants: Claude Sonnet 5 · Claude Opus 4.8 · Claude Haiku 4.5

  • reasoning-divisor-countsonnet-5: passopus-4-8: passhaiku-4-5: pass
  • coding-merge-intervalssonnet-5: passopus-4-8: passhaiku-4-5: pass
  • constraint-json-schemasonnet-5: blockedopus-4-8: blockedhaiku-4-5: blocked
Raw JSON receipt

Task record

Per-task statuses exactly as recorded — including blocked, which means the harness could not verify that task for anyone that day. No aggregation into a single rate: with this few rounds, a percentage would imply precision the data does not have.

reasoning-divisor-count

n=3 · 2026-07-01
  • sonnet-5pass
  • opus-4-8pass
  • haiku-4-5pass

coding-merge-intervals

n=3 · 2026-07-01
  • sonnet-5pass
  • opus-4-8pass
  • haiku-4-5pass

constraint-json-schema

n=3 · 2026-07-01
  • sonnet-5blocked
  • opus-4-8blocked
  • haiku-4-5blocked

Price landscape — third-party figures

Every model the external snapshot prices, on licence-cleared data only. Emerald points have been through the harness at least once. This is deliberately a price chart, not a capability chart — a capability axis needs more measured rounds than exist today.

$0.1$0.2$0.5$1$2$5$10$20$0.5$1$2$5$10$20$50$100$200Input price, USD per million tokens (log)Output price, USD per million tokens (log)Claude Haiku 4.5 — in $1 / out $5 per M tokens · harness-measuredClaude Haiku 4.5Claude Opus 4.8 — in $5 / out $25 per M tokens · harness-measuredClaude Opus 4.8Claude Sonnet 5 — in $2 / out $10 per M tokens · harness-measuredClaude Sonnet 5Claude Fable 5 — in $10 / out $50 per M tokensClaude Opus 4.5 — in $5 / out $25 per M tokensClaude Opus 4.6 — in $5 / out $25 per M tokensClaude Opus 5 — in $5 / out $25 per M tokensClaude Sonnet 4.5 — in $3 / out $15 per M tokensClaude Sonnet 4.6 — in $3 / out $15 per M tokensGemini 3.5 Flash — in $1.5 / out $9 per M tokensGemini 3.7 Flash — in $0.75 / out $3.75 per M tokensGPT-5.2 Pro — in $21 / out $168 per M tokensGPT-5.5 — in $5 / out $30 per M tokensGrok 4.3 — in $1.25 / out $2.5 per M tokensGrok 4.6 — in $2 / out $6 per M tokensKimi K2.6 — in $0.95 / out $4 per M tokensMicrosoft Phi-4 (open-weight family) — in $0.13 / out $0.5 per M tokensMistral Large 3 — in $0.5 / out $1.5 per M tokensQwen3.7-Max — in $2.5 / out $7.5 per M tokens
harness-measured (3)priced, not yet measured (16)n=19 · prices from models.dev [1] · retrieved 2026-08-31

Cross-reference

The same figures as a table. The receipt column records whether first-party evidence exists for a model — it is a fact, not a score.

ModelInput $/M [1]Output $/M [1]ContextHarness receipt
Claude Haiku 4.5$1$5200K1 round
Claude Opus 4.8$5$251000K1 round
Claude Sonnet 5$2$101000K1 round
Claude Fable 5$10$501000Knone yet
Claude Opus 4.5$5$25200Knone yet
Claude Opus 4.6$5$25200Knone yet
Claude Opus 5$5$251000Knone yet
Claude Sonnet 4.5$3$15200Knone yet
Claude Sonnet 4.6$3$15200Knone yet
Gemini 3.5 Flash$1.5$91000Knone yet
Gemini 3.7 Flash$0.75$3.751049Knone yet
GPT-5.2 Pro$21$168196Knone yet
GPT-5.5$5$301000Knone yet
Grok 4.3$1.25$2.51000Knone yet
Grok 4.6$2$6500Knone yet
Kimi K2.6$0.95$4262Knone yet
Microsoft Phi-4 (open-weight family)$0.125$0.516Knone yet
Mistral Large 3$0.5$1.5256Knone yet
Qwen3.7-Max$2.5$7.51000Knone yet

Data sources & licences

  1. [1]models-devMITactiveretrieved 2026-08-31
  2. [2]openrouter-modelsOpenRouter ToS — redistribution not yet verifiedheld — licence unverifiedNo figures from this source appear on this page.
  3. [3]openrouter-rankingsOpenRouter rankings terms — redistribution not yet verifiedheld — licence unverifiedNo figures from this source appear on this page.

Named but never ingested — their terms do not permit republishing figures here:

Vendor-published benchmarks — Claude Sonnet 5

Released 2026-06-30

These are Anthropic's own numbers at launch, not measured by this harness — cited for context only.

Sonnet 5 edges the flagship Opus 4.8 on knowledge work while running at roughly 40% of the list price.

Pricing: $2 / $10 per million input/output tokens; Anthropic cancelled the September increase (checked 2026-09-07)

August 2026 Wave — Public-Report Orientation

Compiled: 2026-08-17
Synthesized from public vendor reports — not a harness result. No round was dispatched and no receipt exists for this entry. Treat the lane calls as orientation, not measurement.
No harness tally — vendor-reported strengths onlyModels: Grok 4.6 · Gemini 3.7 Flash · DeepSeek-V4-Pro · Qwen3.8-27B

Vendor reports point to a strong agentic/speed wave. These lane calls are orientation from public benchmarks, not measured results — treat them as hypotheses until harness rounds run.

The Proving Ground Methodology

How the harness executes evaluations to eliminate cherry-picking bias.

1

Standardized Task Envelopes

Every model receives identical constraints, schemas, and stop conditions without conversational preambles.

2

Multi-Model Blind Evaluation

Outputs are judged anonymously by non-contender frontier models using rigid rubric scoring.

3

Maker ≠ Checker Cross-Verification

Generated solutions must be verified by a model from an independent provider before pass certification.

4

Durable Proof Receipts

Every execution logs exact token metrics, latencies, AST verification outputs, and JSON receipts.

Read complete verification rules in tools/arena/README.md

Tactical Routing Guidelines

Working hypotheses from the measured rounds and the vendor reports above — labelled per source, revised as rounds accumulate.

Swarm / Queen orchestration + real-time tools

Grok 4.6 (Hermes)

Agentic post-training + reliability for loops, memory, and external data.

Fast volume agentic + coding speed

Gemini 3.7 Flash

Large lifts in coding/agentic benchmarks + leading output speed at intro price.

Cost-sensitive or local multi-agent systems

Qwen3.8-27B or DeepSeek-V4-Pro

Open, efficient, strong agentic results on consumer or budget infra.

Deep judgment + brand-craft

Claude Opus/Fable-class

Retained leadership on complex situational judgment and high-craft output.

Verify / checker (mandatory CROSS-MODEL-GATE)

Rotate provider different from maker (Gemini 3.7 or Qwen)

Independent verification catches harness-specific artifacts.

Heavy multi-step work, any model

Enforce contracts structurally (schemas, evals, human gates)

Output discipline degrades under load across all frontier models.

Caveats & Safety Guardrails

  • n small per task — directional signals. Re-test in your own harness.
  • August 2026 wave entries are vendor-reported benchmarks only — no harness rounds have run for that wave yet.
  • Pricing introductory for Gemini 3.7 Flash; weights for GLM-5.3 delayed.
  • Everything measured in harness context (Claude Code / multi-CLI) where possible.

Frequently Asked Questions

How do the new August 2026 models change routing for multi-agent systems?

Grok 4.6 for orchestration. Gemini 3.7 Flash for speed/volume. Open models (Qwen, DeepSeek) for cost and local. Claude remains for deepest judgment and craft. Always CROSS-MODEL-GATE verify.

Starlight Model Arena • built and run by Frank's multi-agent research system • 2026