Qwen3.7-Max
Alibaba’s closed-weight agent flagship: top-5 intelligence, 1M context, and 35-hour autonomy at half the Western-frontier price.
Read the full Qwen3.7-Max analysisContext
1M
Max output
66K
Input /1M
$1.48
Output /1M
$4.43
Live pricing via OpenRouter
Best for
- High-volume, verifiable long-horizon agentic workloads
- Teams on the Anthropic Messages protocol wanting a cheaper drop-in to trial
- Context-heavy agents that benefit from the 90% cached-input discount
Watch out
Closed-weight / API-only (no self-host), undisclosed architecture, and a verbose reasoner whose effective cost-per-task exceeds the per-token rate; trails Opus 4.8 on aggregate and hardest coding.
For creators. Drive long autonomous coding/optimization runs and batch refactors via Claude Code or Qwen Code; route expensive-failure work to Opus.
Benchmarks
| aa intelligence index | 56.6 |
| swe bench pro | 60.6 |
| swe bench verified | 80.4 |
| terminal bench 2 0 | 69.7 |
| gpqa diamond | 92.4 |
| livecodebench | 91.6 |
| mmlu pro | 89.6 |
| humanitys last exam | 41.4 |
Capabilities
- 1M-token context for long-horizon agentic work
- Documented 35-hour autonomous tool-loop run (vendor demo)
- Native Anthropic Messages-compatible protocol (runs in Claude Code / Qwen Code)
- 90% cached-input discount, suited to context-heavy agents
- Top-5 Artificial Analysis Intelligence Index at roughly half Western-frontier price
Compare Qwen3.7-Max
Claude Fable 5 vs Qwen3.7-Max
Qwen3.7-Max leads its peer group (Kimi, DeepSeek) on hard agentic coding and matches Fable 5 on context — at a quarter of the price. Fable 5 keeps a ~19-point SWE-Bench Pro lead. Value leader vs ceiling, both closed.
Qwen3.7-Max vs DeepSeek V4
Qwen3.7-Max has the higher raw intelligence but is closed and API-only; DeepSeek V4 is open-weight MIT, cheaper, and self-hostable. Capability vs control.