GPT-6 Astra
A candidate for demanding work across research, software and creative tools. Evaluate the complete handoff.
Read the full GPT-6 Astra analysisContext
1.05M
Max output
128K
Input /1M
$10.00
Output /1M
$50.00
Workloads to evaluate
- Evaluate multi-step research with source and artifact checks
- Compare repository work against the current model baseline
- Test creative-tool coordination and editable handoffs
Watch out
Vendor-reported benchmarks are not local results. Twelve synthetic evaluation cases are defined but unrun. Standard API pricing is $10/$50 per million input/output tokens up to 272K input; longer context and other processing tiers differ. No calibrated LLM judge or live connector success rate is claimed.
For creators. Coordinate documented tools, then inspect outputs and editable sources. Adobe, Canva and HeyGen have separate access requirements. Astra does not natively output audio or video.
Evaluation status: not run
12 synthetic decision cases defined; 0 model observations. These cases do not measure live connector execution or creative artifact quality. No calibrated judge or production promotion is recorded.
Benchmarks
OpenAI-published launch results; not FrankX measurements. Percentage units are named in the rows. The AA Intelligence Index is an index score. Source and comparison context.
| automationbench percent | 41.4 |
| terminal bench 4 percent | 64.6 |
| deepswe v1 1 percent | 74.1 |
| aa intelligence index v4 1 1 | 54.7 |
Capabilities
- Text input/output and image input; audio and video are not native model outputs
- API reasoning effort: low, medium, high, xhigh, max
- Async tool calling; the application executes tools and manages pending jobs
- ChatGPT, Work and Codex availability depends on client, plan, workspace and rollout
- A separate openai/gpt-6-astra-fast variant exists at 2x price ($20/$100 per 1M) for lower latency with no stated capability difference — not a default upgrade
- OpenAI's own deployment-safety documentation rates Astra at the 'Critical' level of cyber capability; the publicly released model ships with classifiers that decline cyber-exploitation, bio/chem, violence and fraud requests