Quick Answer
GPT-6 Astra is OpenAI's flagship reasoning model, released September 3, 2026. It costs $10 per million input tokens and $50 per million output tokens on the standard API tier, with cached input at $1 and cache writes at $12.50. That is 2.5 times GPT-5.6 Sol's promotional rate and matches Anthropic's Fable 5.1 on both headline numbers. Batch and Flex halve those rates, Fast mode doubles them, and any prompt past 272K input tokens reprices the entire request.
On September 3, 2026, OpenAI shipped its most capable model and priced it like one. Then it told buyers to stop counting tokens. “Pricing tokens doesn’t make any sense,” OpenAI president Greg Brockman told reporters at the launch briefing, arguing buyers should judge the price per completed task instead.
The tokens, of course, will still show up on the invoice. For the people who see that invoice rather than the headline, the launch raises a specific question: is a model priced at 2.5 times your current flagship a cost increase, or a cost reduction with a bigger sticker on it?
The honest answer is that it depends on your workloads, and OpenAI has given you until November 21, 2026 to figure it out. That’s when GPT-5.6 Sol’s promotional pricing is no longer guaranteed. Below: what the rate card says, what it leaves out, and how to decide before the deadline decides for you.
What is GPT-6 Astra?
GPT-6 Astra is OpenAI’s new flagship reasoning model, released September 3, 2026, built for long-horizon agentic work: computer use, coding, research, and multi-step tasks that previous models abandoned halfway through. It ships with a 1 million token context window and, according to OpenAI, a new frontier in speed, accuracy, and safety.
The launch numbers back up the agentic framing. On OSWorld 2.0, a standard benchmark for autonomous computer use, Astra both scores higher than GPT-5.6 Sol and finishes faster, which is the whole basis of the per-task pricing argument below. It creates documents, spreadsheets, and presentations from your templates and adjusts when you change direction mid-task, which is the part previous models handled by starting over.
Astra also arrives wrapped in cybersecurity caution. It’s the first model OpenAI has rated Critical for cyber risk under its Preparedness Framework, which is why enterprises in Daybreak, OpenAI’s gated access program, got it first. The model went through a voluntary US administration review before release, and eligible API customers can run it with Zero Data Retention.
One operational detail matters more than the press coverage suggests: Astra ships with safety monitoring that can interrupt work it flags as suspicious. In ChatGPT and Codex you get a review prompt. In the API, a flagged task can stop outright, and OpenAI concedes the filters currently trip on innocuous work. Worth knowing before you wire Astra into an unattended pipeline that runs at 3 a.m.
Report
Finance needs to prove AI’s return: CloudZero report
260 senior finance leaders (more than half CFOs) told us why the speed of seeing AI spend, not the size of it, separates who pulls ahead on AI from who gets burned.
How much does the GPT-6 Astra API cost?
The GPT-6 Astra API costs $10 per million input tokens, $1 per million cached input tokens, $12.50 per million cache writes, and $50 per million output tokens on the standard tier at short context. Those are OpenAI’s published rates as of the September 2026 launch, ahead of general API availability.
| Standard tier (per 1M tokens) | Short context | Long context (over 272K input) |
|---|---|---|
| Input | $10.00 | $20.00 |
| Cached input | $1.00 | $2.00 |
| Cache writes | $12.50 | $25.00 |
| Output | $50.00 | $75.00 |
Two footnotes on that card deserve promotion to headlines:
- Cache writes bill at 1.25 times the uncached input rate, so aggressive caching strategies carry an upfront charge before the $1 reads pay it back.
- And regional data residency endpoints add a 10% uplift for models released on or after March 5, 2026, which includes Astra.
If you run OpenAI models through a cloud provider, Azure and Bedrock rates are billed by that provider and can differ from OpenAI’s direct pricing. The model name in the API is gpt-6-astra, and the full OpenAI pricing lineup now stretches across four flagship models and four service tiers.
What do Batch, Flex, and Fast mode change?
Batch and Flex processing cut Astra’s rates by 50%, to $5 per million input and $25 per million output, while Fast mode doubles standard rates to $20 and $100. Same model, four service tiers, which is three more answers to “what does it cost” than most budget owners wanted.
| Service tier (per 1M tokens, short context) | Input | Output |
|---|---|---|
| Batch | $5.00 | $25.00 |
| Flex | $5.00 | $25.00 |
| Standard | $10.00 | $50.00 |
| Fast mode | $20.00 | $100.00 |
Fast mode buys up to 2.5 times standard processing speed for exactly 2 times the money, which is at least an honest trade. It’s also unavailable for Astra with EU data residency, so European teams needing residency guarantees run standard processing whether they like it or not.
The practical takeaway: anything asynchronous belongs in Batch or Flex. A nightly summarization job running at standard rates is a donation, not an architecture decision.
How does long context change the bill?
Crossing 272K input tokens reprices the entire request at 2 times the input and cache rates and 1.5 times the output rate. Not the overflow. The whole request. It’s less a pricing tier than a trapdoor, and agentic workloads that accumulate context are the ones most likely to fall through it.
The math gets uncomfortable fast. A request with 280K input tokens and 20K output costs about $7.10, while the same work trimmed to 272K input costs roughly $3.72. Eight thousand tokens of extra context nearly doubled the bill.
Astra’s headline 1 million token window is real, but the affordable part of it ends at 272K. Teams doing long-document analysis or letting agent loops hoard context should treat that threshold as a budget line, with compaction and summarization as the enforcement mechanism.
How does Astra compare to GPT-5.6 and other frontier models?
Astra costs 2.5 times GPT-5.6 Sol’s promotional rate of $4 input and $20 output, and exactly matches Anthropic’s Fable 5.1 at the top of the market. The two frontier labs now carry identical list prices, which quietly deleted price shopping as a routing strategy.
| Model (per 1M tokens, standard) | Input | Output |
|---|---|---|
| GPT-6 Astra | $10.00 | $50.00 |
| Claude Fable 5.1 (Anthropic) | $10.00 | $50.00 |
| GPT-5.6 Sol (promo through Nov 21, 2026) | $4.00 | $20.00 |
| GPT-5.6 Terra | $2.00 | $12.00 |
| GPT-5.6 Luna | $0.20 | $1.20 |
Elsewhere in the market, launch coverage pegs Meta’s Muse at $1.25 input and $4.25 output, with Google’s introductory Gemini 3.8 Flash rate at $0.75 and $3.75. Those are different capability classes, but they anchor the routing question: most production traffic never needed a frontier model in the first place.
The GPT-5.6 family covers that cheaper tier well, and our GPT-5.6 pricing breakdown walks through Sol, Terra, and Luna in detail. On the Anthropic side, Claude’s current pricing and the Mythos tier rates map the equivalent ladder. The sane pattern hasn’t changed: route the routine 90% of calls to a mid-tier model and reserve frontier rates for the tasks that fail without them.
Is Astra actually cheaper per task?
For agentic workloads, possibly yes. Astra finishes OSWorld tasks in roughly 40 minutes against Sol’s 75, scoring 72.6% to Sol’s 65.7%, and OpenAI claims its best Astra configuration beats Sol’s at about 57% lower estimated cost per task on the DeepSWE coding benchmark. Fewer retries, smaller invoice. That’s the entire argument for the 2.5x sticker.
The argument has a boundary, though. Per-task savings only materialize when the task was long, multi-step, and failure-prone to begin with. A chat completion that Sol already handled in one attempt doesn’t get cheaper on Astra. It just gets more expensive, with better vibes.
There is also a measurement problem. That 57% is OpenAI’s own math on one benchmark, and cost per task is a metric its invoice does not report. Your bill still arrives in tokens, so proving the claim for your own workloads means tracking spend at the workflow level, not the account level. Worth doing regardless of which model wins.
Independent benchmarks add a useful footnote: on the Artificial Analysis Intelligence Index, Astra scores 61 at maximum reasoning effort, well above the reasoning-model median of 36 but only a nudge past GPT-5.6 Sol on general reasoning. The gains concentrate in computer use, coding, and cybersecurity. If your workloads live elsewhere, you are paying frontier rates for capabilities you don’t call.
How do you get GPT-6 Astra in ChatGPT?
Astra launched September 3 for enterprises in Daybreak, OpenAI’s gated access program, with API access and the Plus, Pro, Business, and Enterprise plans following over the coming days, plus availability through AWS Bedrock and Azure. Pro, Business, and Enterprise tiers also get GPT-6 Astra Pro, the heavier reasoning variant.
Subscription prices themselves didn’t move at launch. Existing plan rates and limits are covered in our full guide to what ChatGPT costs, and the Codex side of the house has its own pricing. The pattern from past launches is predictable: flat subscription price, new model, tighter usage caps on the expensive one. Watch the caps, not the price.
What are the pros and cons of GPT-6 Astra?
Astra’s case rests on task economics and capability gains, and its case against rests on the sticker, the fine print, and where the gains actually land. Both cases are legitimate, which is rarer in AI launches than it should be.
Where Astra earns the rate:
- Task completion economics. OpenAI claims roughly 57% lower estimated cost per completed task than Sol on DeepSWE, driven by faster finishes and fewer tokens per outcome
- Genuine capability jump in computer use, coding, science, and cybersecurity, with a 72.6% OSWorld score
- Better boundary discipline: in OpenAI’s scope tests without production safeguards, Astra exceeded its authorized perimeter in 0% of runs against Sol’s 48.2%
- A 1M token context window with Zero Data Retention support for eligible API customers
- Price parity with Fable 5.1, so choosing between frontier labs is finally about fit rather than discounts
Where it doesn’t:
- The 2.5x step over Sol lands on every workload, including the ones that gained nothing from the upgrade
- General reasoning barely moved against Sol on independent benchmarks, so chat-heavy workloads pay more for similar answers
- The 272K long-context threshold reprices whole requests, punishing exactly the agentic patterns Astra is sold for
- EU data residency loses Fast mode, and every residency endpoint pays the 10% uplift
- Safety monitoring pauses runs for review in ChatGPT and can stop flagged API tasks outright, a feature for compliance and a surprise for unattended automation
What should finance teams do before November 21?
Segment your OpenAI traffic by task type before Sol’s promotional window closes. The Astra decision is really three separate decisions filed under one model name, and they have different answers. Long-horizon agentic work is the migration candidate. Routine completions should drop down the ladder to Terra or Luna. Everything asynchronous moves to Batch.
That triage only works if you can see spend at the workload level. Four service tiers, two context lengths, thresholds that reprice whole requests, and savings that only show up per task: account-level billing tells you nothing here. Tracking OpenAI spend per feature, per customer, and per agent run is what turns Brockman’s price-per-task framing from a slogan into a number you can defend in a budget review.
CloudZero’s OpenAI integration pulls that usage into the same view as the rest of your cloud and AI spend, so Astra’s unit economics stop being a press-release claim and start being your data. From there, the usual discipline applies, and our guides to OpenAI cost optimization and the OpenAI cost calculator cover the tactics.
If you’d rather see it than read about it, take the self-guided tour or book a demo and bring your ugliest OpenAI invoice.