Quick Answer
Claude Opus pricing is $4 per million input tokens and $20 per million output tokens on Claude Opus 5.5, the current model, with cache reads at $0.20 and batch jobs at $2/$10. Opus 5 and the legacy 4-series bill at $5/$25. The 1M context window carries no surcharge.
Every Opus model Anthropic shipped in 2026 held the same line: $5 in, $25 out, per million tokens. Opus 4.6 in February, 4.7 in April, 4.8 in May, Opus 5 in July. Four releases, one price.
On September 22, 2026, the line broke. Claude Opus 5.5 arrived at $4/$20, a 20% cut, with cache reads down 60% to $0.20 per million, and Anthropic estimating typical workloads run about 40% cheaper because the model also finishes tasks in fewer tokens.
A price cut is only as real as your metering, though. CloudZero’s Chief Product Officer Scott Castle documented how LiteLLM undercounts Claude spend when calls route through a gateway: the same model call arrives fully attributed on one path and near-blank on the next. If your Claude telemetry has holes, you have no way of knowing whether September 22 reached your bill or evaporated somewhere between the API and the invoice.
Verifying that is a spend-visibility problem, and it is the thread running through this whole guide.
How does Claude Opus pricing work?
Claude Opus bills per token: input (what you send) and output (what the model writes back), metered separately and quoted per million tokens (MTok). Four modifiers move the effective rate, and per Anthropic’s pricing docs, they stack rather than replace each other:
- Prompt caching. A cache hit on Opus 5.5 costs 5% of the standard input rate, $0.20 per MTok; Sonnet 5.5 matches that ratio, while the legacy models, including Opus 5 and every Opus from 4.5 through 4.8, charge 10%. Writes bill at 1.25x input for the 5-minute duration and 2x for the 1-hour duration.
- Batch processing. Asynchronous jobs completed within 24 hours run 50% off input and output. Anthropic publishes the batched figures directly: $2/$10 on Opus 5.5.
- Fast mode. A research preview that doubles the rate for up to 2.5x speed: $8/$40 on Opus 5.5. It runs on Anthropic’s first-party surfaces only, not on any AWS, Google Cloud, or Microsoft Foundry path.
- US-only inference. For data residency on Claude 4.6 and later, a 1.1x multiplier applies to every token category, including cache reads and writes. Global routing stays at standard rates.
Because the modifiers stack, a batched, cached, US-pinned request multiplies all three. And one meter hides inside another: extended thinking bills at output rates whether or not you read the reasoning tokens. Budget for the thinking, not just the answer.
The Opus line also includes the 1M token context window at flat rates. There is no long-context surcharge at any Opus generation from 4.6 onward, which matters more than it sounds once agents start packing that window on every step.
2026 State of AI Spend Report
Why Finance Caps What It Can’t See
Learn how finance leaders are managing AI spend, and how they’re approaching FY2027
How much does Claude Opus 5.5 cost?
Claude Opus 5.5 costs $4 per million input tokens and $20 per million output tokens at standard rates. It launched September 22, 2026 as the current Opus model, deployable as claude-opus-5-5 on the Claude Platform and through Amazon Web Services, Google Cloud, and Microsoft Foundry, per Anthropic’s announcement.
| Meter | Opus 5.5 rate (per MTok) | Versus Opus 5 |
|---|---|---|
| Input | $4.00 | 20% lower |
| Output | $20.00 | 20% lower |
| Cache read | $0.20 | 60% lower |
| Cache write, 5-minute | $5.00 | 1.25x input |
| Cache write, 1-hour | $8.00 | 2x input |
| Batch input / output | $2.00 / $10.00 | 50% off standard |
| Fast mode input / output | $8.00 / $40.00 | Research preview, up to 2.5x speed |
The sticker cut understates the real one. Anthropic’s launch page says the model “performs at the level of Claude Fable 5.1 on most work” while finishing tasks in fewer tokens, which is how a 20% rate cut becomes an estimated 40% workload cut.
The customer evidence Anthropic published points the same way: Box measured Opus 5.5 using a third of the tokens Opus 5 did, and Noyan Tokgozoglu, Global Head of AI Engineering, reported agentic coding workloads costing 40 to 50% less at matched quality.
For agents, the cache-read cut is the bigger story than the headline rate. Long-running agents reread the same context on every step, so Claude Opus 5.5 pricing rewards exactly the workload pattern that made earlier Opus bills frightening. If your agents run on Opus, re-run your token math this week rather than at renewal.
What do older Claude Opus models cost?
Every Opus generation from 4.6 onward remains deployable, and every one of them now costs more than the newer model:
| Model | Release date | Input / output (per MTok) | Cache read | Batch |
|---|---|---|---|---|
| Claude Opus 5.5 (current) | Sep 22, 2026 | $4.00 / $20.00 | $0.20 | $2.00 / $10.00 |
| Claude Opus 5 | Jul 24, 2026 | $5.00 / $25.00 | $0.50 | $2.50 / $12.50 |
| Claude Opus 4.8 | May 28, 2026 | $5.00 / $25.00 | $0.50 | $2.50 / $12.50 |
| Claude Opus 4.7 | Apr 16, 2026 | $5.00 / $25.00 | $0.50 | $2.50 / $12.50 |
| Claude Opus 4.6 | Feb 5, 2026 | $5.00 / $25.00 | $0.50 | $2.50 / $12.50 |
The 2026 Opus lineup prices in reverse. The older the model, the more you pay per token: Opus 4.6 through Opus 5 all bill $5/$25, while the newest model bills $4/$20. A team pinned to Opus 4.8 or Opus 5 pays 25% more per token than the current model charges, 150% more per cache read, for output Anthropic’s own benchmarks place below it.
A version pin needs a reason: a compliance requirement, or a migration ticket with a date on it. Without one of those it is just inertia with an invoice.
One genuine migration trap, so the ticket does not backfire: Anthropic notes that models from Claude 4.7 onward use a newer tokenizer producing roughly 30% more tokens for the same text. Teams still on Opus 4.6 should model token counts, not just rates, because a lower per-token price applied to more tokens can quietly cancel itself in a forecast that only changed one variable.
Earlier generations, Claude Opus 4.5 (the model in that viral $12,000 story) and Claude Opus 4.1 before it, predate the 2026 line and sit on Anthropic’s deprecation track; if a workload still runs on either, the migration conversation is overdue on quality grounds before price even comes up.
And the treadmill isn’t stopping: Anthropic’s September 22 announcement said Sonnet 5.5 and Haiku 5.5 would follow in the coming weeks. Both shipped: as of October 2026 Sonnet 5.5 runs $2/$10 and Haiku 5.5 runs $0.10/$0.50 for prompts up to 100K tokens, with Sonnet 5 and Haiku 4.5 now listed as legacy models. Bookmark the change log at the end; the dates on this page move when Anthropic’s do.
Is Claude Opus worth paying for?
Claude Opus is worth paying for when your own evaluations show it completes work that cheaper tiers fail, and it isn’t when you’re paying flagship rates for tasks Sonnet 5 or Haiku 4.5 handle on the first attempt. That answer sounds obvious, and almost nobody operationalizes it.
The internet’s verdict on Claude Opus cost is a horror genre. One developer’s widely read post-mortem describes burning $12,000 in three days on Opus 4.5, agents looping unattended, and the story has hundreds of imitators on Reddit. Read closely and the pattern is never the rate card. It’s unmetered usage: nobody set a budget, nobody watched the loop, nobody could attribute the spend until the invoice was the first anyone heard of it.
Which is why “is Opus expensive” is the wrong question. CloudZero founder and CTO Erik Peterson has been consistent on this: AI spend isn’t an expense to minimize, it’s an investment that has to show a return, and returns require knowing what each dollar bought.
$12,000 of Opus that ships a feature customers pay for is cheap. $120 feeding an agent loop nobody reviews is expensive. The rate card can’t tell you which one you are; cost per outcome can.
The full family economics, consumer plans included, live in CloudZero’s Claude pricing guide, and the cross-vendor picture in the LLM API pricing comparison. The tier decision itself gets its own section, because it is the question most Opus budgets actually turn on.
Claude Opus vs Sonnet: when is paying 2x worth it?
Claude Opus 5.5 costs exactly twice Claude Sonnet 5.5 on every meter: $4 against $2 on input, $20 against $10 on output, and $0.20 against $0.10 on cache reads. The September 22 cut narrowed a gap that had been 2.5x when Opus 5 ran $5/$25 against Sonnet’s $2/$10. The 2x is worth paying when a task fails on Sonnet or needs multiple attempts, because two Sonnet retries cost the same as one Opus pass that works.
| Workload | Route it to | Why |
|---|---|---|
| Classification, extraction, summarization, routing | Haiku 4.5 or Sonnet 5 | High volume, low ambiguity; Opus adds cost, not accuracy |
| Production app logic, most coding, drafting | Sonnet 5 | The default tier; escalate only on measured failures |
| Long-running agents, complex refactors, multi-step reasoning | Opus 5.5 | Fewer retries and fewer tokens per task; the 5% cache reads compound across agent steps |
| Frontier evaluation work where Opus falls short | Fable 5.1 | 2.5x Opus rates; Anthropic’s own line is that Opus 5.5 matches it on most work, so demand evaluation proof first |
The honest version of this comparison is an evaluation harness, not a vibe. Run both on 50 of your real tasks, count completions and retries, and divide spend by outcomes. Teams that do this usually discover their traffic is 80% Sonnet-shaped with an Opus-shaped core, which is a much cheaper answer than either model alone.
The frontier tier above Opus has its own guide in the Fable and Mythos pricing breakdown, and the Claude Code angle, where Opus arrives bundled in plan limits instead of raw tokens, lives in the Claude Code pricing guide and the Codex vs Claude Code comparison.
What does a real Claude Opus workload cost?
Worked example at standard rates: a support assistant handling 10,000 conversations daily at 500 input and 300 output tokens each, which is 150M input and 90M output tokens per month.
| Model | Monthly cost (standard) | Monthly cost (batched) |
|---|---|---|
| Claude Fable 5.1 | $6,000 | $3,000 |
| Claude Opus 5 | $3,000 | $1,500 |
| Claude Opus 5.5 | $2,400 | $1,200 |
| Claude Sonnet 5 | $1,200 | $600 |
| Claude Haiku 4.5 | $600 | $300 |
Same traffic, same month, and staying pinned to Opus 5 instead of moving to Opus 5.5 costs $600 a month for no capability you are using, before efficiency gains are counted. Add the fewer-tokens-per-task effect and the cache-read cut, and cache-heavy agentic workloads see materially more than the sticker 20%.
Two adjustments the clean math skips. Buying Opus through a cloud marketplace changes the invoice, the endpoint premiums, and who inside your company can even see the spend; the AWS route has its own economics, covered in Claude on AWS Bedrock.
And per-token rates assume the tokens are being counted at all, which the gateway problem from the opening says is not a safe assumption. Opus budgets rarely fail on the rate. They fail on the metering.
If you’d rather run this math on your actual bill than on a hypothetical support assistant, a free cloud cost assessment does it with your own numbers.
How do you reduce Claude Opus costs?
Five levers, ordered by how much they typically move the bill:
- Migrate off pinned versions. Opus 4.6 through Opus 5 cost 25% more per token than Opus 5.5 and use tokens less efficiently. This is the rare optimization that improves quality and price in the same change; model the tokenizer effect if you’re coming from 4.6.
- Batch everything that can wait. The 50% discount now lands on the lower base, taking batched Opus 5.5 to $2/$10. After every price change, re-claim the discount, because teams that set batch routing once keep paying against last quarter’s rates.
- Restructure prompts for caching. At $0.20 per MTok, Opus 5.5 cache reads reward stable prompt prefixes more than any previous Opus. The failure mode is silent: reorder a prompt so the prefix changes and every cached token bills at full input rates without a single alert. Instrument hit rates; don’t assume them.
- Route by evaluation, not by habit. Classification and extraction go to Haiku or Sonnet; Opus takes the tasks where your evaluations show the lift. Escalate on evidence.
- Meter before you optimize. None of the above is verifiable if your gateway drops attribution on some Claude calls. Fix the telemetry first, then pull the levers, then confirm each one showed up in the numbers.
How does CloudZero connect Opus spend to AI ROI?
Anthropic’s invoice reports what Claude Opus cost last month. It cannot say which product feature drove the tokens, which customer the spend served, or whether the September 22 cut actually reached your effective rate. Rate and return are different questions, and the second one is the one boards ask.
CloudZero was built for the second question. The native Anthropic integration ingests token-level Claude usage across every model on this page, the AI Hub normalizes it alongside AWS, Azure, GCP, and direct OpenAI spend, and custom Dimensions map every Opus dollar to cost per feature, customer, and team without tags.
The moments this matters are the ones Opus pricing keeps producing. A migration to Opus 5.5 lands, and anomaly detection comparing the last 36 hours against 12 months of history confirms the drop reached production, or names the workload where it didn’t.
An agent’s cache hit rate collapses on a Tuesday, and the spend spike routes to the engineer who owns it while it’s still a Tuesday problem. It is the operating model behind the customer results CloudZero publishes. Curious what that view looks like before talking to anyone? The self-guided tour walks through it, no sales call attached.
AI adoption is compounding fast enough that this stops being optional; the scale of it is documented in CloudZero’s AI statistics roundup, and the full governance framework lives in the AI spend management guide. Knowing the Opus rate card is public information. Knowing your Opus unit economics is a competitive answer only your own metering can produce.
To see your own Claude spend mapped to the features and customers it serves, schedule a demo.