Contents
How does Claude pricing work? What are the Claude subscription plans? What is Claude API pricing by model? Subscription vs API: which actually costs less? What drives your Claude costs? How do you calculate Claude API costs? How does Claude pricing compare to alternatives? What does Claude Code cost? How do you reduce Claude costs? How does CloudZero track and optimize Claude spend? Frequently asked questions about Claude pricing

Quick Answer

How much is Claude? Claude pricing has two sides. Subscription plans run from $0 (Free) to $200/month (Max 20x), with Team seats from $20/seat/month and Enterprise at $20/seat plus usage. Claude API pricing is billed per million tokens: $1/$5 (Haiku 4.5), $2/$10 (Sonnet 5, introductory), $5/$25 (Opus 5), and $10/$50 (Fable 5) for input/output. Model choice is the single biggest lever on your bill.

One CloudZero customer runs more than 50 LLMs in production. They saved over $1 million on AI spend, not by negotiating better rates, but by catching runaway spend patterns early enough to stop them. The rate card never changed. What changed was knowing where the money went.

Keep that story in mind, because it is the whole game with Claude. Anthropic publishes some of the clearest rates in AI, and this guide covers every one of them, verified against Anthropic’s own pricing pages. But a rate card tells you what a token costs. It cannot tell you what your team is paying, or whether that spend is earning its keep.

As Erik Peterson, CloudZero’s founder and CTO, puts it, the question is not just “How much did we spend?” but “Was it worth it?”

Most teams cannot answer either question, and the pool of money at risk keeps growing: Gartner forecasts worldwide AI spending will reach $2.5 trillion in 2026. In CloudZero’s 2026 AI ROI survey of 260 finance leaders, 55% said they ran over their AI budget, and 32% blew past it by more than 20%. Token-billed services like the Claude API are a leading reason why.

How does Claude pricing work?

Claude bills two completely different ways, and the answer to “how much is Claude” depends on which one you mean.

  • Subscriptions are flat monthly fees for using Claude through the apps: web, desktop, mobile, and Claude Code in your terminal. Send ten messages or a thousand, the price is the same, within usage limits that reset on a rolling five-hour window.
  • The API is a meter. You pay per token processed, with separate rates for input tokens (what you send) and output tokens (what Claude generates). No monthly fee, no ceiling. This is where production spend lives, and where bills surprise people.

The two models meet in one place worth knowing about. Paid plans can turn on usage credits, which let you keep working past your plan limits at standard API rates. Heavy users effectively run both billing models on one account, often without realizing it.

What is MTok?

MTok means one million tokens, the unit Anthropic uses for API pricing. A token is roughly three-quarters of an English word, so one MTok is about 750,000 words. A rate of $5/MTok input means every million tokens you send costs $5.

What are the Claude subscription plans?

Claude plans span seven tiers as of 2026. Every paid tier includes Claude Code, Claude Cowork, Claude Design, and Claude Science, all drawing from one shared usage pool.

PlanPriceBest for
Free$0Trying Claude, everyday questions
Pro$17/month billed annually, $20 monthlyIndividuals, daily work, light Claude Code use
Max 5x$100/monthHeavy individual use, 5x Pro usage
Max 20x$200/monthPower users living in Claude Code, 20x Pro usage
Team Standard$20/seat/month annual, $25 monthlyTeams of 2 to 150, more usage than Pro
Team Premium$100/seat/month annual, $125 monthlyTeam seats with 5x Standard usage
Enterprise$20/seat/month plus usage at API ratesLarge organizations, compliance and admin controls

If you are comparing Claude pricing plans side by side, three details most guides to Claude subscription plans get wrong or skip entirely:

  • Annual billing changes every number. The Claude subscription price drops from $20 to $17/month on Pro billed annually, and Team Standard drops from $25 to $20/seat. This is why competing articles contradict each other on Claude AI pricing: they are quoting different billing cycles without saying so.
  • Claude enterprise pricing is a hybrid. Enterprise runs $20/seat/month billed annually plus usage at API rates, so cost scales with which models your teams run and how hard they run them. In exchange you get SCIM, audit logs, custom data retention, a compliance API, and a 500K context window on the default model.
  • Usage pools across everything. Chat, Claude Code, and Cowork draw from the same allowance. A developer who burns the pool in the terminal by lunch has burned it in chat too. That is manageable for one person and chaos for a 40-seat Claude subscription, which is why per-person visibility matters on Team and Enterprise plans.

What is Claude API pricing by model?

Claude model pricing is billed per MTok across four current-generation models, sometimes searched as Anthropic Claude API pricing. One ratio to memorize: output tokens cost 5x input across the entire lineup.

ModelInput (per MTok)Output (per MTok)Cache readBest for
Claude Fable 5$10.00$50.00$1.00Frontier intelligence, long-running agents
Claude Opus 5$5.00$25.00$0.50Complex agentic coding, enterprise work
Claude Sonnet 5$2.00$10.00$0.20Production workloads, coding, agents
Claude Haiku 4.5$1.00$5.00$0.10High-volume, fast, cost-efficient tasks

What the table does not say out loud:

  • The Sonnet 5 intro window is closing. Teams that modeled budgets on $2/$10 get a 50% rate increase on September 1, 2026. If your workload runs on Sonnet, forecast on $3/$15 starting now, not on the invoice that surprises you in October.
  • Opus pricing has held at $5/$25 across five straight releases. Opus 4.5 through Opus 5 all carry the same rate, a 67% cut from the Opus 4.1 era ($15/$75, now retired). The model-by-model view lives in CloudZero’s Claude Opus 4.8 pricing guide.
  • Fable 5 is the premium tier at $10/$50. It is Anthropic’s most capable generally available model, sharing rates with the limited-availability Mythos 5. CloudZero’s Claude Mythos pricing breakdown covers both.
  • The tokenizer matters as much as the rate. Models from Opus 4.7 onward use a tokenizer that produces roughly 30% more tokens for the same text. Migrate from Sonnet 4.6 to Sonnet 5 and your effective cost per request changes even where the rate card looks identical. This is the kind of shift that hides inside a fixed price and shows up three weeks later in a reconciliation meeting nobody enjoys.
  • No long-context surcharge. Models from 4.6 onward include the full 1M token context window at standard rates. A 900K-token request bills at the same per-token price as a 9K one, a structural difference from OpenAI’s long-context pricing.

Two premium modifiers round out the picture: fast mode (research preview) runs Opus 5 at 2x standard rates ($10/$50) for significantly faster output, and US-only inference adds a 1.1x multiplier on models from 4.6 onward.

Subscription vs API: which actually costs less?

This argument is playing out right now in Reddit threads titled “$200 subscription vs $7,470 of API usage” and “Claude subscriptions are up to 36x cheaper than API.” Both headlines are true. The answer is entirely about volume.

Here is the math at standard Sonnet rates ($3/$15), comparing Max 20x at $200/month against the same consumption billed through the API:

Monthly usage profileInput tokensOutput tokensAPI costCheaper option
Light (occasional chat, small tasks)5M1M$30API or Pro
Moderate (daily coding sessions)60M4M$240Max 20x
Heavy (agentic workflows, all day)200M (60% cached)10M$426Max 20x
Very heavy (parallel agents)500M (60% cached)25M$1,065Max 20x, by 5x

Heavy row math: 80M uncached input at $3 is $240, 120M cache reads at $0.30 is $36, 10M output at $15 is $150. Total $426 against a $200 flat rate.

The pattern is clean. Subscriptions are flat-rate arbitrage for heavy interactive users, which is why developers living in Claude Code gravitate to Max 20x and treat the Claude premium tiers as a discount. The API wins for light usage, for automation that needs no seat, and for anything requiring per-request cost attribution. Production applications run on the API regardless, because a subscription cannot serve your customers.

One trap from the field: if an ANTHROPIC_API_KEY environment variable is set in a developer’s shell, Claude Code bills at API rates and ignores the subscription entirely. More than one team has discovered this on the invoice rather than in the docs.

What drives your Claude costs?

Five factors determine what you actually pay, in rough order of impact.

  • Model choice. Haiku 4.5 to Fable 5 is a 10x spread on both input and output. Routing classification and extraction to Haiku while reserving Opus 5 for genuinely hard reasoning is the single largest cost decision you will make.
  • Prompt caching. Cache reads cost 10% of the base input price. A 5-minute cache write costs 1.25x base, and as Anthropic’s own pricing docs note, caching “pays off after just one cache read.” A 1-hour write costs 2x and needs two hits. For agents with big, stable system prompts, caching routinely cuts input spend by more than half. It is also the lever that silently breaks: reorder a prompt so the stable content no longer leads, and every “cached” token quietly bills at 10x, invisible until someone instruments the spend.
  • Batch processing. The Batch API halves both input and output rates for asynchronous work completed within 24 hours. Opus 5 drops to $2.50/$12.50, Haiku to $0.50/$2.50. Nightly pipelines, bulk classification, and content generation all belong in the queue.
  • Output length. Output is 5x input on every Claude model, which makes unconstrained responses the most expensive tokens you buy. Capping response length attacks the priciest part of the bill first.
  • Extended thinking and tools. Thinking tokens bill as output, at output rates. Server-side tools add their own charges, like web search at $10 per 1,000 searches. Agentic workloads stack both fast, which is why AI agent costs blindside teams that budgeted for chat-style usage.

How do you calculate Claude API costs?

The formula behind all Claude token pricing, for any model:

Cost per call = (input tokens / 1,000,000 x input rate) + (output tokens / 1,000,000 x output rate)

Worked example: a support assistant on Haiku 4.5 sends a 400-token cached system prompt plus a 100-token user message and returns a 300-token response.

  • Uncached input: 100 / 1,000,000 x $1.00 = $0.0001
  • Cached read: 400 / 1,000,000 x $0.10 = $0.00004
  • Output: 300 / 1,000,000 x $5.00 = $0.0015
  • Total per call: roughly $0.0016

At 10,000 calls per day, that is about $493/month. The identical workload on Opus 5 without caching runs about $1,875/month. Same traffic, 3.8x the bill. Anthropic’s own customer support agent guide works a similar example at roughly $37 per 10,000 tickets on Haiku.

CloudZero’s Anthropic integration runs this math continuously against your real usage, projecting monthly spend by model, feature, customer, and team instead of by spreadsheet estimate.

How does Claude pricing compare to alternatives?

Anthropic pricing sits mid-pack at equivalent tiers, with one structural advantage the table cannot show: no long-context surcharge.

ModelProviderInput (per MTok)Output (per MTok)
Claude Fable 5Anthropic$10.00$50.00
Claude Opus 5Anthropic$5.00$25.00
GPT-5.5OpenAI$5.00$30.00
Claude Sonnet 5 (intro)Anthropic$2.00$10.00
GPT-5.4 StandardOpenAI$2.50$15.00
Gemini 3.1 ProGoogle$2.00$12.00
Claude Haiku 4.5Anthropic$1.00$5.00
GPT-5.4 MiniOpenAI$0.75$4.50
DeepSeek V3.2DeepSeek$0.28$0.42

At the flagship tier, Opus 5 matches GPT-5.5 on input and undercuts it 17% on output. Below Haiku, OpenAI’s Nano tier and DeepSeek undercut everything Anthropic offers, which matters for very high-volume classification work. For the full field, see CloudZero’s LLM API pricing comparison and the OpenAI API pricing guide.

Claude also runs on Amazon Bedrock and Google Vertex AI, where regional endpoints carry a 10% premium over global ones and billing lands on your cloud invoice instead of Anthropic’s. Same model, three bills that look nothing alike. CloudZero’s Bedrock vs Claude Platform comparison covers when each path makes sense.

What does Claude Code cost?

Claude Code pricing is bundled: it is included in every paid plan, from Pro at $17/month through Enterprise, sharing the plan’s usage pool. Run through the API instead, it bills at standard model token rates with no monthly fee. There are no separate Claude Code plans to buy.

The plan-by-plan breakdown, including where each tier’s limits actually land for daily coding work, lives in CloudZero’s dedicated Claude Code pricing guide.

How do you reduce Claude costs?

Six levers, in the order most teams should pull them:

  1. Route by task complexity. Haiku for classification and extraction, Sonnet for production workloads, Opus 5 for hard reasoning, Fable 5 only where evaluations prove the lift. Tiered routing commonly cuts total spend 60% or more.
  2. Cache aggressively, then verify the hits. Structure prompts so stable content leads, and monitor the cache hit rate. A silent cache miss is a 10x price increase on that content, and nobody sends you a notification.
  3. Batch everything that can wait. 50% off both token types for any workload that tolerates a 24-hour window.
  4. Constrain output. Set response limits and instruct brevity. Output is 5x input on every model.
  5. Forecast the Sonnet 5 change now. Model September budgets on $3/$15, not the intro rate.
  6. Give every dollar an owner. Untracked token spend grows unchecked, and 30% of finance leaders in CloudZero’s survey still reconcile AI spend manually. Per-team, per-feature, and per-customer attribution is what turns the other five levers from advice into practice. CloudZero’s AI spend management guide covers the framework.

How does CloudZero track and optimize Claude spend?

Anthropic’s Console shows total usage by workspace and API key. That answers what you spent. It cannot answer which product feature, customer, or team generated the spend, or whether the spend produced anything worth having.

CloudZero was the first cloud cost platform to integrate directly with Anthropic’s Usage and Cost Admin API. The Anthropic integration ingests token-level consumption (uncached input, cache hits, output, tool usage) by model, workspace, and API key. CloudZero’s CostFormation allocation engine then maps every dollar to the business dimensions that drive decisions: cost per customer, per feature, per team, per environment. No manual tagging required, which matters because nobody tags AI experiments perfectly.

Claude spend lands in the AI Hub alongside OpenAI, AWS, Azure, GCP, and the rest of the stack, so a Bedrock-hosted Claude workload and a direct API workload roll into one normalized view instead of three billing formats.

Anomaly detection compares the last 36 hours of spend against 12 months of history, and when a prompt change or traffic spike pushes consumption outside normal, the alert goes to the engineer who owns the feature. Before the invoice arrives, not after.

That is the difference between reading the rate card and knowing your AI unit economics. The rate card says Opus 5 costs $5/MTok. Unit economics says your AI support feature costs $0.11 per resolved ticket, and your enterprise customers generate 4x the Claude spend of your mid-market ones. Only one of those numbers can run a business.

Schedule a demo to see how teams like Duolingo, Coinbase, and Grammarly connect AI spend to business outcomes. You can also get a free cloud cost assessment to find out where your AI spend stands today, or take a self-guided tour to explore cost per customer and AI allocation in action.

Frequently asked questions about Claude pricing