Contents
What is AI COGS, and where does it sit in the P&L? How much margin compression is AI actually causing? What does AI margin compression look like in practice? Why does AI COGS behave differently from hosting? What does compression do to Rule of 40 and valuation? Will falling model prices fix AI gross margins? How do CFOs respond to AI margin compression? Frequently asked questions about AI gross margin

Quick Answer

AI gross margin is what remains of SaaS profitability after inference, model routing, and AI infrastructure land in cost of revenue. The numbers have moved. AI products averaged 45% gross margin in 2025 and are projected near 53% in 2026, against the 70% to 85% that SaaS built its valuations on. The compression is real, measurable, and manageable for companies that can see their cost to serve.

Something unusual is happening in the disclosure record, and it started with the regulator rather than the filers. In February 2026 the SEC asked Doximity to “quantify where possible” the higher AI infrastructure investment it had named on an earnings call. Doximity’s written response committed to “enhanced disclosures regarding the impact of AI-related expenditures on cost of revenue and gross profit” beginning with its 2026 Form 10-K, while reporting that for the nine months ended December 2025 “such expenditures were not material.” Wix’s CFO told investors in May that Creative Subscriptions gross margin “was stable as AI costs remained minimal.” The disclosure is arriving because regulators are asking for it, and so far almost nobody has answered with a number.

The benchmark data explains why so few can. ICONIQ’s July 2026 State of AI report shows gross margins “jumping from 45% in 2025 to a projected 53% in 2026, and 59% in 2027”: improving, but still far below the classic software profile. Two-thirds of the companies surveyed report improved per-query unit economics, credited to “managing inference costs, improved model routing strategies, and revenue growth that creates cost leverage.” Bessemer’s February 2026 pricing playbook is blunt about it: “Every AI query incurs real compute costs. Companies see 50-60% gross margins vs. 80-90% for SaaS.” Its 2025 AI benchmarks go further at the extreme, putting the fastest-scaling “supernova” cohort at roughly 25% gross margin and often negative.

So the question for anyone running a SaaS P&L isn’t whether AI compresses gross margin. It’s how much, where it lands in your statements, and what separates the companies managing the compression from the ones just reporting it.

The disclosures, the benchmarks, and the two-year trend all point at the same diagnosis, and it isn’t the size of anyone’s model bill. It’s that a margin which moves with individual customer behavior can’t be read off a P&L that doesn’t track customers individually.

What is AI COGS, and where does it sit in the P&L?

AI COGS is the direct, variable cost of serving customers with AI: inference calls, model routing, vector databases, evaluation runs, and the AI infrastructure behind customer-facing features. The classification test is the customer: when product usage triggers the cost, it belongs in cost of revenue, exactly like hosting. When employees trigger it through internal copilots and productivity tools, it’s operating expense.

That line matters more than it used to, because that is where the compression becomes visible or gets hidden. Companies that bundle inference into a general “infrastructure” line report cleaner margins today and inherit harder analyst conversations later; companies that classify honestly report the compression and, increasingly, earn credit for the discipline. The disclosure trend from earnings season points in one direction.

The line items hide in different places. API inference bills arrive from model vendors directly, cloud AI services land inside the cloud invoice where they are easy to misfile as generic infrastructure, and vector databases and embedding pipelines read like storage until you trace what triggers them.

Evaluation runs are the genuinely debatable item: the testing that keeps AI features safe to ship. The emerging disclosure practice treats customer-facing eval as cost of revenue, since the product cannot function without it.

The full picture, item by item:

Cost itemWhat triggers itClassificationWhere it hides
API inference calls for customer-facing featuresCustomer product usageCost of revenueDirect invoices from model vendors
Cloud AI services (Bedrock, Azure OpenAI, Vertex)Customer product usageCost of revenueInside the cloud invoice, misfiled as generic infrastructure
Model routing and gateway computeCustomer product usageCost of revenueCompute line, alongside application services
Vector databases and embedding pipelinesCustomer queries and ingestionCost of revenueReads as storage until you trace the trigger
Evaluation runs for shipped featuresProduct operationCost of revenue (emerging practice)The debatable item
GPU infrastructure serving the productCustomer product usageCost of revenueReserved or on-demand compute
Internal copilots and employee AI toolingEmployee usageOperating expenseR&D or G&A
Model training and fine-tuning for future featuresRoadmap workOperating expense (R&D)R&D compute
Experimentation and prototypingEngineering explorationOperating expense (R&D)R&D compute

The classification exercise is worth an afternoon with your controller, because every item filed wrong distorts two numbers at once: the gross margin investors read and the OpEx efficiency story next to it.

CloudZero has argued for years that COGS is a blunt instrument for SaaS and that unit costs tell the real story. AI sharpens that argument considerably: a COGS line that swings with customer behavior isn’t just blunt, it’s unreadable without the per-customer view underneath it. The fundamentals of SaaS COGS still apply; the stakes just changed.

How much margin compression is AI actually causing?

Against the traditional 70% to 85% target covered in our SaaS gross margin guide, the 2026 picture stacks up in three tiers:

Company typeAI’s roleAverage gross margin, 2026
AI-augmented SaaSInternal copilots, minimal customer-facing AIAbout 80%, largely intact
AI-enabled SaaSAI features inside a traditional product60% to 79%
AI-nativeThe model is the product50% to 59%

The independent benchmark data complicates the picture rather than confirming it. Benchmarkit’s 2026 SaaS and AI-native metrics report finds median software gross margin holding above 80% and stable across four years, concluding that “industry-wide AI infrastructure costs have not yet compressed software margin at the median.” The compression is real in AI-native cohorts and not yet visible in the public-SaaS median. Which of those two describes your business is not something you can read off an aggregate COGS line, which is the actual problem.

Two things are true at once, and holding both is the whole point. AI attaches a real variable cost to every user interaction, so some compression is structural and will not fully reverse. But it responds to management: ICONIQ’s average moved from 45% in 2025 to a projected 53% in 2026, and two-thirds of the companies surveyed credit better inference-cost management and model routing for improving per-query unit economics.

Companies aren’t stuck with their first margin, provided they can see what’s driving it. Where your targets should sit within the bands is covered in our SaaS gross margin benchmarks.

The CloudZero read on the bands: they sort companies by visibility maturity as much as by business model. Every band holds teams that know their cost per feature and teams reading aggregate COGS, and our view is that the first group tends to climb through its band while the second rides it down. The band you’re in is mostly category. Where you sit inside it is mostly instrumentation.

For AI-native founders, the bands are not a problem to escape; they are the business model to underwrite. A 55% gross margin with visible, improving unit economics and growth that compensates on Rule of 40 math is fundable. The unfundable version is the same margin presented as temporary, with no per-feature evidence of what would change it.

What does AI margin compression look like in practice?

Illustrative math: picture one $100M ARR SaaS company walking through the tiers.

StageRevenueTraditional COGSAI COGSGross margin
Pre-AI baseline$100M$20M$080%
AI features shipped$100M$20M$8M (inference, routing, vector DB)72%
AI becomes the core workflow$100M$20M$16M64%
Recovery levers applied$100M$20M~$10M~70%

Sixteen points of compression, no pricing change, no headcount change, and each stage shipped a more valuable product than the one before. That’s the trap in miniature. The compression isn’t caused by failure. It’s caused by success that nobody priced or instrumented, which is why the same walk looks completely different at companies that did both.

The walk also shows why averages deceive. At stage two, an aggregate 72% hides customers being served at anywhere from 40% to 79% margin depending on their usage, and the spread widens at stage three. The aggregate is a report; the per-customer distribution is a decision.

Read the same walk in unit terms and it changes character. Stage three’s $16M is alarming as an aggregate and legible as a set of unit costs: cost per workflow, per feature, per customer, some earning their COGS and some not. Aggregate compression says cut. Unit compression says where.

That fourth row is the one most versions of this story leave out. Apply the standard recovery levers to stage three: route most traffic to right-sized models, cache the stable context, batch the asynchronous work, and add a usage component to pricing. The same company plausibly lands near $10M of AI COGS at 70% gross margin, six points recovered without shipping less AI.

That trajectory isn’t wishful; it’s roughly what ICONIQ’s improvement from 45% to a projected 53% traces at the market level. Compression arrives with the first AI feature. Whether it persists depends on whether anyone measures which features are causing it.

Who feels it worst: products with unlimited plans and enthusiastic users, free tiers with AI features enabled, and any vertical where one power account can consume like a cohort. If your pricing has no usage component and your product has agents, you’re running that walk with the recovery stage removed. The fix starts with knowing which of your accounts are the cohort-sized consumers, a five-minute query once allocation exists.

Why does AI COGS behave differently from hosting?

Because it scales with enthusiasm, not with headcount, and it does not average out. Traditional hosting costs grew sub-linearly: more customers, better utilization, declining unit cost. AI inference scales linearly or worse with usage, and ICONIQ’s 2026 State of AI data captures the uncomfortable direction: model inference climbs from 20% of AI product costs pre-launch to 23% at general availability and at scale, while talent’s share falls from 32% to 26%, because adoption grows consumption faster than efficiency shrinks it.

Cost behaviorTraditional hostingAI COGS
Scaling with volumeSub-linear, utilization improves with scaleLinear or worse, each interaction carries a variable cost
Unit cost over timeDeclines as you growFalls per token, but total spend rises as consumption grows
Per-customer varianceNarrow, similar plans cost similarly to serveOrder-of-magnitude spread on identical plans
Primary driverHeadcount and customer countUsage enthusiasm, prompt habits, agent call depth
Cost-to-serve cadenceAnnual finance exerciseMonthly, with a wider error bar
LeversCapacity planning, reserved commitmentsModel right-sizing, caching, batching, context discipline, pricing

AI cost to serve also varies per customer in ways hosting never did. Two accounts on identical plans can generate order-of-magnitude different inference costs based on prompt habits, feature mix, and whether their workflows chain agent calls. The classic cost to serve question (“what does this customer cost us?”) used to be an annual finance exercise. AI made it a monthly one with a wider error bar.

Ben Murray of The SaaS CFO frames the shift in one line: “If SaaS is about margin efficiency, AI is about value density.” The inference cost drivers are the mechanics underneath the finance problem: agent multiplication, context growth, model choice, and token volume. They’re all architectural, which means they’re all manageable, though only from inside a per-feature view.

There’s also a reason this problem feels familiar to anyone who lived through the last one. Cloud did this to COGS two decades ago: infrastructure went from fixed to variable, finance lost line of sight, and unit economics had to be rebuilt around consumption. CloudZero exists because of that first wave.

AI is the second wave, with steeper per-customer variance and shorter feedback loops, and the playbook transfers: allocate the variable cost to the units the business sells, then manage the units. What doesn’t transfer is the timeline. Cloud gave companies years to build the discipline. AI compresses margins in quarters.

What does compression do to Rule of 40 and valuation?

It breaks peer comparison. A SaaS business growing 25% at an 80% gross margin scores meaningfully better on Rule of 40 math than the identical business at 67%, even with the same free cash flow profile, which is why Bessemer urges AI founders to “focus less on legacy benchmarks like Rule of 40 or gross margin expansion, and more on unit economics that balance growth with compute efficiency.” Sophisticated investors increasingly compare within gross-margin bands rather than across them.

There’s a structural read gaining currency, too. Bessemer’s Talia Goldberg puts it plainly: “COGS (cost of goods sold) is the new CAC. In SaaS, the constraint was customer acquisition cost. In AI, the constraint is compute cost.” The dollars migrate rather than disappear, but valuation frameworks built on where the dollars sit do not forgive the move.

The practical implication for operators is narrower and more useful: know your compression, name its cause, and anchor the narrative in unit numbers, because the companies doing that are getting analyst credit while the bundlers get questions. Our margin analysis guide covers the reporting mechanics, and the disclosure choice compounds: a clean AI cost line this year becomes a credible efficiency trend next year, which is a valuation asset no bundled line can produce.

The diligence questions have already standardized, so it’s worth knowing them before someone across the table asks.

Diligence questionWhat it’s actually testing
What’s your inference cost as a share of revenue, and its trend?Whether you measure AI COGS at all
What does your per-customer cost distribution look like, not the average?Whether allocation exists below the aggregate
How has pricing responded to AI COGS?Whether compression was a decision or an accident
How mature is your model routing?Whether you can capture price deflation
Can you show cost per feature for the AI capabilities you sell hardest?Whether engineering and finance read the same numbers

None of the five can be answered from an aggregate COGS line.

Will falling model prices fix AI gross margins?

They’ll help, and they won’t save anyone. The input-side tailwind is real: per-token prices keep collapsing, vendor inference economics keep improving, and fast: SemiAnalysis reported in May 2026 that Anthropic’s gross margins on inference infrastructure rose from 38% to over 70% while its ARR went from $9B to over $44B, and yesterday’s frontier rates keep becoming today’s bargain tier.

The reason deflation alone does not restore margins is consumption. Cheaper inference gets used for more things, agents multiply calls, and features grow richer, so total AI COGS tends to grow even as its unit price falls, the same paradox our inference cost guide walks through in depth.

Companies that capture the deflation are the ones whose routing notices the price drops and whose measurement shows which features to scale. Without both, the savings get absorbed by consumption before anyone sees them.

How do CFOs respond to AI margin compression?

Four moves, in rough order of dependency:

Classify honestly, then instrument beneath the line

Put customer-facing AI costs in cost of revenue, break out inference where it’s material, and build AI unit economics underneath: cost per customer, per feature, per workflow, with AI cost monitoring watching the trend between board meetings. The aggregate margin tells you compression happened; the unit view tells you what to do Tuesday. Our SaaS unit economics guide covers the framework.

Price to recover, deliberately

Compression that pricing never responds to is a choice, usually an accidental one. Whether the answer is usage components, credits, or premium AI tiers is a pricing-model decision with its own tradeoffs, and it only works when cost per customer is known first.

Route and cache to the tailwind

Four savings levers routinely recover meaningful points of margin: model right-sizing, caching, batch processing, and context discipline. They’re covered in our AI cost optimization guide, and every one of them requires knowing where the spend concentrates.

Report the story you can defend

Six points of compression with a named cause, a unit-cost trend, and a recovery plan reads as strategy. The same six points bundled into “infrastructure” reads as drift.

The board version of that conversation is worth rehearsing, because it’s coming either way. The weak deck says margins fell because AI is expensive. The strong deck names the feature. Drift found that the feature serving custom messages to site visitors was hurting its margins more than anything else in the product, cut that feature’s cost by 80%, and took $2.4 million out of cloud spend in the process. Same compression, two very different rooms. The second deck exists because someone instrumented the spend before the board asked.

All four moves share a dependency, and it’s the one this entire conversation keeps arriving at: per-customer, per-feature cost visibility. That’s CloudZero’s ground. The platform allocates AI spend of every kind (API, cloud AI services, and GPU infrastructure) alongside everything else in the cloud bill, down to the customers and features that drive it, so gross margin stops being something you discover at close.

One design principle matters more than any feature: engineering and finance have to read the same numbers. Margin is compressed by architectural decisions and recovered by architectural decisions, so a margin view that only finance can see recovers nothing.

Leading global organizations such as Upstart, Progress Software, and Drift run margin-sensitive products on exactly that view, where “our margin compressed” comes with “here’s which feature, which customers, and here’s the plan.” Request a demo to see your AI cost to serve by customer and feature.

Frequently asked questions about AI gross margin