Quick Answer
Indirect costs are expenses that keep your product running but cannot be traced to a single customer, feature, or workload: shared clusters, observability, staging, and now AI inference. In 2026, companies will spend $23.3 billion running inference on cloud AI infrastructure, more than the $19 billion they spend training models there. Unallocated, that spend overstates margins and misprices products. The fix is allocating 100% of spend to business drivers.
In August 2026, Gartner put a number on a shift finance teams had been feeling all year. Companies will spend $23.3 billion running AI inference on cloud AI infrastructure in 2026, more than the $19 billion they’ll spend training models there. In that market, operating AI is now more expensive than building it.
Inference is what a model does every time it answers a request in production. The bill for it never stops.
That crossover quietly changes whose problem AI spend is. Training is a project, with an owner, a budget line, and an end date. Inference is a meter that runs every time a feature, an agent, or an employee calls a model, and the charge lands on one shared bill.
So a spending category growing nearly seven times faster than the overall IT budget (96% against 14.2%, on Gartner’s 2026 forecasts) now behaves like office rent. Everyone consumes it. Nobody owns it, and at most companies nobody allocates it.
We’ve watched this pattern take shape across the billing data of hundreds of cloud businesses, and documented it in our State of AI Costs research. The companies that treat AI spend as a first-class allocation problem can answer the question their board is now asking: what did we get for it? The companies that leave it in a shared bucket are reporting margins they can’t defend.
That gap between who consumes AI and who answers for it is the story of indirect costs in 2026.
What are indirect costs?
Indirect costs are the shared or supportive expenses that keep your platform running but can’t be cleanly assigned to one cost driver. The defining test is traceability. A cost that exists because a specific customer or feature exists is direct. A cost that exists because the business exists is indirect. They’re the connective tissue of the business, essential to everything and attributable to nothing.
Traditional accounting files them under overhead: rent, utilities, administrative salaries. A cloud business carries a different list. Shared Kubernetes clusters, data pipelines, logging systems, and internal developer platforms all support every product without belonging to any of them.
Then there’s the newest entry. A company-wide LLM subscription that powers search, support automation, and three internal copilots is an indirect cost by any classical definition. It just grows much faster than anything else in the overhead bucket.
How do indirect costs work in traditional accounting?
In standard cost accounting, indirect costs are pooled and then spread across output using an indirect cost rate, also called an overhead rate:
Indirect cost rate = total indirect cost pool / total allocation base
If a company pools $2 million of indirect cost against an $8 million direct labor base, the indirect cost rate is 25%. Every direct dollar carries 25 cents of overhead.
Two distinctions get confused here. Indirect is not the same as fixed: a shared inference bill is indirect and variable at once. And under GAAP, indirect production costs are absorbed into inventory rather than expensed in the period they occur, while general and administrative overhead is expensed as incurred. Cloud businesses inherit the same mechanics with different pools.
The defining test is traceability. A cost that exists because a specific customer or feature exists is direct. A cost that exists because the business exists is indirect.
Report
Finance needs to prove AI’s return: CloudZero report
260 senior finance leaders (more than half CFOs) told us why the speed of seeing AI spend, not the size of it, separates who pulls ahead on AI from who gets burned.
Direct vs indirect costs: what’s the difference?
Direct costs can be attributed to delivering a specific product or service, while indirect costs support operations across teams and resist clean attribution. Both belong in your cost of goods sold if the margin number is going to be honest.
| Dimension | Direct costs | Indirect costs |
|---|---|---|
| Traceability | Tied to one customer, feature, or workload | Shared across teams, features, or the whole business |
| Behavior | Scale roughly with usage | Persist through churn; grow in steps |
| Cloud examples | Compute for a customer environment, per-feature database storage, embedded services like Auth0 or Stripe | Shared clusters, logging pipelines, staging environments, company-wide LLM subscriptions |
| Ownership | Clear team or product owner | Often no single owner |
| Margin impact | Visible in unit economics | Hidden until allocated |
In practice, the direct side is the easier half of the ledger. It scales predictably, engineers can optimize it at the code level, and it dominates most reported COGS figures. When usage doubles, the bill roughly doubles, and everyone can see why.
Indirect spend follows different physics. It doesn’t fall when customers churn, and it doesn’t rise in a straight line when usage grows. A staging environment costs the same whether you closed ten deals last quarter or none.
The margin damage comes from that asymmetry. Every unallocated indirect dollar makes your gross margin look better than it is, right up until someone asks you to defend the number in a board meeting or a pricing negotiation.
Why is AI spend the fastest-growing indirect cost of 2026?
AI spend has become the fastest-growing indirect cost because it combines record growth with the classic shared-cost profile: metered consumption, many consumers, and no natural owner. No other expense category checks all three boxes at once.
The growth is the headline. Gartner forecasts worldwide AI spending of $2.59 trillion in 2026, up 47% year over year, a figure the firm already revised upward once, from $2.52 trillion in its January outlook. Gartner expects infrastructure alone to account for over 45% of the total.
The enterprise layer tells the same story at smaller scale:
- Spending on AI models and platforms will reach $64 billion in 2026, up 63.4% from $39 billion in 2025, with generative AI model spending growing 117% on its own
- AI-optimized infrastructure as a service will grow 96% to $42 billion this year, heading for $66 billion in 2027
- 55% of that infrastructure supports inference in 2026, a share Gartner expects to hit 59% in 2027
- Spending on AI-optimized servers will rise 49%, making up 17% of all AI spending
The growth concentrates in inference, serving, and platform usage. These are consumption charges. Gartner analyst Hardeep Singh names two drivers of that growth, continued LLM training demand and “the rapid operationalization of AI across enterprise applications and workflows.” The second one is a precise description of spend going indirect.
Why is AI spend indirect by default?
AI spend defaults to indirect because the billing unit (a token, a GPU-hour, a seat) rarely matches the business unit you manage (a feature, a customer, a team). Three patterns drive the mismatch.
Shared subscriptions come first. One OpenAI API account can feed a search feature, a support bot, and a personalization engine at the same time. The invoice shows tokens consumed. It says nothing about which product consumed them, or whether the consumption earned its keep.
Agentic workloads compound the problem. An agent completing a task can spawn sub-agents, retry failed steps, and burn a variable number of tokens for the same nominal unit of work. Costs that vary per execution resist the flat allocation formulas finance teams inherited from the data center era.
Seat-based AI tools add a third layer. Coding assistants, meeting summarizers, and ChatGPT seats get purchased team by team, often outside central procurement. Individually small, they accumulate into a company-wide line item that no single budget explains.
The result mirrors what happened with shared cloud infrastructure a decade ago, compressed into a much shorter window. Unless someone deliberately allocates it, AI spend makes AI features look cheaper than they are, and the distortion flows straight into product profitability and pricing decisions.
What happens when nobody allocates AI spend?
Unallocated AI spend breaks the AI ROI calculation before it starts, because a return needs a cost basis and a shared bucket doesn’t provide one. Finance leaders have started to say this out loud.
Gartner analyst Arunasree Cheparthi notes that “enterprise AI budgets are coming under greater scrutiny,” with spending shifting toward providers that can demonstrate clear value. John-David Lovelock, in Gartner’s May forecast, points at where the market still sits: “Enterprises have yet to really flex their spending potential.”
Both observations point at the same gap. Boards approved the experimentation era. The next round of AI budget requests will be judged against results per feature, per team, and per customer, and only allocated spend can produce those denominators.
Getting the allocation right requires knowing the other indirect categories AI spend sits alongside, because the same machinery handles all of them.
What are examples of indirect costs in the cloud?
Beyond AI, indirect cloud costs cluster into six established categories, each essential to operations and none attributable to a single revenue stream.
| Category | What it includes | Why it stays indirect |
|---|---|---|
| Shared infrastructure | Multi-team Kubernetes clusters, VPCs, load balancers, service mesh, caching layers | Serves many workloads; rarely tagged to one |
| Internal environments | Staging, QA, sandbox, and test environments | No end user to attribute them to |
| Observability and security | Monitoring, logging, tracing, alerting, compliance tooling | Protects value rather than producing it |
| Data platform | Shared warehouses, ETL pipelines, BI dashboards, data lakes | Mixed ownership across product, finance, marketing |
| Core services | Auth, billing, notifications, feature flags, shared SDKs | Scales with total usage, not one feature |
| Networking and management overhead | NAT gateways, cross-region traffic, backups, egress, CDN | Baked into the operational layer of the cloud |
Two quieter layers sit underneath the technical six. Internal SaaS tooling (collaboration, project tracking, CRM, design) rarely appears in COGS conversations despite scaling with headcount. And people-driven costs, from governance reviews to incident response, shape cloud efficiency without ever showing up on a cloud bill. Mature allocation models fold both layers in, because a margin number that ignores them is only mostly true.
None of these categories is new, and finance teams have workable playbooks for most of them. What’s changed is the mix. When a fast-growing, usage-metered AI layer joins a stack of slow-moving shared services, the old habit of parking 20% of spend in “shared” stops being a rounding error.
Why do unallocated indirect costs distort your margins?
Unallocated indirect costs inflate reported margins, misprice products, and erode accountability, and each effect compounds as the business scales.
- The margin illusion comes first. When shared infrastructure and AI inference sit in a miscellaneous bucket, your cost per customer reads better than reality. A product line can look comfortably profitable while shared spend quietly consumes the spread.
- Pricing inherits the error. Without the true cost to serve, you cannot price with confidence, and you cannot tell how much discount a renewal can absorb before the account goes underwater. Deals get won at margins that only exist on paper.
- Accountability erodes next. Costs that are shared but unowned become financial technical debt: every team consumes the platform, no team answers for it, and the bill grows in the dark. Engineers can’t optimize spend they never see attributed to their work.
- Scale then multiplies everything above. More regions, more environments, more model calls. Gartner’s Lovelock describes the AI buildout as “the largest infrastructure project ever attempted by humanity” in the firm’s latest IT spending forecast, which projects $6.37 trillion in worldwide IT spending this year, up 14.2%. An unallocated slice of a number growing that fast is a blind spot with a growth rate.
How do you allocate and manage indirect costs?
Allocating indirect costs takes five steps: define shared cost pools, choose allocation bases people trust, push to 100% allocation, build accountability through showback, and automate the loop so it survives contact with real usage.
1. Define and group shared cost pools
Name the major categories of indirect spend before trying to divide them: platform services, observability, data platform, internal tooling, and AI inference as its own pool. In CloudZero, tagged, untagged, and untaggable spend is detected automatically and grouped into dynamic cost pools, which makes hidden spend visible before allocation math begins.
Giving AI its own pool matters more than it looks. Blending inference into “data and analytics” hides the one category your board will ask about by name.
2. Choose allocation bases people trust
Distribute each pool using an allocation basis, the measure you divide a pool by so each team’s share reflects what it actually consumed. Pick one the affected teams accept as fair.
| Cost pool | Default allocation basis | Fallback when usage data is missing |
|---|---|---|
| Platform and shared infrastructure | Compute hours or vCPU-hours consumed | Proportional to each team’s direct costs |
| Core services (auth, billing, notifications) | API calls per consuming service | Request-weighted share of total traffic |
| AI inference and model APIs | Tokens consumed per feature or customer | Share of feature-level request volume |
| GPU and training infrastructure | GPU-hours per job, tagged to the requesting team | Reserved-capacity share per team |
| Observability and security | Ingested log or metric volume per service | Proportional to each service’s compute spend |
| Data platform | Query or compute credits per consumer | Active dashboards or pipelines per team |
| Internal SaaS and seat-based AI tools | Assigned seats per team | Headcount |
Consistency beats mathematical perfection here. A simple rule every team understands outperforms an elegant formula nobody trusts, and the full menu of approaches is in our guide to cost allocation methods.
3. Allocate 100% of your spend
Most teams stop at the easy 60% to 80% of directly attributable spend. High performers push to 100%, tying every dollar, shared and AI included, to a business driver. Whatever remains in “shared” is a blind spot in your COGS, your cost per customer, and your AI ROI math.
Perfect tagging is not the prerequisite it appears to be. Remitly allocated 50% more of its cloud spend without tagging by mapping costs through business logic instead of resource metadata.
4. Build accountability with showback, then chargeback
Visibility changes behavior before any money moves. Showback gives each team a clear view of its share of common costs without transferring financial responsibility, and spending patterns start improving on visibility alone. Chargeback, where teams absorb what they consume, comes later as maturity grows.
The showback versus chargeback decision depends on culture and incentives. The sequence, visibility first and ownership second, rarely changes.
5. Automate and refine continuously
Manual allocation cannot keep pace with cloud billing, and token-metered AI usage moves faster still. Spreadsheets break, tags drift, and a formula tuned in January misses the agent workloads that launched in April.
Automate ingestion and allocation, then review the model quarterly. Architectures evolve, new shared services appear, and an allocation model is only as good as its resemblance to how teams work this quarter.
How does CloudZero turn indirect costs into AI ROI?
CloudZero ingests spend from any AI, cloud, or SaaS provider, allocates 100% of it in hours regardless of tagging quality, and maps each dollar to the customer, product, and team that drove it. That last step is what converts an overhead bucket into a numerator and denominator you can put in front of a board.
For the AI layer specifically, AI Hub unifies spend across model providers and GPU infrastructure, so a shared inference bill resolves into cost per feature and cost per customer.
Dimensions, CloudZero’s business-context layer, slices all of it by the concepts you manage: teams, products, features, customers, and even individual AI models. Combined with the allocation engine, the same machinery answers both halves of the question finance keeps asking: what does each AI feature really cost, and what is it returning?
The proof shows up in customer numbers. PicPay allocated 100% of its cloud costs across 20 business units and 2,000+ engineers, even in poorly tagged environments, and has since taken $18.6 million out of its annual cloud spend. Drift improved its COGS by $2.4 million. Duolingo’s engineers use the same platform to see their own costs, and a deeper treatment of the AI-specific version lives in our guide to AI cost management.
Your indirect costs are already shaping your margins. The only question is whether you can see them doing it.
See it against your own bill. Start with a free AI and cloud cost assessment to find out how much of your spend is still sitting in a shared bucket.