Contents
The $40 million question nobody at Klarna agrees on Four buckets, one framework Unit economics is what makes the buckets real A worked example What finance leaders actually say about this How CloudZero closes the gap See what your AI spend is actually producing FAQs

Quick Answer

Proving AI business value means sorting every AI investment into one of four buckets - revenue growth, cost avoidance, productivity gain, or risk reduction - then tracking spend at the unit level (per feature, customer, or team) so each dollar has a traceable return. Most companies measure one bucket well and leave the rest unattributed. That gap is why the same AI deployment can look like a $40M win and a public reversal at the same time.

The $40 million question nobody at Klarna agrees on

In February 2024, the fintech company Klarna launched a customer service AI assistant built with OpenAI. Within its first month, it handled 2.3 million customer chats, roughly two-thirds of all customer service volume, doing work Klarna said was equivalent to about 700 full-time agents.

The company projected close to $40 million in annual profit improvement. Resolution time dropped from around 11 minutes to under 2. For a full year, this was the most-cited AI customer service deployment in the industry, repeated on earnings calls and in competitors’ pitch decks.

Then the story splits, depending on which Klarna statement you read.

AccountThe finding
The walkbackKlarna leaned too hard into cost. “What you end up having is lower quality,” Entrepreneur CEO Sebastian Siemiatkowski told Bloomberg in May 2025, and the company began reversing course and hiring human customer service agents again.
The earnings reportCustomer service costs per transaction dropped 40% over two years, from $0.32 per transaction in Q1 2023 to $0.19 in Q1 2025, CX Dive according to Klarna’s Q1 2025 release. Consumer satisfaction remained steady after the AI assistant rollout, CX Dive presented as a success.
The denialSiemiatkowski later pushed back on “reversal” framing: the AI was still handling about 1.3 million errands per month, equivalent to about 800 people Big Technology, up from the original 700 figure. Klarna’s statement: “Klarna is not reversing on AI… Our AI assistant now performs the work of over 800 full-time roles (not 700), and that number continues to grow.”

All three can be technically true at once, and that’s exactly the point.

Klarna measured cost avoidance cleanly and reported it with confidence. It never built an equally rigorous way to measure the risk and quality bucket, the thing that made outside observers read the story as a failure even while the company’s own cost data said otherwise. 

Without a framework that tracks every bucket with the same rigor, one set of facts can support three different headlines, and even the company living the story can’t settle which one is correct.

That’s the exact failure mode the rest of this article is built to catch.

Four buckets, one framework

Every legitimate AI investment produces value in one of four measurable ways. Sort a proposed AI feature into one of these before spending a dollar, not after, or you don’t have an investment case yet. You have an idea.

BucketWhat it looks likeThe test
Revenue growthA feature that helps upsell, retain, or open a new segmentCan you trace a dollar of revenue back to it?
Cost avoidanceWork that used to need a person, a vendor, or more infrastructure, and now doesn’tWhat would this have cost without AI this quarter and not hypothetically?
Productivity gainExisting people shipping more at the same headcountWhat is your team shipping now that it wasn’t a year ago, at the same size?
Risk reductionFraud caught, compliance held, an outage avoidedThe hardest to quantify, because nothing happening is hard to put a number on?

Klarna’s case only had one of these four columns filled in. Cost avoidance was measured well. 

Risk, the brand and retention cost of quality slipping on the interactions that mattered most, was never tracked with the same discipline, which is exactly why the same deployment reads as both a $40 million win and a public reversal depending on who’s asked.

Unit economics is what makes the buckets real

Buckets alone don’t prove anything. What actually connects an AI dollar to an outcome is unit economics: cost per customer, cost per feature, cost per transaction, tracked consistently enough that spend and outcome sit next to each other instead of living in separate reports.

CloudZero’s own cloud cost optimization framework already treats “measure business value” as a required step, not an afterthought, and the mechanics carry over directly to AI spend: allocate cost to the feature or customer that generated it, then hold that number next to the outcome it produced. 

CloudZero’s CostFormation approach does this by tagging AI spend with the metadata that actually matters, which feature, which customer, which team, so a single aggregate AI bill stops being one number and becomes several, each attributable to a bucket above.

A worked example

Say a SaaS company adds an AI-powered search feature.

The AI spend for the quarter comes in at $18,000. Broken down by unit economics: $6,000 attributable to a 4% lift in trial-to-paid conversion among accounts using the feature (revenue growth), $5,000 attributable to a support-ticket reduction that avoided one additional hire this quarter (cost avoidance), and the remaining $7,000 sitting unattributed, because usage data on that portion is incomplete.

That last number is the part most companies leave out. Proving value doesn’t mean every dollar comes back clean. It means being able to say exactly how much of your AI spend is accounted for and how much isn’t, instead of reporting one unexplained total.

What finance leaders actually say about this

In CloudZero’s 2026 AI ROI survey of 260 finance leaders, 34% said they could not produce a credible ROI number for their AI spend, and 43% said they face board or investor demands for AI ROI proof they currently cannot provide. Separately, 64% said tying spend to outcomes would fundamentally change how they invest, and 66% said their boards already tie AI funding to demonstrated returns.

Read those together and the pressure is coming from both directions at once. Two-thirds of boards are asking a question that a third of finance leaders cannot currently answer, which is the exact gap Klarna’s story played out in public.

How CloudZero closes the gap

The mechanics above, mapping spend to buckets, tracking it at the unit level, are exactly what CloudZero’s OpenAI spend tracking work is built on, and the same approach extends across every AI vendor a company runs, not just one.

CloudZero’s own data on this problem is blunt: only 32% of companies allocate even half their AI bill to the teams, features, or customers that actually drove it.

Most organizations are trying to prove value with two-thirds of their spend sitting in an unattributed bucket, which is the same blind spot that made Klarna’s cost data and its public narrative pull in opposite directions.

The fix isn’t a better spreadsheet. It’s metadata at the point of spend, tagging cost by feature, customer, and team as it happens, so the unit economics exist automatically instead of getting reconstructed by hand every quarter.

That’s the same foundation covered in the cloud cost optimization pillar and in how CloudZero approaches reducing AI spend without losing the ability to say what that spend bought.

See what your AI spend is actually producing

A number without a bucket isn’t proof of anything, and neither is a headline. If your AI spend can’t currently be traced to revenue, cost avoidance, productivity, or risk, that’s the gap worth closing before the next budget cycle, not after.

to see your AI spend mapped to real outcomes, run a free cloud cost assessment to find out where your spend is currently unattributed, or take the self-guided product tour to see how the mapping works.

FAQs