Quick Answer
Cost per AI outcome is your total attributed AI spend divided by the business results it produced: resolved tickets, converted leads, merged pull requests. It includes the cost of failed attempts, sits at the top of the AI unit-cost ladder, and it's the number that makes vendor outcome pricing, ROI claims, and build-versus-buy decisions comparable.
Earlier this year, CloudZero’s SVP of Engineering, Bill Buckley, told a group of engineering leaders his monthly AI budget per engineer: $5,000 in Claude Code tokens. Another leader in the same thread was capping engineers at $200. A 25x gap, between two orgs that look more alike than different.
Bill’s math is a hiring trade: a fully loaded U.S. engineer runs about $250,000 a year, and $5,000 a month per engineer is $60,000 a year. Across four engineers, that’s $240,000, roughly one new hire. Token spend, as he put it, trades directly against “payroll, recruiting, onboarding, and the year of ramp” a new hire needs. But the post’s sharpest section is its question, the one he says nobody’s asking: what did you get in return?
His post carries the cautionary tale that makes the question urgent: Uber burned through its entire 2026 AI budget in four months as its coding agent went from 1% to 8% of code changes. Spend that scales with success is exactly the spend that needs a per-result denominator.
That question has a name. Cost per AI outcome is the return question made computable: not what did AI cost, but what did each result cost.
We wrote this guide the way we run the math on ourselves, and it covers how to define an outcome you can count, how to compute its cost honestly (failures included), and how to use the number for the decisions it exists to settle.
What is cost per AI outcome?
Cost per AI outcome is attributed AI spend divided by business results delivered over the same period. It’s the top rung of the AI unit-cost ladder: a query costs tokens, a task chains queries, and an outcome is what the task was for. The rungs below are engineering metrics; the outcome rung is where AI spend meets the P&L.
The distinction from its neighbors is what gives it teeth. AI inference cost prices a single model call, and cost per task prices a unit of agent work whether or not it succeeded. Cost per outcome counts only successes in the denominator, so failures, retries, and abandoned attempts all live in the numerator, where they belong.
| Rung | What it prices | How failures count | Who uses it |
|---|---|---|---|
| Cost per query | One model call or request | Not applicable, a query is a single call | Engineering |
| Cost per task | One complete unit of agent work, successful or not | Failed tasks count in the denominator | Engineering and platform |
| Cost per AI outcome | One business result that met your written definition | Failures live in the numerator, never the denominator | Finance, pricing, vendor decisions |
That’s the design, not a technicality.
An agent that attempts everything and resolves little has a fine cost per task and a terrible cost per outcome, and only the second number tells the truth.
It’s also the rung ROI arguments need: hours saved and tickets deflected are numerators, and cost per outcome is the denominator that turns each one into a return. Our position on this is on the record: AI ROI is an allocation problem, because a return you cannot attribute is a return you cannot defend.
Report
Finance needs to prove AI’s return: CloudZero report
260 senior finance leaders (more than half CFOs) told us why the speed of seeing AI spend, not the size of it, separates who pulls ahead on AI from who gets burned.
Which AI vendors charge per outcome?
Vendors moved first. The outcome is now a billable unit with published prices attached:
| Vendor | Billing unit | Rate | Bills failed attempts? |
|---|---|---|---|
| Fin | Per outcome | $0.99, 50-outcome monthly minimum | No, but a handoff to a human counts as an outcome |
| Salesforce Agentforce | Per conversation | $2.00 | Yes, a conversation bills whether or not it resolves |
| Zendesk AI agents | Per automated resolution | $1.50 committed / $2.00 pay-as-you-go | No, resolution verified by a separate LLM check |
List prices from each provider’s first-party pricing page, checked September 2026. Enterprise rate cards, partner-cloud listings, negotiated rates, and regional pricing differ.
Salesforce paid $3.6 billion for Fin, formerly Intercom, closing the acquisition in September 2026, which is what a pricing model looks like when the market validates it. Per-conversation and per-outcome are different animals: a conversation bills whether or not it resolves anything. Fin’s $0.99 outcome is usually a resolution, but its own pricing page also counts a configured handoff to a human as a billable outcome, so read any vendor’s definition of the billable unit before you compare rates.
And these are seller prices, not your costs. Setting prices like these is the monetization question we cover in pricing AI agents.
Vendors know their cost per outcome to the cent, because their margin depends on it. Most buyers evaluating those rate cards don’t know their own. Outcome-based pricing only tilts in your favor once both sides of the table have a number.
How we run this math at CloudZero
We’re not neutral observers here, so read this section as a practitioner’s log with receipts.
The $5,000-per-engineer budget above is our spend side, published with the reasoning. The return side is the part we’ve built our product around, and the releases are public: connecting tokens to outcomes, because counting tokens is not enough, and allocating spend to outcome inside the platform, alongside the broader financial control plane for AI spend announced on our blog.
The methodology predates AI.
We’ve published step by step how CloudZero measures cost per customer on our own product, and customers run the same discipline: CleverTap uses unit economics in practice, and Drift cut $2.4 million from its AWS bill working with CloudZero. Cost per AI outcome is that same muscle applied to the newest spend on the bill.
The honest admission that comes with practicing this: the hard part was never the division. It’s the attribution underneath, which is why the outcome metric is only as good as the AI cost allocation feeding it.
How do you define an outcome you can count?
Write the definition before you measure anything, because counting rules move this number more than model choice does.
Zendesk’s documentation defines its billable unit as a request “resolved by the AI agent, without any escalation to a human agent,” verified by a separate LLM check. Every clause in that sentence is a decision you have to make too:
- What closes an outcome? Explicit customer confirmation, a re-contact window (72 hours and 5 days produce very different counts), or the agent’s own judgment.
- What about human assists? An outcome the AI carried 90% of the way is a partial AI outcome, a human outcome with AI costs attached, or its own category. Pick one, apply it everywhere.
- What never counts? Abandoned conversations, spam, and duplicates need explicit exclusion, because vendor billing disputes cluster right here.
The same rigor applies whether you’re counting your own agent’s outcomes or auditing a vendor’s invoice. A definition you didn’t write is a definition someone else’s revenue team wrote.
A workable outcome definition fits in five fields, versioned like code:
- Outcome name and owner: what result, and who answers for the count.
- Closing condition: the event that makes it final, with its time window stated.
- Human-assist rule: how partially automated results count, applied uniformly.
- Exclusions: abandoned, spam, duplicate, and test traffic, listed explicitly.
- Review date: when the definition gets revisited, because silent drift is how the metric gets gamed.
How do you calculate cost per AI outcome?
The formula is short, and the honesty lives in one denominator choice:
Cost per AI outcome = attributed AI spend for the workload ÷ outcomes achieved in the same period
Spend covers every attempt: successful runs, failed runs, retries, and the workload’s share of shared infrastructure. Outcomes count only what met your written definition. Dividing all attempts by only successes is what keeps the metric honest, because the failures were real spend in pursuit of the same results.
A worked example, chained from our agent cost analysis:
Each resolution attempt runs about 12 model calls totaling 36,000 input and 6,000 output tokens, which on Claude Sonnet 5 ($2 and $10 per million tokens) costs $0.132 per attempt. If the agent fully resolves 70% of attempts:
$0.132 per attempt ÷ 0.70 resolution rate = about $0.19 per resolved ticket
The 30% that escalate are not wasted in the ledger’s eyes; their cost is carried by the resolutions, which is exactly what dividing by outcomes does. At 10,000 monthly resolutions, that’s roughly $1,890 in model spend, before platform, retrieval, and engineering costs join the numerator.
Getting from token logs to that number is the instrumentation problem the phrase tokens to outcomes names, and it’s the join our Shipped releases above exist to automate.
What decisions does cost per AI outcome improve?
Three decisions get dramatically easier once the number exists:
- Build versus buy, priced honestly. The worked example’s $0.19 per resolution against Fin’s $0.99 looks like a 5x gap, and the gap is the price of everything the vendor carries: platform, maintenance, model risk, the engineering you didn’t hire. Sometimes that’s a bargain, sometimes it isn’t, but without your own cost per resolution the comparison is a coin flip. The same math turns vendor negotiations from vibes into arithmetic.
- ROI claims that survive finance. “The agent saved 4,000 hours” is a numerator. “Each resolution costs $0.19 against $6 to $8 of agent handling time” is a return, computable monthly and defensible in front of a CFO. This is Bill’s unanswered question, answered.
- Autonomy and architecture choices. Anthropic’s production finding that multi-agent systems consume around 15x the tokens of chat becomes actionable at this rung: a 15x cost architecture needs outcomes worth 15x more, and cost per outcome is the number that checks it.
Totals justify budgets. Outcome costs justify the program. The rung matters because it is where the CFO’s margin question and the engineer’s architecture question turn out to be the same question, asked in different units.
What’s a good cost per AI outcome?
A good cost per outcome sits comfortably below the value per outcome, and no external benchmark can stand in for that ratio. A resolved support ticket that would have cost $6 to $8 in agent time can afford $0.99 and still return well; a lead-qualification outcome feeding a $50,000 pipeline can afford far more; a free-tier feature’s outcomes have to cost pennies.
The comparisons that hold up are internal: your trend over time, your spread across outcome types, and your number against the vendor rate for the same result. One caution as the metric spreads: a falling cost per outcome with a loosening outcome definition is an illusion, and it’s the most common way this number gets gamed. Freeze the definition, then optimize the cost.
How do you track cost per AI outcome over time?
Track it monthly at minimum, per outcome type, with the definition version pinned to every reading. The prerequisite is the same one every unit cost shares: spend attributed to the workload and outcomes logged against the same workload, joined over the same window. That join is precisely what we ship, and precisely what a spreadsheet reconstruction of it quietly gets wrong.
Two practices keep the trend meaningful. Cohort by outcome type, because a blended average across resolutions, leads, and merged PRs hides every insight worth having. And pair each team’s outcome cost with its volume, so a team scaling outcomes 10x gets credit for the growth instead of a lecture about the total.
Vendors already price the outcome; the advantage goes to buyers who can price their own. Request a demo to see CloudZero connect AI spend to outcomes on your data.