I spend most of my week talking to companies about AI ROI.
A few months ago, that was still a weirdly specific conversation. Now it’s everywhere. CloudZero spends a lot of time in that conversation, so I’m glad the market is talking about it. But the conversation tends to start, and stall, in the wrong place.
There are two ideas I keep coming back to:
- Non-developer AI ROI is the bigger story, and in many cases it’s easier to measure the ROI of non-engineering AI projects than engineering AI projects.
- Developer AI ROI won’t ultimately be measured in DORA metrics. It’ll be measured in business outcomes.
That doesn’t mean developer productivity is fake. It’s very real. It just means the market bet is too big for developer productivity alone to carry the math.
The $1T AI ROI question
As the recent 20VC and SaaStr conversation framed it, the AI buildout only makes economic sense if demand catches up to the scale of infrastructure being built.
Hyperscalers are spending roughly $700B a year on CapEx. AI revenue is still well under $100B. Add power and operating costs, and the industry probably needs something closer to $1T in annual revenue to justify the investment.
If customers are going to spend $1T on AI, they need more than $1T of value back. Total U.S. labor spend is roughly $20T, so the math starts pointing toward something like 7-8% labor replacement, or equivalent productivity gains, just to make the buildout pencil.
To be clear, I don’t think AI is a bubble. I think the opposite. But the AI boom can be real and still have a very high economic bar to clear. And that bar won’t be cleared only by making developers faster.
Report
Finance needs to prove AI’s return: CloudZero report
260 senior finance leaders (more than half CFOs) told us why the speed of seeing AI spend, not the size of it, separates who pulls ahead on AI from who gets burned.
Developers matter, but they’re not the whole story
Software developers are roughly 1% of the U.S. workforce. Even if you widen the lens to include security, sysadmin, IT, and adjacent technical roles, you still don’t get close to the part of the labor market that has to change for the AI math to work.
That’s why the non-developer side matters so much.
JPMorgan has said 81% of its workforce is non-technical. Microsoft has said 65%. OpenAI has said that non-developers, including analysts, marketers, operators, designers, researchers, investors, and bankers, make up about 20% of overall Codex users and are growing more than 3x as fast as developers.
The first wave of AI ROI showed up in engineering because AI came from technologists and was first adopted heavily by software companies. Companies measured it with the workflows they already knew. Code generation was obvious. PRs went up. Tickets moved faster. DORA metrics gave everyone a familiar scoreboard.
But the second, much larger wave, will hit when AI changes how businesses actually run: sales, finance, marketing, operations, legal, customer success, and recruiting.
Riding that wave doesn’t mean good tooling. It means radically updated company design.
Sales already has a scoreboard
Last week, a colleague asked me how we should measure non-developer AI ROI. His assumption was that it would be harder.
My reaction was the opposite: sales has never had the luxury of fuzzy measurement. The ROI is in the comp plan.
Engineering had to elevate proxy metrics: deployment frequency, change failure rate, lead time, time to restore service. Those are useful metrics for running an engineering team, but a team can improve all of them and still build the wrong product. They can ship faster and still miss the revenue plan.
Sales can’t. GTM has quota. We have activity metrics too, but those have never been the metrics that matter most in board meetings and investor decks. The test is cleaner: Does AI create better pipeline, improve conversion, shorten cycles, increase deal sizes, and ultimately drive more revenue per rep?
Improved sales velocity is ROI. Anything else is activity.
What this looks like in practice
In my org, there’s one crowning ROI metric: sales velocity per AE.
More qualified opportunities. Higher conversion. Larger deals. Faster cycles. Better management visibility. If AI moves those numbers, the next headcount plan looks different from the last one.
So I look for outliers: Where does AI spend correlate with better conversion? Where is a rep getting unusual leverage? Where is a manager seeing risk earlier? Where did a workflow move from interesting to repeatable?
When we see that pattern, we inspect the workflow.
Who built it? What does it cost per run? Does the quality hold? Can the rest of the team use it, or is it trapped inside one power user’s process?
If the answers check out, we scale it.
One of my reps built a demo-prep skill that costs $5.21 per run. He went 5-for-5 on meetings prepped with it. Five meetings is a signal, not a sample size, but the discipline matters more than the anecdote: Attribute the cost, measure the outcome, and decide whether to scale.
I have an AI system that briefs me every morning on what moved, what’s at risk, and what needs my attention. The information finds me instead of the other way around. That’s not AI usage. That’s management leverage.
One of my reps spent $1,100 the night a new model dropped, running experimental loops. Same rep who automated huge parts of our renewal motion for pennies a night.
Nobody panicked, because the AI spend was tagged by purpose. We could have a conversation instead of a lockdown.
That distinction matters. Hard caps don’t protect you from waste. They mostly kill the workflows that are starting to work. The better answer is to scale what works, harden what shows promise, and trim what burns spend without moving a number.
Allocate capital based on outcomes, not activity
Most companies are going to measure AI the way they measure SaaS adoption or engineering output. Seats provisioned. Active users. Chat volume. PR velocity. Token spend.
Those aren’t useless, they’re just not ROI.
A CFO shouldn’t be satisfied with, “500 people are using AI.” A board shouldn’t be impressed by a usage chart if nobody can connect that usage to revenue productivity, margin, labor leverage, customer outcomes, or speed of execution.
The question shouldn’t be, “Are our people using AI?”
Or, “Is AI boosting activity?”
It should be, “Where is AI changing the metrics that move the business?”
Revenue. Margin. Labor leverage. Customer outcomes. The things we already put in OKRs.
AI spend isn’t just a tool budget anymore. It’s not even only a technology initiative.
It’s becoming a capital allocation problem.
The winners will know where AI is replacing work, where it’s creating leverage, where it’s just creating expensive noise, and how much intelligence they’re actually getting per dollar.
We don’t need a new scoreboard. We need the discipline to connect AI spend to the one the business already gets paid on.
OKRs, not DORA.