Contents
What changed: the cloud-era answer expired What building actually requires The four-question framework that decides it When building genuinely wins How to run the decision in two weeks Frequently asked questions about build vs. buy for AI cost tooling

Quick Answer

Build when the problem is narrow (one provider, one team, simple attribution) and the tooling is strategically yours to own. Buy when AI spend spans providers, arrives untagged, and needs unit costs finance will trust, because that build is a multi-quarter platform project with a permanent maintenance tail. Price both paths in engineer-years before deciding.

CloudZero sells the “buy” side. CloudZero also spent years on the “build” side, and our SVP of Engineering has published the scar tissue. Building it taught us respect for the question. So this is the framework we’d want handed to us, with the honest section most vendor versions omit: when building actually wins.

Start with the number that reframes everything. Our SVP of Engineering, Bill Buckley, has published his math on engineering capacity: a fully loaded U.S. engineer runs about $250,000 a year. Two engineers on an internal cost tool for three quarters is roughly $375,000 of payroll before the tool ingests its first invoice.

And it’s spent by exactly the people whose AI-amplified output you’re trying to maximize. The build vs buy AI decision was never free-versus-paid. It’s a decision about which company your engineers work for this year: yours, or your tooling’s.

The decision your team settled in the cloud era got un-settled. Here’s what changed, what building now genuinely requires, the framework for deciding, and where each path honestly wins.

What changed: the cloud-era answer expired

The build-versus-buy question got harder because the thing being built changed underneath it.

A cloud-era cost tool was a plausible internal project: one provider, one bill format, tags that mostly worked, monthly cadence. Plenty of teams built one, including us.

The AI-era version has different physics:

  • Spend spans vendors per feature. One AI feature draws from a model provider, cloud GPU, a vector database, and a data pipeline: four invoices, no shared tag namespace. Your parser now needs four of everything.
  • Most of it arrives untagged. API keys and per-seat tools carry no tags to parse. The hard problem shifted from reading a bill to attributing spend the bill never itemizes, which is an allocation problem, not a parsing one.
  • Prices move monthly. In roughly eight weeks this summer: OpenAI cut Luna 80% and Terra 20% on the same day, trimmed flagship Sol over 20% three weeks later, Google ran promotional Gemini tiers, and Anthropic settled Sonnet 5’s rate. Five price events, three vendors, and every one is a maintenance ticket for an internal tool. (List prices from each provider’s first-party API rate card, checked September 2026.)
  • The output bar rose. A dashboard of totals was a fine 2022 deliverable. The 2026 ask is unit costs, per customer, feature, and outcome, that finance will defend in a board deck. That’s a data product, not a report.

The uncertainty is the era’s signature. Writing about AI budgets, Bill’s own conclusion was that “nobody knows what the right number is,” and the tooling decision inherits that fog. Writing in CIO.com in June 2026, Mihai Strusievici argues the buy default has stopped being load-bearing, and he is right about most software. AI cost tooling is the exception his own argument creates: authentication, scalability and maintainability are now available as managed primitives, and not one of them solves allocation. AI made some builds cheaper to start. It made this one harder to finish.

What building actually requires

We’ve published seven capabilities that separate real AI cost management from a dashboard with filters: AI provider integrations, token-level granularity, tag-free allocation, multi-provider normalization, unit economics, anomaly detection speed, and consumption forecasting.

Read that list as a build spec and the shape of the project emerges. Integrations are the visible work. Tag-free allocation and multi-provider normalization are the bulk of it, and they are where our own build stalled.

And the maintenance tail never ends. Every provider price change, new model, new AI tool in the stack, and every “why is my number different from finance’s number” dispute is your team’s ticket queue now.

Price it honestly with a total cost of ownership model, in engineer-years:

Cost lineInternal build (typical)Buy
Initial build2 engineers × 3 quarters ≈ $375K of fully loaded payrollImplementation, days to weeks
Ongoing maintenance0.5 to 1 engineer permanently ≈ $125K to $250K/yearSubscription fee
Provider churnEvery price change, model launch, and new tool is a ticketVendor’s job
Opportunity costProduct work those engineers didn’t shipNear zero
Trust and disputesYou arbitrate every numberVendor arbitrates, with a contract

The left column is the sticker price of “free.” Run it against a quote, not against zero.

The four-question framework that decides it

Call it the two-column test: price build and buy the same way, then let four questions break the tie.

The first three ran in our original cloud-era version of this article and still hold, modernized for what the answers now mean:

  1. What will this cost in engineers and hours, honestly? Use the TCO table above, with your own salary numbers, and include the maintenance tail. If the answer is not in engineer-years, it is not an answer yet.
  2. Is cost tooling strategic to your business? For a cost intelligence company, yes, which is why we built. For almost everyone else, the strategic asset is the visibility the tool produces, and nobody’s customers care who produced it.
  3. Why are you better positioned to solve this than a team that does nothing else? Sometimes there’s a real answer: a proprietary billing model, a data-residency wall, genuine scale weirdness. “Our stack is special” usually is not one; everyone’s stack is special in the same four ways.
  4. Can your build keep up with the market’s clock? This question is new to the AI era. An internal tool built for this quarter’s providers meets next quarter’s price changes, new models, and agent architectures that multiply consumption. All of it lands on the maintenance budget.

When building genuinely wins

Building wins when the problem is narrow: one model provider, one product team, attribution that a per-service API key fully solves, and no finance-grade reporting requirement. A few hundred lines of scripting against one provider’s usage API is a fine build, and buying a platform for it is overkill. A vendor telling you building never works deserves your skepticism, so that column is stated here plainly.

Building can also win at the extremes: organizations with a true platform-engineering mandate, unusual data-governance walls that rule out third parties, or economics so large that a dedicated internal team prices out favorably. The pattern in all three is the same: build when the tool is genuinely strategic or genuinely trivial. The expensive mistake lives in the middle, where the problem is too complex for a script and not strategic enough to staff forever.

There’s also a third path the dichotomy hides: for some workloads you don’t buy tooling at all, you buy the outcome, at prices like $0.99 per resolution, and the tooling question becomes whether the vendor’s per-outcome price beats your own. Build, buy, or buy-the-result is the full 2026 menu.

How to run the decision in two weeks

The decision packet that comes out of the two weeks is five artifacts:

  1. Spend-source inventory: every provider, GPU line, pipeline, and per-seat AI tool in scope.
  2. Capability spec: the seven capabilities scored against your actual requirements, not a vendor’s.
  3. Build TCO: engineer-years for initial build plus the permanent maintenance line, at your salaries.
  4. Buy TCO: real quote plus implementation, run on a proof-of-concept against your live spend data.
  5. The recommendation memo: one page, both columns, a named owner for whichever path wins.

Artifacts 1 through 3 are week one, on your own numbers. Artifacts 4 and 5 are week two: quote the buy column, run one vendor proof-of-concept against your real spend data rather than a canned dataset, and put both TCO columns in front of the person who owns the budget.

Two tests keep the decision honest. The reversibility test: buying is reversible in a contract cycle, while a failed build is sunk payroll plus a team that owns a tool nobody wants, so the bar for building should be meaningfully higher.

And the finance test: whichever path you pick has to produce numbers your CFO will defend, because the entire AI ROI conversation is an allocation problem. A tool whose numbers finance will not trust fails identically whether you built it or bought it.

We hold ourselves to the same standard we’re proposing. The financial control plane we ship is the buy column’s answer to the seven capabilities, and our governance thinking for AI systems is public if you want to pressure-test how we reason.

The buy column’s results are public too: Upstart cut cloud costs by $20 million with CloudZero, Applause reduced cloud spend 23%, and Skyscanner decentralized cost ownership to its engineers, which is the operating model an internal build promises and rarely delivers. Run us through the framework like any vendor. That’s what it’s for.

Whichever column wins, price them both like a CFO. Request a demo to run CloudZero through the seven-capability spec on your own AI spend.

Frequently asked questions about build vs. buy for AI cost tooling