AI costs are changing. As noted by research from EY, outputs that cost just $0.04 in 2023 now cost $1.20, a 30x increase over just three years.
It’s worth noting that task operations and complexity have also changed. In 2023, the process was simple. Users input a question, retrieval engines found relevant data, and AI models returned a response. Today, many tasks are handled by orchestrated AI agents capable of much more complex reasoning and analysis. Models can also retry and refine questions to improve outputs and reduce the time between questions and answers.
The result is a challenging market for enterprise AI budgeting. Spend too little, and outputs won’t keep pace with staff and customer queries. Spend too much, and AI return on investment (ROI) becomes almost impossible to achieve.
While model choice and business use case impact cost over time, the biggest contributor to rising budgets is AI tokens. Here’s what businesses need to know about the token management that can help control AI agent and LLM spend.
Let’s talk tokens
NVIDIA defines AI tokens as “tiny units of data that come from breaking down bigger chunks of information.” But what does this mean, exactly?
Think of a token as the smallest unit of data processed by AI tools. All inputs and outputs are made up of tokens, and each of these tokens comes with a cost assigned by the LLM provider.
While there’s no hard-and-fast rule for the size of tokens, AI providers typically use words as benchmarks, with one token representing four letters or numbers. For example, the words “token budgets” contain 13 characters, including the space. If every four characters equals one token, the two words represent three tokens.
Tokens make up both parts of the AI process. When a user asks a question, all the words processed by the model are part of tokens. All output words are also tokenized. Consider the following question and answer.
Q: What is a token?
A: A token is the smallest unit of data processed by AI tools.
The question has 16 characters, or four tokens. The answer has 59 characters, or approximately 15 tokens. While the question was significantly shorter than the answer, businesses are charged for both, meaning the total number of tokens used is 19.
A token budget is the amount allotted for token spending each month. It should be based on intended usage, but also account for potential cost overruns. Token monitoring solutions are essential to notify teams if they are approaching budget limits, and budgets should be regularly reviewed to ensure they account for necessary token volumes.
Report
Finance needs to prove AI’s return: CloudZero report
260 senior finance leaders (more than half CFOs) told us why the speed of seeing AI spend, not the size of it, separates who pulls ahead on AI from who gets burned.
Interface vs. API AI options
While all AI models use tokens, companies aren’t always charged per token. This is because there are two broad types of AI available: interface-based and API-based.
Interface-based options are typically free for simple queries and available as per-seat plans for enterprise use. For example, the Claude Team plan includes single sign-on (SSO), admin controls, and shared projects, at $25 per seat per month on monthly billing or $20 on annual billing, with a five-seat minimum.
They provide access to web or application-based AI interfaces that do not directly connect with enterprise applications or services. In effect, this type of AI is more surface-level. It’s great for questions and answers, but companies can’t use it to automate or orchestrate key functions.
API-based plans, meanwhile, are charged on a per-token basis. Choosing an API plan lets teams deploy and connect AI with the services and solutions of their choice. For example, companies might connect AI tools with ERP systems to track supply chain trends and improve procurement efficiency.
Claude offers multiple API models, each with its own cost. The company’s fastest and most efficient model, Haiku 4.5, costs $1 per million input tokens and $5 per million output tokens. Its next-generation intelligence model, Fable 5, comes in at $10 per million input tokens and $50 per million output tokens.
Common challenges in token management
Any business using an API-driven AI approach needs a complete accounting of token use. Several common challenges, however, can undermine this effort. They include:
Unsupervised services
If AI-based services are left to run unsupervised, they may constantly run queries and produce answers. Both of these processes are token-intensive, meaning companies may get a shock on their next AI bill.
Inefficient processes
Queries that must be run multiple times to return correct or usable answers aren’t free. Every subsequent query comes with another set of input and output costs. The more inefficient the process, the more companies pay.
Unprepared data
If data is not properly vetted, deduplicated, and compressed, AI models may take additional time and require additional tokens to parse this information and identify trends. While pre-AI preparation is costly, it’s significantly less expensive than using tokens to do the job.
Tips for managing token budgets
Tokens cost money. More tokens = more money, and ideally, a greater return on investment. Given the sheer number of tokens used by even small and midsize businesses, however, companies need concrete strategies to keep token budgets under control.
1. Assign tokens to the right cost line
As noted by the Boston Consulting Group (BCG), token costs should not be part of the IT budget, nor do they all belong on the same line.
Instead, BCG recommends that tokens used to build reusable capacity, such as designing agents or improving workflows, should be considered investments similar to CapEx. Tokens used for internal functions should be classified as operating expenses, while those that are part of customer-facing services are reported as costs of goods sold (COGS).
There’s a catch that makes this harder than it sounds. A provider invoice arrives as one number. Splitting it across three P&L lines means knowing which tokens built an agent, which ran an internal workflow, and which served a paying customer, and no billing dashboard tracks that distinction. That allocation is what CloudZero does: mapping token spend to the feature, team, and customer behind it, so the CapEx, OpEx, and COGS split becomes a report rather than an estimate.
BCG ties all three lines back to one measure: return on AI, or economic return divided by the combined cost of human intelligence and tokens. Whichever line the spend sits on, that ratio is what tells you whether it earned its place there.
2. Consider cost in the context of performance
Lower per-token costs aren’t always cheaper. Consider an AI model that charges $1 per million tokens but only resolves 20% of queries on the first try, compared to one that charges $2.40 per million tokens but has an 80% success rate.
Spending $1 on the cheaper model leaves 4/5 of queries unanswered. To reach 100% coverage, businesses must spend another $4, for a total of $5 on five million tokens.
In the case of the more expensive model, meanwhile, $2.40 gets companies 4/5 of the way there, leaving just 1/5 of queries to resolve. Another 48 cents covers the remainder, meaning companies spend just $2.88 on 1.2 million tokens to get the job done.
3. Capture granular token use
To effectively set and manage token budgets, businesses need to know exactly how many tokens they use. Not a ballpark, not an approximation, and not an educated guess. Exactly, down to the last token. In practice, this requires AI cost management solutions capable of breaking down input and output token use by model version and caching, and reporting this data in real time.
AI budgets: more than just a token gesture
AI models matter, but tokens are the biggest budget line for companies using API-based solutions. With token costs changing as AI options evolve, businesses need to know exactly how many tokens they use, how much they spend, and how close (or far) this number is to ideal token budgets.
Control AI and LLM spend with granular token monitoring. Schedule a CloudZero demo today.