Why this matters
If you set the budget for your team’s AI agent work, or answer to someone who does, you need a rough idea of what a job will cost before it starts. That’s hard to get. Stanford researchers found the same agent, given the same task, can use up to 30 times more tokens from one run to the next, and you usually find out afterward.
Most developers just run the job. The people who think about cost are the ones asked about the bill: the engineering lead who sets spend limits on agent runs, and the finance partner who forecasts them. Gartner predicts AI coding costs will pass the average developer’s salary by 2028, so that group is growing.
If you run an agent now and then, you don’t need any of this. But the lesson applies to anyone who uses AI. A model’s guess about its own work isn’t a number to plan around, and as agent work grows from a few runs a week into a line on the budget, someone on your team will need one that is.
For them, an estimate is only useful if it holds still. If you get a different answer every time you ask, you can’t plan around it, and a spend warning built on it goes off at random.
The obvious move is to ask the AI how many tokens a job will take. That’s the one thing it’s worst at.
What we built
Here’s what we found while building CloudZero’s Model Rightsizer: an AI model can’t tell you what its own work will cost. Ask the same model to estimate the same job three times and you’ll get three different answers, often 10 to 30% apart. Stanford’s researchers found the same thing, and added that models tend to guess low.
Model Rightsizer, our free, open-source tool for Claude Code, is now built around that finding. Describe the work you’re about to do, and it tells you which model should handle each piece, with a token budget next to each one. New in this release: the model no longer guesses that budget. It rates each piece on three things it can judge well: how many tool calls it needs, how much new content it writes, and how many other files it has to stay consistent with. Fixed math turns those ratings into a budget, so the same work gets the same budget every time.
Work that’s hard in all three ways at once gets a budget that reflects all three. That’s the work most likely to run long. On real tasks the team hadn’t tuned it on, the budget called 3 of 4 correctly, and the miss ran over by about 4%.
Both directions matter. If you set a budget too low, the agent gets cut off mid-job, so you pay to run it again. If you set it too high, you approve spend you didn’t need.
How it works
Install it free, or update the plugin if you already have it:
/plugin marketplace add cloudzero/cloudzero-claude-marketplace
/plugin install model-right-sizer@cloudzero
Before your next agent job that’s bigger than a quick fix, ask Model Rightsizer for the plan first. Think of it as getting a contractor’s quote before anyone knocks out a wall.
Say an agent is about to move a service to a new API, and the change touches a dozen files. Describe the job and you get it back in pieces, each with a model and a budget. The one-file edits get small models and small budgets. The piece that has to keep all twelve files consistent gets the biggest budget. You know where the money goes before you spend it, so you can split that piece, give it more room, or do it yourself.
You don’t need to own a budget to benefit. You stop losing afternoons to reruns, and you can compare two ways of splitting a job before you pick one. If you do own the budget, your spend limits rest on a number that holds still, and you can show your lead or finance partner a plan with a number next to each piece.
One caveat: Sonnet budgets are the most thoroughly measured. Opus and Haiku budgets are scaled from Sonnet for now, and the team is measuring them directly next.
See it in the docs → https://docs.cloudzero.com/docs/ai-model-right-sizer