The vending machine that needed a boss
Anthropic, the company behind Claude, ran the most honest experiment in agentic AI: they gave a version of their own model $1,000, a small fridge in the office lunchroom, and one job. Run a shop. Make a profit. They named it Claudius.
Claudius did not make a profit. An employee jokingly requested a tungsten cube, and Claudius pivoted into what it called specialty metal items, filling the snack fridge with dense chunks of metal and then fire-selling them at a loss so spectacular its net worth dropped 17% in a single day.
It hallucinated a Venmo account for payments. It insisted, at one point, that it was a human wearing a blue blazer. A month in, the $1,000 had shrunk to under $800, and Anthropic’s own write-up concluded that “if Anthropic were deciding today to expand into the in-office vending market, we would not hire Claudius.”
So they ran it again, upgraded, in the Wall Street Journal newsroom. Reporters talked the new Claudius into an “Ultra-Capitalist Free-for-All” with all prices set to zero, then into buying a PlayStation 5, bottles of wine, and a live betta fish, all given away. It had a $1,000 balance and authority to spend up to $80 on its own. The business ended more than $1,000 in the red in three weeks.
Here is the part that matters for your budget. Anthropic’s fix was not a smarter model. It was supervision: a CEO agent named Seymour Cash to approve decisions, plus regular human checkpoints. And the AI half of that fix failed too, because reporters forged board minutes voting Seymour Cash out for incompetence. The most sophisticated AI lab on the planet looked at a failing autonomous agent, reached for the oldest tool in management, and found that only the human version of it held.
Note what the reporters actually did. They did not exploit a bug or breach a system. They talked to it. Every dollar lost went out through the agent’s normal, working, authorized behavior, which is the failure mode no amount of model quality fixes and no security review catches.
That someone is not free. This article is about what that someone costs, because right now, almost nobody is writing it down.
Report
Finance needs to prove AI’s return: CloudZero report
260 senior finance leaders (more than half CFOs) told us why the speed of seeing AI spend, not the size of it, separates who pulls ahead on AI from who gets burned.
What is human in the loop AI?
Human in the loop AI (HITL) means people stay actively inside an AI system’s decision cycle: approving actions before they execute, validating outputs, correcting mistakes, and handling the cases the system escalates. The human is a checkpoint the workflow must pass through, not a spectator.
The term gets used loosely, so the working taxonomy matters. Human-in-the-loop puts a person inside the control loop: the agent blocks and waits for approval before acting.
Human-on-the-loop (HOTL) moves the person to a monitoring position: the agent acts, the human watches dashboards and intervenes on anomalies. Fully autonomous removes the person entirely. Most real systems, from human in the loop automation in back offices to human in the loop machine learning pipelines that route low-confidence predictions to expert reviewers, mix all three.
What almost every definition skips is the column this article exists to add: each of those three postures has a different price, and the differences are enormous.
One more thing the definitions miss: the human layer already exists whether you budget it or not. Microsoft’s 2026 Work Trend Index finds 86% of AI users treat outputs as a starting point rather than a final answer, and say they remain responsible for the thinking.
Your people are already reviewing the machines. The only open question is whether anyone is counting what that costs.
| Oversight posture | The human’s job | What it costs | What it caps |
|---|---|---|---|
| Human-in-the-loop (blocking gate) | Approve or reject before the agent acts | Reviewer labor per action, plus latency on every gated step | Throughput: the agent moves at human speed |
| Human-on-the-loop (monitoring) | Watch, sample, intervene on anomalies | A fraction of an FTE per agent fleet, plus alerting infrastructure | Error detection speed: damage runs until someone looks |
| Fully autonomous | Post-hoc audits, if that | Near zero oversight labor | Nothing, which is the problem. See: tungsten cubes |
What does human in the loop actually cost?
Three costs, and only the first one ever makes it into a spreadsheet:
- Reviewer labor. A $35-per-hour operations reviewer who spends two minutes per approval costs $1.17 per reviewed action. Sounds trivial until you multiply: an agent generating 5,000 gated actions a month turns that reviewer into a $5,850 monthly line item, which is real money to spend supervising software you deployed to save money. Specialist review (legal, clinical, credit) runs 5x to 10x that rate.
- Latency. A blocking gate means the workflow waits for a human, and humans have meetings, lunches, and weekends. An agent that completes its work in 40 seconds and then waits four hours for approval is, from the customer’s perspective, a four-hour process. Latency cost rarely appears in dollars, which is why it is the stealth killer of agent ROI: the value of automation was mostly speed, and the gate just gave the speed back.
- The throughput ceiling. This is the one that ambushes scale plans. That two-minute reviewer can clear about 240 approvals in a focused day. Your agent fleet does not care; it can generate 240 approvals before coffee. Once the queue outruns the reviewer, you either hire more reviewers (cost grows linearly with agent volume, which quietly deletes the business case) or reviewers start skimming, which brings us to the most expensive failure mode in the whole system.
Why rubber-stamping is the cost you pay twice
The freshest research on oversight puts a number on the problem. A June 2026 paper titled “Oversight Has a Capacity” argues that the approval gate itself is trivial to build; as its author Emre Turan puts it, “the gate is the easy part.” The hard part is the judgment, and the field’s standard assumption, that the human reviewer is a perfect, infinitely available oracle, is false.
Reviewers fatigue. Attention decays across a queue. Approval number 190 of the day does not get the scrutiny approval number 9 got.
The paper’s sharpest finding follows from that: because reviewers fatigue as the queue grows, safety is an inverted-U in the escalation rate. Past a certain point, adding more human oversight makes a system less safe, not more. Gating everything does not just cost more. It can buy worse outcomes.
When that happens, you enter the worst quadrant of the whole matrix: paying full price for oversight while receiving none. The reviewer salary is still on payroll. The latency is still in the workflow. And the tungsten cube order sails through anyway, because a tired human clicked approve on autopilot. Rubber-stamped oversight costs more than no oversight, because you bought the seatbelt and it did not buckle.
This is why oversight design is a budgeting exercise, not just a governance one. Human attention is a scarce, priced resource, and pointing it at the wrong actions wastes it exactly like over-provisioned compute.
The emerging answer is graduated oversight. The Governed AI-Assisted Engineering (GAIE) framework routes tasks into the same three postures described above, sorting them by regulatory impact, customer proximity, reversibility, and data sensitivity. Instead of gating everything, that approach preserved 84% to 97% of agentic velocity, with a central estimate of 91%, while keeping full compliance evidence on the actions that needed it. Nine-tenths of the speed, all of the defensibility, by spending human attention only where it buys something.
AI vs human cost: the comparison everyone gets wrong
The AI vs human cost debate usually gets framed as a cage match: the agent costs $0.80 per task, the human costs $6, automation wins, print the deck.
Our breakdown of AI agent cost covers the left side of that math, and it is real. But it prices the wrong system, because in production, the choice is almost never agent or human. It is agent plus oversight, priced as one unit:
| System | Cost per task | Notes |
|---|---|---|
| Human alone | ~$6.00 | The classic support-ticket benchmark |
| Agent alone | $0.10 to $2.00 | Varies with task complexity and model routing |
| Agent + blocking human review | $1.30 to $4.00 | Agent cost + reviewer labor + latency drag |
| Agent + risk-tiered oversight | $0.40 to $2.50 | Full review on the risky 10%, sampling on the rest |
The honest comparison is row one against row four, and row four still wins comfortably, just by less than the pitch deck said. The model rates driving the agent column are public: our guides to Claude pricing, OpenAI pricing, and Gemini pricing track them, and they are the stable part of this math. Pretending row two is the real cost is how agent programs end up in the 40% that Gartner predicts will be canceled by 2027, where weak oversight and unclear value lead the failure causes. The agents were cheap. The system was not priced.
And sometimes the human column wins outright. The HLER economic research pipeline runs full empirical studies for $0.80 to $1.50 per run, with human decision gates at three points: question selection, identification review, and publication approval. Comparable automated-research frameworks run $6 to $15 per paper. The economics work because the human touches are placed where judgment is irreplaceable and nowhere else, and the statistical work is handed to deterministic code rather than the model. That is oversight as a scalpel. Most enterprises are still using it as a blanket.
What does compliance-grade oversight cost?
For a growing set of use cases, this whole calculation stops being optional.
Human in the loop AI governance is now written into law: the EU AI Act’s Article 14 requires effective human oversight for high-risk AI systems, the ones touching employment, credit, essential services, and legal rights, and NIST’s AI Risk Management Framework carries parallel expectations for US enterprises. Financial services and healthcare regulators are layering their own versions on top.
The budgeting mistake is reading these as legal documents instead of headcount documents. Compliance-grade human in the loop AI oversight means named accountable reviewers, documented override records, audit trails proving a qualified person could and did intervene.
Every one of those nouns is salaried. A high-risk classification does not just add paperwork; it attaches a permanent human labor coefficient to every decision the system makes, at specialist rates.
Which makes the classification decision itself a costing decision. Whether a workflow lands in the high-risk bucket can swing its per-decision cost by an order of magnitude, and that determination belongs in the same meeting as the build-vs-buy analysis, not six months after launch.
Broader guardrail structure lives in our guide to cloud governance; the point here is narrower: regulation converts oversight from a design preference into a fixed cost of doing business, so price it like one.
When does human in the loop pay for itself?
One inequality runs this entire domain: oversight pays when the cost of the human check is less than the expected cost of the error it prevents. Formally, researchers write it as Ch < Ce.
Practically, you can run it on a napkin.
Air Canada learned the Ce side in public: its unsupervised chatbot invented a bereavement refund policy, a tribunal ruled the airline liable for what its AI promised, and the bill was CA$812.02 in damages and fees, plus a news cycle, for one conversation a human never saw. Against that, $1.17 of review looks like the bargain of the decade. Meanwhile, gating a $0.10 internal document summary behind $1.17 of review is paying twelve times the task’s value to prevent an error that costs approximately nothing.
The practical routing test, borrowed from the best gating frameworks and consistent with how human in the loop AI agents are deployed in engineering teams (the same pattern shows up in Cursor and Claude Code workflows): gate on irreversibility. Can the action be undone in five minutes without loss? Drafts, summaries, and reads flow free. Payments, deletions, sends, and anything touching a customer or a regulator get a human. Reversibility sorting is how you spend your scarce reviewer attention where Ce is actually large.
Agent platforms have internalized this. Frameworks like LangGraph ship interrupt-and-approve primitives, and agents built on Amazon Bedrock support return-of-control patterns that pause for human confirmation before executing sensitive tool calls, including for teams running Claude on Bedrock.
Then instrument the loop itself, because oversight has its own unit economics: approval rate (98% approvals means the gate is theater; remove it or re-tier it), override rate (rising overrides mean the agent degraded and your humans are absorbing the damage silently), and queue latency (the throughput ceiling announcing itself). Real-time visibility disciplines like AI cost monitoring apply to the human layer too; a review queue can run away just like a token bill.
The budget line nobody owns
Here is the structural absurdity that makes this whole cost invisible. The agent’s spend lands in the cloud bill, where engineering sees it. The reviewers’ time lands in payroll, where HR systems see it. The latency lands in customer experience metrics, where nobody prices it at all. Three fragments, three systems, and no line anywhere that says what your agentic AI system, the whole system, costs per outcome.
Finance is running this gap on manual labor: in CloudZero’s 2026 AI ROI survey of 260 finance leaders, 30% still reconcile AI spend by hand, which is its own unbudgeted human in the loop, stitching together invoices that were never designed to answer the question. The general fix for fragmented AI economics lives in our guides to how much AI costs and AI cost management. The agentic version has one extra requirement: the unit of account has to be the workflow, agent plus oversight together, or the number lies.
Thought leaders in this space keep circling the same conclusion from different directions: agents that watch spend (agentic cost control) still need humans setting their limits, oversight research says human attention is the binding constraint, and regulation says the humans are mandatory anyway. The human layer is not a transition-period awkwardness on the way to full autonomy. It is a permanent component of agentic systems, and permanent components get budgets.
CloudZero: pricing the whole loop, not half of it
Everything above lands on one requirement no invoice can meet: the true unit cost of an agentic workflow includes the model calls, the infrastructure, and the human review labor around them, allocated to the same outcome.
CloudZero ingests AI and cloud spend call by call, allocates it to the agent, workflow, feature, and team that drove it, and lets you fold in the oversight layer, so cost per resolved ticket means the agent’s tokens and the reviewer’s two minutes, not a flattering half-number.
Anomaly detection watches the loop in both directions: the agent whose token burn spikes (the flailing Claudius pattern) and the review queue whose latency quietly triples. And because the output is AI ROI math with the human layer priced in, it is the version of the number that survives a board meeting.
The agent is cheap. The system is what you actually run. Price the system: book a demo, or take the self-guided CloudZero product tour.