Your AI Agent Can Spend Real Money While You Sleep: The Case for Hard Budget Caps
Simon Willison's argument for default hard budget caps on metered services topped Hacker News this week. Why AI coding agents create a new cost risk, how Uber reportedly burned its AI budget in four months, and five guardrails to build in from day one.
Amir Ali Liaqat · Founder & CEO, DesignsToDeploy

The story in one paragraph
On October 3, 2026, developer Simon Willison published a post titled "We're going to need default hard budget caps on pretty much everything." Within a day it was near the top of Hacker News with hundreds of points and hundreds of comments, one of the most-discussed engineering threads of the week. His argument is simple and worth taking seriously: coding agents have made it trivially easy to deploy software, and just as easy for a runaway loop to burn real money overnight while nobody is watching. Every metered service those agents can reach, he argues, should ship with a hard spending cap turned on by default.
Why AI agents change the cost equation
Cloud bills have surprised people since cloud computing existed. What agents change is the distance between intent and spend. A human developer provisioning infrastructure pauses, at least briefly, to think about what it costs. An agent running at 2 AM does not pause at all. It calls paid APIs, writes to object storage, spins up auto-scaling compute, and keeps going until the task is done, or until something stops it. The session ends. The meter keeps running.
The failure mode Willison describes is not a crash or a bug. It is success without supervision: the demo the agent shipped last night kept running while you slept, and the warning email that arrived at 2 AM changed nothing. Soft budgets notify. They do not shut the meter off.
Soft caps notify. Hard caps stop.
The distinction is the whole argument. A soft cap says: after $X, email me. It is a notification, not a control. An unattended application can keep consuming paid services after the alert arrives, and you discover the cost in the morning.
A hard cap says: after $X, cut it off and return errors. The service pauses. Nobody likes downtime, and Willison acknowledges that providers resist hard caps because returning errors can break a customer's app. His counter is blunt: most people would take a paused service over a surprise bill above $10,000. And the cap is the default. Customers who understand the risk can explicitly opt out and accept overages.
“A warning email at 2 AM does not stop the meter. A hard cap does.”
Proof it is real: the Uber story
If this sounds like a corner case, consider what happened at Uber. According to reporting by The Information, confirmed by Uber CTO Praveen Neppalli Naga, Uber rolled out Claude Code in December 2025. By February, 32% of engineers were on agentic coding tools. By March, 84%. Somewhere in that sprint, the company burned through its entire 2026 AI budget in four months. Bloomberg reported the response in June: a hard cap of $1,500 per employee per month, per tool. (Via CIO, September 2026.)
The most interesting part of the story is not the overrun. It is what Uber's COO said when asked whether the spending was working: it was very hard to draw a line between the usage statistics and actual productive output, like producing 25% more useful consumer features. That is a measurement complaint, not a cost complaint. Uber did not cap spending because tokens are expensive. It capped spending because it could not price what the tokens were buying. You cannot defend a number you cannot connect to anything.
The clouds are moving
The industry is responding. Google Cloud introduced Spend Caps in July 2026. AWS announced optional monthly project spending limits in September 2026: reach the selected limit and the project pauses, with separate budgets allowed for separate projects. The controls are still young, some are in limited rollout, and AWS's version carries operational qualifications worth reading before you rely on it. But the direction is clear. Hard limits are becoming a product feature, not a workaround.
Five guardrails for anything agents touch
If you are building or buying software that AI agents operate, here is the checklist we recommend:
- Default-on hard spending caps on every metered service. The cap pauses, it does not email. Opt out only if you accept the risk.
- Separate monthly budgets per project. Demos, experiments, and production each carry their own limit, so a runaway experiment cannot eat the production budget.
- Alert on spend rate, not just totals. A $2-per-minute spike caught in real time is worth more than a monthly total reviewed after the damage.
- Idempotency keys and retry budgets. Every money-touching call must include an idempotency key so retries never double-charge, and retries get a per-run budget.
- Human approval gates on real-money actions. Deploy, buy, scale, or charge: nothing fires without a person signing off.
What this means for your product
At DesignsToDeploy, we build these guardrails into the products and infrastructure we ship for clients, because AI speed should never mean surprise bills. Cloud setup, CI/CD, cost alerts, and budget caps are part of the same discipline as testing and code review. If you are building with AI agents, or planning to, build the cost safety in from day one. It is far cheaper than the incident that teaches it.
Have a project in mind? Reach us at designstodeploy@gmail.com, call or WhatsApp +92 309 0886518, or start at www.designstodeploy.dev.
Sources: Simon Willison's blog, "We're going to need default hard budget caps on pretty much everything" (Oct 3, 2026), via entrelligence.com, ycroaster.com, redreamality.com, and muddy.jprs.me summaries and Hacker News discussion; CIO, "The wrong million tokens" (Sep 25, 2026), reporting The Information and Bloomberg on Uber's AI budget; AWS September 2026 builder-experience announcement; Google Cloud Spend Caps (July 2026).
1 / 4Keep reading

Top 5 Web Hosting Providers for 2026
The right hosting makes your website faster, safer, and more reliable. Here are five popular providers worth comparing in 2026 — plus what to look for before you buy.
Read article
AI Agents Now Write Half the Code: What the JetBrains 2026 Survey Means for Your Business
JetBrains surveyed 15,000+ professional developers and found 47% of code is now fully generated by AI agents. Here is what the agent-coding shift means for teams, tools, and your next project.
Read article
Vinext 1.0: Next.js Apps Can Now Run on Vite. What It Means for Your Team
Cloudflare shipped Vinext 1.0, a production-ready reimplementation of the Next.js API on Vite, freeing Next.js apps to deploy to Workers, Node, Vercel, Netlify, AWS, or Deno. What it supports, the honest fine print, and what we recommend.
Read article