Agent token bills are a new kind of surprise cost
A developer's month with GLM 5.3 Flash included a painful story: picking the wrong model for a prototype burned 450M tokens, about $150 and 5kWh, almost overnight. The thread's advice was to pair a cheap coder with a good planner and a separate reviewer. On the Anthropic thread, users complained about hitting rate limits within 1h45 on afternoons and said a variable-cost 'extra usage' tier is hard to justify. A commenter on Offrun said the thing they want is per-account rate limit headroom shown before dispatching work, not after.
The AWS billing glitch fits the mood. Users got estimates of $286 million, $15 billion and even $284 billion on hobby accounts. It was a bug, but the panic shows how little trust people have in metering that scales without a ceiling.
The pattern: usage-based pricing combined with autonomous loops means cost control is now a core feature, not an afterthought.
So what?
If you sell agent-driven features, hard spend caps, pre-flight cost estimates and per-task budgets are features customers will ask for. If you buy them, put those limits in place before you let an agent run unattended.
Read these
One month coding with GLM 5.3 Flash
AWS: Inaccurate Estimated Billing Data – $1.7 billion
Tell HN: Anthropic no longer allowing Claude Code subscriptions to use OpenClaw
Show HN: Offrun – manage every coding agent from one workspace