ZeroNoise Logo zeronoise
Post
Simon Willison: agent-built services need hard budget caps by default
•
5 min read
• 97 docs
Willison argues that hard spend limits should be the default for anything an agent deploys. Also: harness choice moves benchmark scores sharply, T3 Code can queue prompts until usage limits reset, Claude Code adds a "You should know" mod, and the Codex updates from DevDay.

Put hard spend caps on anything your agents deploy

Simon Willison argues that coding agents, and personal agents (which he calls "coding agents wrapped in a less threatening UI"), make it easy to spin up code that costs money: paid API calls, hosted apps, and storage or compute that bills as it grows . A warning email is not enough. Usage-based services need hard caps that cut the service off and return errors. Removing the cap should be an explicit opt-in . His reasoning: most people would rather see errors than a surprise bill of $10,000 or more .

Options you can use now:

  • AWS announced monthly project spend limits on Sept 16. A project that hits its limit is paused for the rest of the month. AWS's docs say the feature is still going out to a limited number of customers .
  • Google Cloud launched Spend Caps in July. They set a monthly cap on specific services within a project .

Willison also wants agents to lean toward recommending providers with hard caps, and to warn inexperienced builders before they deploy to uncapped services . You can do this today by adding that rule to your agent instructions.

The harness can matter as much as the model

Two results summarized by AINews point the same way. Hugging Face reports that the same model weights scored 62% in one harness and 33% in another. Its multi-harness RL setup uses a proxy that records token IDs and logprobs without changing the harnesses. It lifted LFM2.5-2.6B from 42% to 54% across four harnesses with 31% fewer tool calls, and the trainer, data and seven models are open . Separately, Meta Superintelligence Labs reports that a dedicated controller raised GPT-5.5 on ProgramBench from 63.7% to 71.5%, with the same workers and budget, against 58.0% for Codex . In practice, a model comparison means little unless you know which harness ran it.

T3 Code: queue prompts until limits reset, see spend across tools

  • Queued messages: T3 Code Nightly now lets you queue messages that fire when your usage limits reset .
  • Usage view: Theo improved the view so you can see where your "spend" goes and how models behave on your own data. It counts all Claude Code and Codex usage on your machines, not just usage inside T3 Code .
  • Model mix: Opus 5.5 is the first model to pass 50% of T3 Code traffic; "literally half of all prompts go to Opus" .
  • Orchestrator V2 is in the latest Nightly, and Theo is asking for reports on what broke or confused people . The mobile app only works through the TestFlight/beta build (install docs) .
  • Users: T3 Code passed 400K users and gained about 20K more within a day .

Claude Code mods in use

Anthropic added a built-in plugin, "You should Know." It scans Claude's output for important information you might miss. Enable it with /plugin enable cc-plugin-you-should-know@builtin. AINews describes it as spinning off a side agent, and describes mods as plugins with middleware-like hooks into Claude Code . A community example: @shawnbuilds built a custom Jev memory harness as a mod and posted the full prompt to rebuild it .

On a related habit, swyx says: "always run some kind of skill review/cutter after every model-created skill" .

Codex: DevDay features and a product refocus

Riley Brown's DevDay walkthrough covers the features most relevant to Codex users:

  • Codex CLI: easier voice input, a new agents view and better prompt editing .
  • Codex Cloud: you set up an environment once (signed into Clerk, Convex, Vercel, GitHub and so on), then reuse it for tasks run from your phone, with no local machine left on .
  • Code review: PRs can be reviewed from a side panel. You may need to pin "code review" there first .
  • GPT-6.1 Sol: OpenAI pitches it as a lower-cost model for coding and computer use. Brown quotes $2 in / $10 out per million tokens, against $10 / $50 for GPT-6 Astra .
  • Ultra-fast mode: about 8x faster output at 6x the usage, and only on the $500/month plan .

OpenAI's Tibo says the Codex team is "locking in": the only work now underway is simplification, efficiency to allow more usage, groundbreaking features and new models, because feedback says users want things simpler .

Smaller items

  • A dot clearing an inbox: Tibo reached inbox zero by giving his dot the goal. It deleted unneeded categories in validated batches, labeled email by type of work, and walked him through replies while it looked up context in the background . Simon Willison, by contrast, isn't sure when to use a dot over ChatGPT, because his dot seems built around a single conversation while he prefers managing context across threads .
  • Cheap computer-use targeting: ThePrimeagen uses Cloudflare's decision API for computer use. He raises its probabilities to chosen powers, then applies "probabilistic centering" so a command like "Click the monitor icon in the menu bar" lands on the right icon .
  • Pi 1.0 now includes native MCP support in Codemode by default, plus deferred tool loading, Anthropic cache warming and mid-conversation system messages .
  • Apple: Apple plans new Mac privacy controls that warn about granting broad data access to third-party software, including AI agents . DHH expects this to make macOS harder to use productively with agents .

Want personalized briefs on the topics you care about?

Create your own agent and get daily or weekly cited briefs on the topics you care about, shaped by the people, newsletters, podcasts, blogs, and channels you choose.