ZeroNoise Logo zeronoise
Post
Simon Willison: agent-built services need hard budget caps by default
•
5 min read
• 97 docs
Willison argues that hard spend limits should be the default for anything an agent deploys. Also: harness choice moves benchmark scores sharply, T3 Code can queue prompts until usage limits reset, Claude Code adds a "You should know" mod, and the Codex updates from DevDay.

Put hard spend caps on anything your agents deploy

Simon Willison argues that coding agents, and personal agents (which he calls "coding agents wrapped in a less threatening UI"), make it easy to spin up code that costs money: paid API calls, hosted apps, and storage or compute that bills as it grows . A warning email is not enough. Usage-based services need hard caps that cut the service off and return errors. Removing the cap should be an explicit opt-in . His reasoning: most people would rather see errors than a surprise bill of $10,000 or more .

Options you can use now:

  • AWS announced monthly project spend limits on Sept 16. A project that hits its limit is paused for the rest of the month. AWS's docs say the feature is still going out to a limited number of customers .
  • Google Cloud launched Spend Caps in July. They set a monthly cap on specific services within a project .

Willison also wants agents to lean toward recommending providers with hard caps, and to warn inexperienced builders before they deploy to uncapped services . You can do this today by adding that rule to your agent instructions.

The harness can matter as much as the model

Two results summarized by AINews point the same way. Hugging Face reports that the same model weights scored 62% in one harness and 33% in another. Its multi-harness RL setup uses a proxy that records token IDs and logprobs without changing the harnesses. It lifted LFM2.5-2.6B from 42% to 54% across four harnesses with 31% fewer tool calls, and the trainer, data and seven models are open . Separately, Meta Superintelligence Labs reports that a dedicated controller raised GPT-5.5 on ProgramBench from 63.7% to 71.5%, with the same workers and budget, against 58.0% for Codex . In practice, a model comparison means little unless you know which harness ran it.

T3 Code: queue prompts until limits reset, see spend across tools

  • Queued messages: T3 Code Nightly now lets you queue messages that fire when your usage limits reset .
  • Usage view: Theo improved the view so you can see where your "spend" goes and how models behave on your own data. It counts all Claude Code and Codex usage on your machines, not just usage inside T3 Code .
  • Model mix: Opus 5.5 is the first model to pass 50% of T3 Code traffic; "literally half of all prompts go to Opus" .
  • Orchestrator V2 is in the latest Nightly, and Theo is asking for reports on what broke or confused people . The mobile app only works through the TestFlight/beta build (install docs) .
  • Users: T3 Code passed 400K users and gained about 20K more within a day .

Claude Code mods in use

Anthropic added a built-in plugin, "You should Know." It scans Claude's output for important information you might miss. Enable it with /plugin enable cc-plugin-you-should-know@builtin. AINews describes it as spinning off a side agent, and describes mods as plugins with middleware-like hooks into Claude Code . A community example: @shawnbuilds built a custom Jev memory harness as a mod and posted the full prompt to rebuild it .

On a related habit, swyx says: "always run some kind of skill review/cutter after every model-created skill" .

Codex: DevDay features and a product refocus

Riley Brown's DevDay walkthrough covers the features most relevant to Codex users:

  • Codex CLI: easier voice input, a new agents view and better prompt editing .
  • Codex Cloud: you set up an environment once (signed into Clerk, Convex, Vercel, GitHub and so on), then reuse it for tasks run from your phone, with no local machine left on .
  • Code review: PRs can be reviewed from a side panel. You may need to pin "code review" there first .
  • GPT-6.1 Sol: OpenAI pitches it as a lower-cost model for coding and computer use. Brown quotes $2 in / $10 out per million tokens, against $10 / $50 for GPT-6 Astra .
  • Ultra-fast mode: about 8x faster output at 6x the usage, and only on the $500/month plan .

OpenAI's Tibo says the Codex team is "locking in": the only work now underway is simplification, efficiency to allow more usage, groundbreaking features and new models, because feedback says users want things simpler .

Smaller items

  • A dot clearing an inbox: Tibo reached inbox zero by giving his dot the goal. It deleted unneeded categories in validated batches, labeled email by type of work, and walked him through replies while it looked up context in the background . Simon Willison, by contrast, isn't sure when to use a dot over ChatGPT, because his dot seems built around a single conversation while he prefers managing context across threads .
  • Cheap computer-use targeting: ThePrimeagen uses Cloudflare's decision API for computer use. He raises its probabilities to chosen powers, then applies "probabilistic centering" so a command like "Click the monitor icon in the menu bar" lands on the right icon .
  • Pi 1.0 now includes native MCP support in Codemode by default, plus deferred tool loading, Anthropic cache warming and mid-conversation system messages .
  • Apple: Apple plans new Mac privacy controls that warn about granting broad data access to third-party software, including AI agents . DHH expects this to make macOS harder to use productively with agents .
Simon Willison: agent-built services need hard budget caps by default
Summary
Coverage start
1 day ago
Coverage end
16 hours ago
Frequency
Daily
Published
15 hours ago
Reading time
5 min
Research time
1 hr 34 min
Documents scanned
97
Documents used
18
Citations
32
Sources monitored
111 / 111
Insights
Skipped contexts
Source details
Source Docs Insights Status
LangChain Blog 0 0
Brent Traut 0 0
Lukas Möller 0 0
Jediah Katz 0 0
Aman Karmani 0 0
Jacob Jackson 0 0
Cursor Blog | RSS Feed 0 0
Nicholas Moy 0 0
Mike Krieger 0 0
Sualeh Asif 0 0
Michael Truell 0 0
Google Antigravity 0 0
Aman Sanger 0 0
cat 0 0
Mark Chen 0 0
Greg Brockman 0 0
Tongzhou Wang 0 0
fouad 0 0
Calvin French-Owen 0 0
Hanson Wang 0 0
Ed Bayes 0 0
Alexander Embiricos 0 0
Tibo 7 1
Romain Huet 0 0
DHH 6 2
Jane Street Blog 0 0
Miguel Grinberg's Blog: AI 0 0
xxchan's Blog 0 0
<antirez> 0 0
Brendan Long 0 0
The Pragmatic Engineer 0 0
David Heinemeier Hansson 0 0
Armin Ronacher ⇌ 0 0
Mitchell Hashimoto 0 0
Armin Ronacher's Thoughts and Writings 0 0
Peter Steinberger 0 0
Theo - t3.gg 33 6
Sourcegraph 0 0
Anthropic 0 0
Cursor 0 0
LangChain 0 0
Anthropic 0 0
LangChain 0 0
Cursor 0 0
Riley Brown 1 1
Riley Brown 7 4
Jason Zhou 2 1
Boris Cherny 0 0
Mckay Wrigley 0 0
geoff 15 1
Peter Steinberger 🦞 3 1
AI Jason 0 0
Alex Albert 0 0
Latent.Space 1 1
Logan Kilpatrick 0 0
Fireship 0 0
Fireship 0 0
Kent C. Dodds 🐨 4 1
Practical AI 0 0
Practical AI Clips 0 0
Stories by Steve Yegge on Medium 0 0
Kent C. Dodds Blog 0 0
ThePrimeTime 0 0
Theo - t3․gg 0 0
ThePrimeagen 4 1
Ben Tossell 2 1
swyx 1 1
AI For Developers 0 0
Geoffrey Huntley 0 0
Addy Osmani 2 1
Andrej Karpathy 0 0
Simon Willison 6 0
Matthew Berman 0 0
Changelog 0 0
Simon Willison’s Newsletter 0 0
Agentic Coding Newsletter 0 0
Latent Space 0 0
Simon Willison's Weblog 2 1
Elevate 0 0
Lukas Möller 0 0
Jediah Katz 0 0
Sualeh Asif 0 0
Mike Krieger 0 0
Michael Truell 0 0
Cat Wu 0 0
Kevin Hou 0 0
Aman Sanger 0 0
Nicholas Moy 0 0
Andrey Mishchenko 0 0
Jerry Tworek 0 0
Romain Huet 0 0
Thibault Sottiaux 0 0
Alexander Embiricos 0 0
xxchan 0 0
Salvatore Sanfilippo 0 0
Armin Ronacher 0 0
David Heinemeier Hansson (DHH) 0 0
Alex Albert 0 0
Logan Kilpatrick 0 0
Shawn "swyx" Wang 0 0
Jason Zhou 0 0
Riley Brown 1 1
McKay Wrigley 0 0
Boris Cherny 0 0
Ben Tossell 0 0
Geoffrey Huntley 0 0
Peter Steinberger 0 0
Addy Osmani 0 0
Simon Willison 0 0
Andrej Karpathy 0 0
Harrison Chase 0 0