ZeroNoise Logo zeronoise
Post
Simon Willison: agent-built services need hard budget caps by default
•
5 min read
• 97 docs
Willison argues that hard spend limits should be the default for anything an agent deploys. Also: harness choice moves benchmark scores sharply, T3 Code can queue prompts until usage limits reset, Claude Code adds a "You should know" mod, and the Codex updates from DevDay.

Put hard spend caps on anything your agents deploy

Simon Willison argues that coding agents, and personal agents (which he calls "coding agents wrapped in a less threatening UI"), make it easy to spin up code that costs money: paid API calls, hosted apps, and storage or compute that bills as it grows . A warning email is not enough. Usage-based services need hard caps that cut the service off and return errors. Removing the cap should be an explicit opt-in . His reasoning: most people would rather see errors than a surprise bill of $10,000 or more .

Options you can use now:

  • AWS announced monthly project spend limits on Sept 16. A project that hits its limit is paused for the rest of the month. AWS's docs say the feature is still going out to a limited number of customers .
  • Google Cloud launched Spend Caps in July. They set a monthly cap on specific services within a project .

Willison also wants agents to lean toward recommending providers with hard caps, and to warn inexperienced builders before they deploy to uncapped services . You can do this today by adding that rule to your agent instructions.

The harness can matter as much as the model

Two results summarized by AINews point the same way. Hugging Face reports that the same model weights scored 62% in one harness and 33% in another. Its multi-harness RL setup uses a proxy that records token IDs and logprobs without changing the harnesses. It lifted LFM2.5-2.6B from 42% to 54% across four harnesses with 31% fewer tool calls, and the trainer, data and seven models are open . Separately, Meta Superintelligence Labs reports that a dedicated controller raised GPT-5.5 on ProgramBench from 63.7% to 71.5%, with the same workers and budget, against 58.0% for Codex . In practice, a model comparison means little unless you know which harness ran it.

T3 Code: queue prompts until limits reset, see spend across tools

  • Queued messages: T3 Code Nightly now lets you queue messages that fire when your usage limits reset .
  • Usage view: Theo improved the view so you can see where your "spend" goes and how models behave on your own data. It counts all Claude Code and Codex usage on your machines, not just usage inside T3 Code .
  • Model mix: Opus 5.5 is the first model to pass 50% of T3 Code traffic; "literally half of all prompts go to Opus" .
  • Orchestrator V2 is in the latest Nightly, and Theo is asking for reports on what broke or confused people . The mobile app only works through the TestFlight/beta build (install docs) .
  • Users: T3 Code passed 400K users and gained about 20K more within a day .

Claude Code mods in use

Anthropic added a built-in plugin, "You should Know." It scans Claude's output for important information you might miss. Enable it with /plugin enable cc-plugin-you-should-know@builtin. AINews describes it as spinning off a side agent, and describes mods as plugins with middleware-like hooks into Claude Code . A community example: @shawnbuilds built a custom Jev memory harness as a mod and posted the full prompt to rebuild it .

On a related habit, swyx says: "always run some kind of skill review/cutter after every model-created skill" .

Codex: DevDay features and a product refocus

Riley Brown's DevDay walkthrough covers the features most relevant to Codex users:

  • Codex CLI: easier voice input, a new agents view and better prompt editing .
  • Codex Cloud: you set up an environment once (signed into Clerk, Convex, Vercel, GitHub and so on), then reuse it for tasks run from your phone, with no local machine left on .
  • Code review: PRs can be reviewed from a side panel. You may need to pin "code review" there first .
  • GPT-6.1 Sol: OpenAI pitches it as a lower-cost model for coding and computer use. Brown quotes $2 in / $10 out per million tokens, against $10 / $50 for GPT-6 Astra .
  • Ultra-fast mode: about 8x faster output at 6x the usage, and only on the $500/month plan .

OpenAI's Tibo says the Codex team is "locking in": the only work now underway is simplification, efficiency to allow more usage, groundbreaking features and new models, because feedback says users want things simpler .

Smaller items

  • A dot clearing an inbox: Tibo reached inbox zero by giving his dot the goal. It deleted unneeded categories in validated batches, labeled email by type of work, and walked him through replies while it looked up context in the background . Simon Willison, by contrast, isn't sure when to use a dot over ChatGPT, because his dot seems built around a single conversation while he prefers managing context across threads .
  • Cheap computer-use targeting: ThePrimeagen uses Cloudflare's decision API for computer use. He raises its probabilities to chosen powers, then applies "probabilistic centering" so a command like "Click the monitor icon in the menu bar" lands on the right icon .
  • Pi 1.0 now includes native MCP support in Codemode by default, plus deferred tool loading, Anthropic cache warming and mid-conversation system messages .
  • Apple: Apple plans new Mac privacy controls that warn about granting broad data access to third-party software, including AI agents . DHH expects this to make macOS harder to use productively with agents .
Simon Willison: agent-built services need hard budget caps by default
Latent.Space
  • Pi 1.0 added native MCP support to Codemode, deferred tool loading, Anthropic cache warming, mid-conversation system messages, and support for non-LLM/image models and virtual-model extensions. Claude Code mods expose middleware-like hooks; its “You should know” plugin runs a side agent that flags important output the user might miss.
  • T3 Code’s rewrite added cross-provider delegate_task, thread forking, mid-thread model switching, subagent lineage views, scheduled tasks, Pi support, and an ACP registry. Cursor Rollouts can identify the PR behind a regression, open an issue, and offer a one-click cloud-agent fix.
  • Agent performance can vary substantially with the harness: identical weights scored 62% in one harness and 33% in another. Multi-harness training raised LFM2.5-2.6B from 42% to 54% across four harnesses and cut tool calls by 31%; the trainer, data, and seven trained models were released openly. Meta Superintelligence Labs also reported that a dedicated controller raised GPT-5.5’s ProgramBench score from 63.7% to 71.5% with the same workers and budget, versus 58.0% for Codex.
  • A user tested Qwen3.8-27B IQ4_XS via llama.cpp on one 24 GB RTX 4090 at a 196,608-token context, reporting 12/12 on a code-review task and 40/43 hidden tests on one DeepSWE task—but the binary pass score was 0, and the author cautioned this did not establish broad parity across the 113-task benchmark.
[AINews] not much happened today
Riley Brown
Profile
  • Codex Cloud provides a way to run coding tasks without leaving a local computer on: configure a cloud environment with services such as Clerk, Convex, Vercel, and GitHub, reuse it across tasks, and work from a phone.
  • Codex CLI gained voice input, an agents view, and improved prompt editing. Codex Code Review can review pull requests from its side panel; the presenter was unsure whether the announced security dashboard was publicly available.
  • OpenAI’s Agents API lets developers run agents in their own products; the video says setup requires a platform project and API key, and agents can use an OpenAI-hosted browser for computer use.
  • OpenAI’s newly announced 6.1 model is positioned as a lower-cost model for coding and computer use, which the presenter described as a practical workhorse based on initial use. Ultra-fast mode was described as roughly eight times faster while using six times the regular usage, and limited to the $500/month plan.
ChatGPT Just Changed A LOT (Here’s Everything New)
Simon Willison's Weblog

Treat hard spend caps as a deployment guardrail for coding-agent-built services: agents lower the friction of creating software that can incur paid API, hosting, storage, and compute charges, so the author argues usage services should stop at a hard limit rather than only send warnings, with removing the cap requiring explicit opt-in. AWS announced monthly project spend limits on September 16, 2026, which pause a project when its limit is reached, although the feature was still rolling out to a limited number of customers; Google Cloud had launched monthly Spend Caps for specific services within a project in July. The author suggests agents could recommend providers with hard caps and warn against deploying to uncapped services.

We're going to need default hard budget caps on pretty much everything
Riley Brown

Riley endorsed a principle for AI-assisted software creation: bring your own point of view and decide what deserves to exist; asking AI to choose without your intention risks plausible, easy-to-approve output, while greater speed makes judgment more important, not less.

Yes [https://x.com/ryolu_/status/2106337039201505453](https://x.com/ryolu_/status/2106337039201505453) convergence to the mean when everyone looks at everyone else to decide what to make, we eventually stop making anything of our own. we co…
Addy Osmani

Claude Code’s “You should Know” plugin scans Claude’s output for important information you might miss, helping keep you informed; enable it with /plugin enable cc-plugin-you-should-know@builtin. Addy Osmani calls it useful .

We're adding a new plugin to Claude Code: You should Know. It scans Claude's output for important information you might miss to help keep… "You should know" is a useful plugin for Claude Code that lets you know important info you might have missed. Enable with: /plugin enable…
ThePrimeagen

ThePrimeagen says he uses Cloudflare’s Decision API for computer use, applying what he calls a “noul” approach: raising probabilities to selected powers and then probabilistically centering them to click the correct icon. His example instruction is “Click the monitor icon in the menu bar.”

I effectively have decision api from cloudflare helping with computer use What I did is I used the idea of a noul, taking probabilities t…
Ben Tossell

Ben Tossell says personal agents do not use Opus 5.5 yet . He calls Opus “the goat” with early OpenClaw .

no personal agents use opus 5.5 (yet) opus was the goat with early openclaw
Kent C. Dodds 🐨

@_justelias says Kody can turn needs that would otherwise require a small service, script, credential, or webhook into capabilities usable by any agent, reducing glue work and avoiding rebuilding your setup for each agent . @kentcdodds endorses its broad applicability .

Once that pattern clicked, I started finding uses for Kody everywhere. Things that used to mean another little service, script, credentia… Once it clicks, you find use cases everywhere! [https://x.com/_justelias/status/2106444646377558402](https://x.com/_justelias/status/2106…
geoff

Geoffrey Huntley argues that the compile/restart/deploy loop could be collapsed, and predicts a shift toward adapting the older “actors/repair the image (don’t rebuild it)” idea for systems designed for LLMs as programmers rather than humans—a conceptual direction for coding-agent system design, not a concrete workflow.

dear young farts out there who have grown up with Docker/Kubernetes and the compile/restart/deploy phases of CI/CD. Did you know that the…
Riley Brown
  • Codex CLI adds voice input, an agents view, and improved prompt editing. Codex Cloud runs coding tasks without keeping a local computer on; configure an environment with services such as Clerk, Convex, Vercel, and GitHub once, then reuse it for cloud tasks from the phone.
  • Codex can review pull requests from its code-review panel. A security dashboard for checking app vulnerabilities was also announced, though the creator was unsure whether it was publicly available yet.
  • OpenAI’s Agents API is presented as a way to run Codex-powered agents inside developer-built products, using a platform project and API key, with computer use in an OpenAI-hosted browser. The announced Decisions API instead targets fast, constrained choices such as routing incoming messages to sales, billing, or support; the creator said it was not released yet.
  • The video says Ultra-fast mode generates output about eight times faster but uses six times the usage of regular mode and is limited to the $500/month Pro plan.
ChatGPT Just Changed A LOT (Here’s Everything New)
swyx

Run a skill review/cutter after every skill created by a model; the post strongly recommends making this a consistent step in coding-agent workflows.

[@BHolmesDev](https://x.com/BHolmesDev) [@aiDotEngineer](https://x.com/aiDotEngineer) 1000% always run some kind of skill review/cutter a…
Jason Zhou

Jason Zhou recommends a Claude Code Mods tutorial and highlights the /jev-memory mod as “pretty cool.” The linked post’s creator says they built a custom Jev memory harness using @typesafeai via @treg_ai and points readers to a full rebuild prompt.

Good Claude code mods tutorial /jev-memory mod is pretty cool And very happy seeing [@treg_ai](https://x.com/treg_ai) used here :) [https… Claude Code Mods make Claude 100% customizable, so I built a custom Jev memory harness built using [@typesafeai](https://x.com/typesafeai…
Peter Steinberger 🦞

Peter Steinberger said OpenClaw’s Android app had been in review limbo for over a week and asked whether anyone at Google could help; the post does not specify the review process or the app’s availability.

Do I know anyone at Google who could help? We're now over a week in review limbo for OpenClaw's Android app.
Riley Brown

Riley proposed a game-hosting platform designed for existing coding agents: a user would paste its link into Codex, Claude, or Grok and ask the agent to make a game multiplayer; the agent would use the platform’s docs to host and list it. He emphasized getting real-time database support, latency, and low-friction access right, and making the platform free initially.

Startup Idea: Make a platform that hosts video games and makes them multiplayer. Don't build an agent that builds video games. Waste of t…
Riley Brown

Riley Brown says he plugged in a ModRetro and asked Codex whether it was compatible with Muse Gadgets, presenting this as a way to make the device smart.

i turned my ModRetro into a smart device. by plugging it in and asking codex "is this device compatible with [https://gadgets.muse.ai/](h…
Theo - t3.gg

Theo said contributors to T3 Code had collectively burned “$10m+ in tokens.” He also reported using up four Claude accounts without hitting any of his Fable limits.

New branding for T3 Code: "We burned the tokens so you don't have to" We've probably burned $10m+ in tokens collectively just contributin… Killed 4 Claude accounts without a single percentage of my Fable limits being hit. Never thought I'd see the day. ![](https://pbs.twimg.c…
Riley Brown

Riley Brown says coding agents have moved beyond routine SaaS apps into multiplayer game mods, adding agents to older hardware, and designing objects for 3D printing, citing Codex and Claude Code as tools used for these projects.

Vibe Coding has completely evolved. It’s gotten so good at vibe coding “SAAS apps” that it’s no longer interesting. Now, people are vibe-…
Tibo

Tibo says current work is limited to simplifications, greater efficiency as usage grows, breakthrough features, or new models; he says user feedback clearly favors making things simpler.

All right, we’re locking in. Only things being worked on are simplifications, more efficiency for more usage, groundbreaking features or …
Theo - t3.gg

Theo asked users to try the latest T3 Code nightly and Orchestrator V2, then report what went wrong, what was confusing, what broke, and what to prioritize next. T3 Code also has a mobile app, but users must install its TestFlight/beta version; Theo linked installation instructions.

So who has installed the latest nightly for T3 Code and given Orchestrator V2 a shot? If you have, tell me what went wrong. What was conf… You can use the mobile app btw! You HAVE TO INSTALL THE TESTFLIGHT/BETA VERSION THOUGH More info here: [https://github.com/pingdotgg/t3co…