ZeroNoise Logo zeronoise
Post
OpenAI and Anthropic both replace per-command approval with an agent reviewer; LangChain shows routing cut Open SWE costs 64%
•
6 min read
• 210 docs
Codex's Auto-review is now free, matching Anthropic's case that per-command approval has become security theater. Also: LangChain's measured model router, Theo's results with Astra Ultrafast, Boris Cherny's prompting advice, and early harness problems with Mistral Large 4.

Let an agent review commands instead of approving each one yourself

OpenAI made Codex's Auto-review free for anyone signed in with a ChatGPT account. Tibo also says it doesn't draw usage from your plan . A second agent reviews every action the primary agent takes. Its only jobs are to block high-risk actions and anything that doesn't match your original intent. This replaces the default sandbox, which asks you to approve everything and wears you down unless you spend time writing rules . To turn it on, go to Settings → Permissions → Auto-review, or pick "Approve for me" from the permissions menu below the composer . Tibo's roundup post adds that it "can be between 2-10% of plan when used," and doesn't say what that figure measures .

Anthropic made the same argument from its own data. Boris Cherny described a study in which contractors solving coding puzzles in Claude Code were sometimes shown injected commands that would have damaged their systems. They approved them "almost all the time," so per-command prompts had turned into "security theater" . Claude Code's auto mode sends each proposed action to a separate model that has none of the conversation's context. Cherny says it gets the call right "almost every time" . Mike Krieger says every Claude Code session now runs through that classifier by default. It asks you only in specific cases, which avoids both dangerously skipping permissions and endless yes-clicking .

For you, this means both major harnesses now offer a reviewer agent as the alternative to full access. If you have been running with permissions skipped, this is the setting to switch.

Model routing: measured savings, and a disagreement over what to route on

LangChain's Open SWE team published a router they built into their harness and measured :

  1. Map your tasks from traces. They sorted LangSmith traces into categories and rated complexity by median number of agent invocations and median LLM cost. They call this the most important step .
  2. Pick models on the cost/intelligence Pareto frontier. They chose GLM-5.3-Flash as the fast tier, GPT-5.6 Sol as balanced and GPT-6 Astra for performance .
  3. Route once, on the first message. The router picks a model from the thread's first message and keeps it for the whole task. It now runs on Jev, a dedicated decision model, after starting on a small LLM .
  4. Track outcomes. They ran an A/B test on merged-PR rate and user thumbs up/down .

Across almost 500 threads, compared with always using the performance model, median cost per thread fell 64% and P90 cost fell 37%. Merged PRs were slightly higher, but the difference wasn't statistically significant . Traffic went 34% to the fast tier, 56% to balanced and 10% to performance . A second test, router versus always using the fast model, was stopped within a day because of user complaints .

OpenAI's new Decisions API (public beta) is aimed at this same job: choosing a model, tool or action in near real time. OpenAI says it decides up to 10x faster than GPT-6 Luna through the Responses API . Theo pushed back: using it to expose the right tools to an agent is "possibly cool," but using it to decide what "level of intelligence" a task needs is "absolutely useless" . LangChain's numbers are one data point against him, though their task mix may not match yours.

Theo on Ultrafast: a different way of working, at a price

Theo's video on Astra Ultrafast includes workflow points that apply whenever inference gets fast:

  • Split-screen your app and the agent, and steer it while it works. Send many small corrections instead of a batch of 20 . He sent five prompts in four minutes, with replies under a minute .
  • Cut slow tool calls. Once generation drops from 10 minutes to 30 seconds, tool calls can double your runtime. He told the agent to stop using the preview browser and just change code, while he watched the dev server .
  • He says staying in the loop got him a better UI from Astra than he gets from Opus, even though Astra is worse at design .

The costs: the main thread came to $36 and a follow-up to $250. GPT-6.1 Sol would have handled the work "just as well" for $12 . He advises against upgrading to the $500 tier for this . Separately, he says the $200 Codex plan "feels reasonable" with 6.1 Sol, while Astra on it is "pretty close to unusable." The $200 Claude Code plan feels meaningfully less limited with Opus .

Cherny: state the goal, the effort and the check

Boris Cherny's prompting advice: talk to Claude like a coworker and don't over-scaffold. Tell it what you want, how much effort to spend, and how it should verify the result . As an example he pointed to his earlier run where Opus 5.5 formally verified the Claude Agent SDK in Lean. A couple of short prompts produced 16 PRs fixing bugs and race conditions. He sometimes pairs Lean with TLA+ to check data flow, concurrency and state management, without knowing either language well . Krieger reports something similar: Fable 5.1 with no extra harness beat Hatch, a builder-plus-adversarial-verifier system that took months to tune. The only piece of Hatch he kept was dynamic workflows .

Mistral Large 4: promising, but configure your harness first

Mistral Large 4 is in API preview: 1T total parameters, 49B active, with open weights promised for the end of October. It costs $1.36/$4.18 per million input/output tokens . Mistral says it beats GLM 5.3 on DeepSWE and Kimi K3 on Terminal-Bench 4, and that it finished #2 behind Opus 5 in a blind Surge coding review. Critics say it trails GLM-5.3 on Artificial Analysis's index . Mistral says many reported failures come from not setting reasoning_effort="high" . Matthew Berman ran into a different problem in OpenCode: with high effort set, it hit about 32K tokens and stopped. Raising the context limit or turning thinking off helped. He found Cursor much easier for running third-party models .

Smaller items

  • Cursor iOS: you can check in on, reply to or start agents running on your computer. They keep going if your phone loses signal. Pair at cursor.com/mobile; Enterprise admins have to enable it .
  • Steinberger's team claw takes work requests from X. Unassigned sessions are open for anyone to claim, and the agent pings whoever last touched the related code. He says the whole setup was one prompt, with the server extending itself through hot-reloadable plugins .
  • T3 Code: Opus 5.5 built and rendered mocks of five UI treatments inside the thread . Theo repeats his advice not to run agents on a Mac unless you have to .
  • Willison had Codex work out how to run Parseable and send it Datasette's OpenTelemetry traces. He then wrote up the working patterns himself as a TIL .
  • DHH: much of the work of getting the most out of agents is now product management, project management and QA .
OpenAI and Anthropic both replace per-command approval with an agent reviewer; LangChain shows routing cut Open SWE costs 64%
Summary
Coverage start
1 day ago
Coverage end
17 hours ago
Frequency
Daily
Published
16 hours ago
Reading time
6 min
Research time
5 hrs 14 min
Documents scanned
210
Documents used
23
Citations
42
Sources monitored
111 / 111
Insights
Skipped contexts
Source details
Source Docs Insights Status
LangChain Blog 0 0
Brent Traut 0 0
Lukas Möller 0 0
Jediah Katz 2 1
Aman Karmani 0 0
Jacob Jackson 0 0
Cursor Blog | RSS Feed 0 0
Nicholas Moy 2 0
Mike Krieger 0 0
Sualeh Asif 0 0
Michael Truell 0 0
Google Antigravity 2 1
Aman Sanger 0 0
cat 0 0
Mark Chen 0 0
Greg Brockman 2 0
Tongzhou Wang 0 0
fouad 0 0
Calvin French-Owen 0 0
Hanson Wang 2 0
Ed Bayes 0 0
Alexander Embiricos 6 3
Tibo 24 6
Romain Huet 0 0
DHH 17 6
Jane Street Blog 0 0
Miguel Grinberg's Blog: AI 0 0
xxchan's Blog 0 0
<antirez> 0 0
Brendan Long 0 0
The Pragmatic Engineer 0 0
David Heinemeier Hansson 1 0
Armin Ronacher ⇌ 5 3
Mitchell Hashimoto 0 0
Armin Ronacher's Thoughts and Writings 0 0
Peter Steinberger 0 0
Theo - t3.gg 53 9
Sourcegraph 1 1
Anthropic 1 1
Cursor 0 0
LangChain 2 2
Anthropic 0 0
LangChain 10 4
Cursor 3 1
Riley Brown 0 0
Riley Brown 13 4
Jason Zhou 4 2
Boris Cherny 6 2
Mckay Wrigley 0 0
geoff 5 0
Peter Steinberger 🦞 5 2
AI Jason 0 0
Alex Albert 0 0
Latent.Space 2 2
Logan Kilpatrick 1 0
Fireship 0 0
Fireship 0 0
Kent C. Dodds 🐨 12 3
Practical AI 0 0
Practical AI Clips 0 0
Stories by Steve Yegge on Medium 0 0
Kent C. Dodds Blog 0 0
ThePrimeTime 0 0
Theo - t3․gg 1 1
ThePrimeagen 0 0
Ben Tossell 1 0
swyx 1 0
AI For Developers 0 0
Geoffrey Huntley 0 0
Addy Osmani 0 0
Andrej Karpathy 0 0
Simon Willison 9 2
Matthew Berman 1 1
Changelog 0 0
Simon Willison’s Newsletter 0 0
Agentic Coding Newsletter 0 0
Latent Space 0 0
Simon Willison's Weblog 10 6
Elevate 0 0
Lukas Möller 0 0
Jediah Katz 0 0
Sualeh Asif 0 0
Mike Krieger 1 1
Michael Truell 0 0
Cat Wu 0 0
Kevin Hou 0 0
Aman Sanger 0 0
Nicholas Moy 0 0
Andrey Mishchenko 0 0
Jerry Tworek 0 0
Romain Huet 0 0
Thibault Sottiaux 0 0
Alexander Embiricos 0 0
xxchan 0 0
Salvatore Sanfilippo 3 2
Armin Ronacher 0 0
David Heinemeier Hansson (DHH) 0 0
Alex Albert 0 0
Logan Kilpatrick 0 0
Shawn "swyx" Wang 1 1
Jason Zhou 0 0
Riley Brown 0 0
McKay Wrigley 0 0
Boris Cherny 1 1
Ben Tossell 0 0
Geoffrey Huntley 0 0
Peter Steinberger 0 0
Addy Osmani 0 0
Simon Willison 0 0
Andrej Karpathy 0 0
Harrison Chase 0 0