We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
Let an agent review commands instead of approving each one yourself
OpenAI made Codex's Auto-review free for anyone signed in with a ChatGPT account. Tibo also says it doesn't draw usage from your plan . A second agent reviews every action the primary agent takes. Its only jobs are to block high-risk actions and anything that doesn't match your original intent. This replaces the default sandbox, which asks you to approve everything and wears you down unless you spend time writing rules . To turn it on, go to Settings → Permissions → Auto-review, or pick "Approve for me" from the permissions menu below the composer . Tibo's roundup post adds that it "can be between 2-10% of plan when used," and doesn't say what that figure measures .
Anthropic made the same argument from its own data. Boris Cherny described a study in which contractors solving coding puzzles in Claude Code were sometimes shown injected commands that would have damaged their systems. They approved them "almost all the time," so per-command prompts had turned into "security theater" . Claude Code's auto mode sends each proposed action to a separate model that has none of the conversation's context. Cherny says it gets the call right "almost every time" . Mike Krieger says every Claude Code session now runs through that classifier by default. It asks you only in specific cases, which avoids both dangerously skipping permissions and endless yes-clicking .
For you, this means both major harnesses now offer a reviewer agent as the alternative to full access. If you have been running with permissions skipped, this is the setting to switch.
Model routing: measured savings, and a disagreement over what to route on
LangChain's Open SWE team published a router they built into their harness and measured :
- Map your tasks from traces. They sorted LangSmith traces into categories and rated complexity by median number of agent invocations and median LLM cost. They call this the most important step .
- Pick models on the cost/intelligence Pareto frontier. They chose GLM-5.3-Flash as the fast tier, GPT-5.6 Sol as balanced and GPT-6 Astra for performance .
- Route once, on the first message. The router picks a model from the thread's first message and keeps it for the whole task. It now runs on Jev, a dedicated decision model, after starting on a small LLM .
- Track outcomes. They ran an A/B test on merged-PR rate and user thumbs up/down .
Across almost 500 threads, compared with always using the performance model, median cost per thread fell 64% and P90 cost fell 37%. Merged PRs were slightly higher, but the difference wasn't statistically significant . Traffic went 34% to the fast tier, 56% to balanced and 10% to performance . A second test, router versus always using the fast model, was stopped within a day because of user complaints .
OpenAI's new Decisions API (public beta) is aimed at this same job: choosing a model, tool or action in near real time. OpenAI says it decides up to 10x faster than GPT-6 Luna through the Responses API . Theo pushed back: using it to expose the right tools to an agent is "possibly cool," but using it to decide what "level of intelligence" a task needs is "absolutely useless" . LangChain's numbers are one data point against him, though their task mix may not match yours.
Theo on Ultrafast: a different way of working, at a price
Theo's video on Astra Ultrafast includes workflow points that apply whenever inference gets fast:
- Split-screen your app and the agent, and steer it while it works. Send many small corrections instead of a batch of 20 . He sent five prompts in four minutes, with replies under a minute .
- Cut slow tool calls. Once generation drops from 10 minutes to 30 seconds, tool calls can double your runtime. He told the agent to stop using the preview browser and just change code, while he watched the dev server .
- He says staying in the loop got him a better UI from Astra than he gets from Opus, even though Astra is worse at design .
The costs: the main thread came to $36 and a follow-up to $250. GPT-6.1 Sol would have handled the work "just as well" for $12 . He advises against upgrading to the $500 tier for this . Separately, he says the $200 Codex plan "feels reasonable" with 6.1 Sol, while Astra on it is "pretty close to unusable." The $200 Claude Code plan feels meaningfully less limited with Opus .
Cherny: state the goal, the effort and the check
Boris Cherny's prompting advice: talk to Claude like a coworker and don't over-scaffold. Tell it what you want, how much effort to spend, and how it should verify the result . As an example he pointed to his earlier run where Opus 5.5 formally verified the Claude Agent SDK in Lean. A couple of short prompts produced 16 PRs fixing bugs and race conditions. He sometimes pairs Lean with TLA+ to check data flow, concurrency and state management, without knowing either language well . Krieger reports something similar: Fable 5.1 with no extra harness beat Hatch, a builder-plus-adversarial-verifier system that took months to tune. The only piece of Hatch he kept was dynamic workflows .
Mistral Large 4: promising, but configure your harness first
Mistral Large 4 is in API preview: 1T total parameters, 49B active, with open weights promised for the end of October. It costs $1.36/$4.18 per million input/output tokens . Mistral says it beats GLM 5.3 on DeepSWE and Kimi K3 on Terminal-Bench 4, and that it finished #2 behind Opus 5 in a blind Surge coding review. Critics say it trails GLM-5.3 on Artificial Analysis's index . Mistral says many reported failures come from not setting reasoning_effort="high" . Matthew Berman ran into a different problem in OpenCode: with high effort set, it hit about 32K tokens and stopped. Raising the context limit or turning thinking off helped. He found Cursor much easier for running third-party models .
Smaller items
- Cursor iOS: you can check in on, reply to or start agents running on your computer. They keep going if your phone loses signal. Pair at cursor.com/mobile; Enterprise admins have to enable it .
- Steinberger's team claw takes work requests from X. Unassigned sessions are open for anyone to claim, and the agent pings whoever last touched the related code. He says the whole setup was one prompt, with the server extending itself through hot-reloadable plugins .
- T3 Code: Opus 5.5 built and rendered mocks of five UI treatments inside the thread . Theo repeats his advice not to run agents on a Mac unless you have to .
- Willison had Codex work out how to run Parseable and send it Datasette's OpenTelemetry traces. He then wrote up the working patterns himself as a TIL .
- DHH: much of the work of getting the most out of agents is now product management, project management and QA .