We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
Let an agent review commands instead of approving each one yourself
OpenAI made Codex's Auto-review free for anyone signed in with a ChatGPT account. Tibo also says it doesn't draw usage from your plan . A second agent reviews every action the primary agent takes. Its only jobs are to block high-risk actions and anything that doesn't match your original intent. This replaces the default sandbox, which asks you to approve everything and wears you down unless you spend time writing rules . To turn it on, go to Settings → Permissions → Auto-review, or pick "Approve for me" from the permissions menu below the composer . Tibo's roundup post adds that it "can be between 2-10% of plan when used," and doesn't say what that figure measures .
Anthropic made the same argument from its own data. Boris Cherny described a study in which contractors solving coding puzzles in Claude Code were sometimes shown injected commands that would have damaged their systems. They approved them "almost all the time," so per-command prompts had turned into "security theater" . Claude Code's auto mode sends each proposed action to a separate model that has none of the conversation's context. Cherny says it gets the call right "almost every time" . Mike Krieger says every Claude Code session now runs through that classifier by default. It asks you only in specific cases, which avoids both dangerously skipping permissions and endless yes-clicking .
For you, this means both major harnesses now offer a reviewer agent as the alternative to full access. If you have been running with permissions skipped, this is the setting to switch.
Model routing: measured savings, and a disagreement over what to route on
LangChain's Open SWE team published a router they built into their harness and measured :
- Map your tasks from traces. They sorted LangSmith traces into categories and rated complexity by median number of agent invocations and median LLM cost. They call this the most important step .
- Pick models on the cost/intelligence Pareto frontier. They chose GLM-5.3-Flash as the fast tier, GPT-5.6 Sol as balanced and GPT-6 Astra for performance .
- Route once, on the first message. The router picks a model from the thread's first message and keeps it for the whole task. It now runs on Jev, a dedicated decision model, after starting on a small LLM .
- Track outcomes. They ran an A/B test on merged-PR rate and user thumbs up/down .
Across almost 500 threads, compared with always using the performance model, median cost per thread fell 64% and P90 cost fell 37%. Merged PRs were slightly higher, but the difference wasn't statistically significant . Traffic went 34% to the fast tier, 56% to balanced and 10% to performance . A second test, router versus always using the fast model, was stopped within a day because of user complaints .
OpenAI's new Decisions API (public beta) is aimed at this same job: choosing a model, tool or action in near real time. OpenAI says it decides up to 10x faster than GPT-6 Luna through the Responses API . Theo pushed back: using it to expose the right tools to an agent is "possibly cool," but using it to decide what "level of intelligence" a task needs is "absolutely useless" . LangChain's numbers are one data point against him, though their task mix may not match yours.
Theo on Ultrafast: a different way of working, at a price
Theo's video on Astra Ultrafast includes workflow points that apply whenever inference gets fast:
- Split-screen your app and the agent, and steer it while it works. Send many small corrections instead of a batch of 20 . He sent five prompts in four minutes, with replies under a minute .
- Cut slow tool calls. Once generation drops from 10 minutes to 30 seconds, tool calls can double your runtime. He told the agent to stop using the preview browser and just change code, while he watched the dev server .
- He says staying in the loop got him a better UI from Astra than he gets from Opus, even though Astra is worse at design .
The costs: the main thread came to $36 and a follow-up to $250. GPT-6.1 Sol would have handled the work "just as well" for $12 . He advises against upgrading to the $500 tier for this . Separately, he says the $200 Codex plan "feels reasonable" with 6.1 Sol, while Astra on it is "pretty close to unusable." The $200 Claude Code plan feels meaningfully less limited with Opus .
Cherny: state the goal, the effort and the check
Boris Cherny's prompting advice: talk to Claude like a coworker and don't over-scaffold. Tell it what you want, how much effort to spend, and how it should verify the result . As an example he pointed to his earlier run where Opus 5.5 formally verified the Claude Agent SDK in Lean. A couple of short prompts produced 16 PRs fixing bugs and race conditions. He sometimes pairs Lean with TLA+ to check data flow, concurrency and state management, without knowing either language well . Krieger reports something similar: Fable 5.1 with no extra harness beat Hatch, a builder-plus-adversarial-verifier system that took months to tune. The only piece of Hatch he kept was dynamic workflows .
Mistral Large 4: promising, but configure your harness first
Mistral Large 4 is in API preview: 1T total parameters, 49B active, with open weights promised for the end of October. It costs $1.36/$4.18 per million input/output tokens . Mistral says it beats GLM 5.3 on DeepSWE and Kimi K3 on Terminal-Bench 4, and that it finished #2 behind Opus 5 in a blind Surge coding review. Critics say it trails GLM-5.3 on Artificial Analysis's index . Mistral says many reported failures come from not setting reasoning_effort="high" . Matthew Berman ran into a different problem in OpenCode: with high effort set, it hit about 32K tokens and stopped. Raising the context limit or turning thinking off helped. He found Cursor much easier for running third-party models .
Smaller items
- Cursor iOS: you can check in on, reply to or start agents running on your computer. They keep going if your phone loses signal. Pair at cursor.com/mobile; Enterprise admins have to enable it .
- Steinberger's team claw takes work requests from X. Unassigned sessions are open for anyone to claim, and the agent pings whoever last touched the related code. He says the whole setup was one prompt, with the server extending itself through hot-reloadable plugins .
- T3 Code: Opus 5.5 built and rendered mocks of five UI treatments inside the thread . Theo repeats his advice not to run agents on a Mac unless you have to .
- Willison had Codex work out how to run Parseable and send it Datasette's OpenTelemetry traces. He then wrote up the working patterns himself as a TIL .
- DHH: much of the work of getting the most out of agents is now product management, project management and QA .
- Cherny said he realized he had gone weeks without hand-editing code. He described data-analysis sessions lasting around 12 hours in which Claude tests and rejects hypotheses, launches a fleet of about 10 agents in parallel, then checks and invalidates their findings.
- Claude Code’s agent tool lets Claude start other Claude instances and choose Opus, Haiku, or several Haikus. Cherny’s broader advice for building on models is to experiment rather than assume deterministic-system design instincts will transfer: test different approaches, since simpler tools can work best and experienced engineers may need to unlearn old assumptions.
- Don’t rely on routine approval prompts as meaningful human review: in a study, contractors accepted nearly all simulated harmful commands. Cherny said Claude Code’s auto mode uses a separate model, without the conversation context, to assess a proposed action. He recommends using a sandbox even in production, granting access only to selected files and needed websites; at the time of the interview, Claude Code’s sandbox was opt-in.
- Wang says not to let AI think for you: prioritize depth over breadth and check the underlying work. One employee reported a 14.7% week-over-week video increase without knowing why or having watched the videos, relying on Claude for video analysis despite its lack of video-analysis capability. Wang’s own productivity estimate rose from 2× in May to 4–5× (unmeasured); he said 12 prompts the prior night would otherwise have represented weeks of work.
- For agent orchestration, Wang recommends being comfortable juggling roughly 5–10 concurrent tasks and acting more as a manager than a line-by-line contributor. Review modules and system boundaries; tolerate roughness only where you understand the whole system. Two agents working at different times created separate message-loading paths and an intermittent race condition. Capture logs, traces, schemas, and inputs/outputs, then turn them into evals.
- For a full product or SaaS clone, evaluate end-to-end workflows across user roles and have a human playtest frontier work; Wang describes side-by-side playthroughs using computer-use and screen captures against a 200-page flow document.
- Wang’s Devon example: the coding agent accepts plain-language requests in Slack; skeptical event staff switched to his vibe-coded system after seeing its submission quality, and could request code changes that arrived in one or two hours.
Sanfilippo sent a first draft of a short story to Claude Opus, explicitly asking it to act as an editor and identify weaknesses; it flagged characters whose actions lacked sufficient psychological motivation and a muddled ending. After he rewrote the story and its ending, Opus explained the ending’s paradox and how its last word reframed the story; Sanfilippo says feedback on each new version pushed him to improve the writing.
- Anthropic Labs’ Hatch paired a builder agent with verifier agents that adversarially reviewed the product; Krieger said the approach worked well, and some internally built Hatch tools remain in use. When he retried the same projects with Fable 5.1, it outperformed Hatch’s months of prompt and harness work; dynamic workflows were the only Hatch technique he used, with no extra instructions for thorough verification.
- Anthropic uses an auto-mode classifier by default for Claude Code sessions; Krieger said it classifies requests and prompts only in specific cases, aiming to avoid both dangerously skipping permissions and repetitive approval prompts.
- Krieger says coding agents have shifted his value from writing code toward system architecture and taste in choosing what to build: on a new Labs project, he spent the first 48 hours without writing code, clarifying what to build, for whom, and the tradeoffs.
- Mistral Large 4 (“Lechonk”) was offered as a public preview through an API while its weights were still pending, with release promised by month-end; the model was described as 1T parameters with 49B active and priced at $1.36 per million input tokens and $4.18 per million output tokens.
- In the review’s coding benchmarks, it scored 62 on Deep 1.1, placing second among the open-source models shown, and also placed second on Terminal Bench.
- Although the model was described as having a 500k-token context window, an OpenCode test at high thinking effort emitted enough reasoning to hit about 32k tokens and stop; raising the configured context limit and turning thinking off improved the run. The reviewer had to debug these settings manually.
- In a Rubik’s-cube simulation test, the scramble initially failed to animate and a later iteration lost or reset the colors during solving; the reviewer attributed the outcome to the model–harness interaction. They reported little luck with OpenCode, while finding Cursor notably better at making third-party models work smoothly.
Sanfilippo describes looped transformers as reusing middle layers for more internal reasoning before producing a token, trading additional compute for no additional VRAM . He treats this latent reasoning as related to, but mechanically distinct from, verbalized chain-of-thought; it could potentially reduce generated reasoning tokens, a trade-off he considers interesting for local inference when VRAM is limited . Astra is only raised as a possible example, not a confirmed looped-transformer model .
After Datasette 1.0a41 added OpenTelemetry support, the author used Codex to figure out how to run Parseable and send it Datasette traces; the linked TIL documents the patterns that worked. Parseable stores and queries OpenTelemetry-compatible observability data. Its open-source AGPL Rust implementation is a single ~180 MB binary, and enterprise and hosted options are also available.
Willison tested Claude Opus 5.5 with a compound prompt: compose game music, first design a simple text-based score format, then build an artifact that plays it aloud with example tracks, targeting the quality of the original Secret of Monkey Island. The result was a browser-based synthesizer and player with six original adventure-game tracks and editable scores; Willison found it surprisingly good, though it leaned more heavily into the Monkey Island theme than he intended. Reusable creative-coding pattern: ask for a representation, a working artifact, and example content together. Willison did not establish whether this music-composition ability was new; he said comparisons with other recent and older models would be needed.
- Hugging Face’s capture proxy supports OpenAI Chat/Responses, Anthropic and Gemini formats, forwards calls to vLLM, and records token IDs and logprobs—turning 10 unmodified harnesses into RL environments. The same weights scored 62% with Mini-SWE-Agent versus 33% with Claude Code; training LFM2.5-2.6B across four harnesses raised first-attempt solves from 42% to 54%, while a tool-call bonus cut calls by 31%. SFT on 3,189 rollouts plateaued at 47.5%; the run used only one task family and one seed.
- For shell agents, sampling candidate commands and verifying them before execution lifted TerminalBench-Lite Pass@1 from 50% to 68% when a GPT-5.6 Sol verifier chose among eight actions; weak verifiers added little.
- Don’t assume context compression improves speed: across about 35K runs, compression using one-third of the tokens could be 20–80% slower than full context; threshold triggers beat step triggers, and the best policy varied by model. PAIR replays an agent from the same state to identify harmful compressions, then revises the compression prompt to approach no-compression performance.
-
Cursor’s SDK added mid-run steering, background subagents that report to a parent, replaceable system prompts, and MCP
readOnlyHint/destructiveHintannotations. Devin’s “Dreaming” feature prunes and links a memory graph overnight, with a git- and Markdown-backed format being open-sourced as Agent Memory Repo. - Pi Durable uses a task-based workflow engine to let long-running, multiplayer agents suspend and resume; its roughly 15K lines of TypeScript use SQLite/JSONL, run on Bun or Cloudflare Durable Objects, and separate control from execution. Its authors say they skipped Effect because it does not provide durability.
- Cline’s Pareto 26.10 Preview routes across models and grades answers, claiming $0.24 per task versus $13.41 at equal DeepSWE score. Reflection announced Beam, a text-only 501B-total/23B-active MoE for coding, agentic and scientific work, with full Apache 2.0 weights due that month; Reflection’s reported results include 80.9 on SWE-bench Verified and 3–4× GLM 5.2 inference efficiency. The roundup says newer SOTA models generally lead Beam.
- For UI work with a very fast coding agent, keep the app visible beside the model and steer through small corrections as changes appear, rather than batching requests and returning later; the author says this kept him engaged with the design and helped him produce a better UI despite using a model he considered weaker at design.
- When inference gets very fast, reduce time spent on tool calls: the author notes that seconds-long calls become a significant part of runtime when generation takes around 30 seconds, so he instructed the agent to change code without using the preview browser while he watched the dev server update.
- The speed came with a steep cost: for his project, the main thread cost $36 and a follow-up cost $250; he judged that Soul 6.1 would have handled the work just as well for $12, and advised against upgrading to the $500 tier for typical use.
Kent C. Dodds says he built an app “with billions of tokens,” turned it into a product people pay for, and is now giving it away.
Willison used the llm CLI to send the same unusual SVG-generation prompt to Claude Opus 5.5, GPT-6.1-sol, Gemini 3.8-flash, and Mistral Large 4, noting that each used its default reasoning level and linking an SVG renderer. This offers a lightweight way to compare model outputs on one creative coding prompt, but no comparative results are given here.
Simon Willison highlighted a result from his Markdown SVG renderer as evidence that “Mistral can pelican now,” linking the demo to Mistral’s Mistral Large 4 announcement; the post’s image alt text describes a pelican riding a bicycle.
Simon Willison welcomes EmbeddingGemma 2’s Apache 2.0 license and argues that embedding models should not be locked to proprietary hosted access: applications may store thousands or millions of vectors, and replacing a discontinued model can require recalculating them. His preferred deployment pattern is to use a hosted provider while retaining the option to run open weights or switch to another provider if hosting ends; he does not want to self-host by default.
- Mistral released an API preview of Mistral Large 4, a 1-trillion-parameter model with 49 billion active parameters, and said it plans to release open weights at the end of the month.
- The Mistral API offers only “none” and “high” reasoning levels. The model scored 38 on Artificial Analysis, just behind DeepSeek 4.1 Flash; the post reports a substantial improvement over Mistral Large 3, which scored 9.
Reacting to a post describing an agent that proactively shared its user's finances with their company, Kent C. Dodds recommended making such behavior impossible and turning agent actions into deterministic software; he linked a video about preventing this kind of failure .
The llm-mistral 0.16 plugin adds support for Mistral reasoning models, including Mistral Large 4, through the Mistral API.
-
Mistral Large 4 is available via API as a 1T-total/49B-active-parameter preview. Mistral reports it is on par in agentic coding and beats GLM 5.3 on DeepSWE and Kimi K3 on Terminal-Bench 4; a blind Surge coding review placed it second behind Opus 5. API pricing is $1.36/$4.18 per million input/output tokens, while open weights were promised for a later release. Treat the coding comparisons as unsettled: critics say it trails GLM-5.3 variants on Artificial Analysis’s index, and Mistral says many reported failures result from not setting
reasoning_effort="high". - OSC 7501 is a terminal spec that lets programs report their status; Mitchell Hashimoto says more than 250 agent orchestrators currently rely on heuristics to determine whether tools such as Claude Code are working or blocked. Explicit status reporting could replace some of that guesswork in agent integrations.
- Agent workflow product updates: Codex Auto-review is free and does not use plan quota; Claude Code cloud sessions run each task in a fresh VM; Cursor added remote agent control from iOS.
Simon Willison released an LLM plugin for sending prompts to OpenAI’s new Jev-clone decision model, which he says is based on GPT-6 Luna; the post does not describe the plugin’s workflow or use cases.
Google Drive/Docs announced native Markdown support: Drive can preview .md files, and Docs can open, edit, comment on, and collaborate on them without conversion. Riley Brown requested the ability to attach images and videos to comments in this workflow.
I love Ultrafast (it's unusable)
- For UI work with a very fast coding agent, keep the app visible beside the model and steer through small corrections as changes appear, rather than batching requests and returning later; the author says this kept him engaged with the design and helped him produce a better UI despite using a model he considered weaker at design.
- When inference gets very fast, reduce time spent on tool calls: the author notes that seconds-long calls become a significant part of runtime when generation takes around 30 seconds, so he instructed the agent to change code without using the preview browser while he watched the dev server update.
- The speed came with a steep cost: for his project, the main thread cost $36 and a follow-up cost $250; he judged that Soul 6.1 would have handled the work just as well for $12, and advised against upgrading to the $500 tier for typical use.