ZeroNoise Logo zeronoise
Post
Huntley: let types and simulators check agent code so cheaper models can do the work
•
6 min read
• 194 docs
Geoffrey Huntley argues that verification built into the language and a simulator let cheaper models handle more coding. Also: T3 Code's large multi-agent overhaul, a practical setup for using a dot with Codex, Theo's critique of NerfBench, and Airbnb's numbers.

Pick tools that check the agent's work, then use cheaper models

Geoffrey Huntley's new essay argues that code no longer has to be easy to read, as long as a model can explain it. The practical advice is about verification and spend. Huntley says a Rust codebase is far more maintainable with agents than a Python one. Compiler errors act as "back pressure" that the LLM picks up and fixes on every loop. Because the language does the checking, he can use cheaper models: "you need less intelligence to stay on the rails" . His own setup:

  • Models: GPT 6.1 Sol on no or low reasoning, plus GLM/Kimi. Match the model to the task, then buy those tokens wherever they're cheapest, for example open-weights GLM on Baseten. For company work, check for zero data retention first .
  • Simulator first: he is having agents rebuild a source-control system, a distributed system written in Rust. He built the simulator before anything else and has the agent validate all its work through it, which he says "has kept the agent on the rails remarkably well" .
  • Reading unfamiliar code: paste the function into an LLM and ask it to "explain this function to me as if you were explaining it to my son or daughter but in Python as a reference" .

Two related notes from Huntley. He says he can leave sol6x unattended on a goal for days, but feels he has to watch Opus 5.5 . He also noticed new sol6x behavior: when a goal needs verification loops, it writes the whole check as one Python script that takes state as variables. The loop then passes state in and runs the script to validate . Separately, a post Huntley shared suggests pointing agents at SCXML statecharts for distributed-systems work and links a demo of an agent porting Pi Durable to statecharts .

T3 Code's large overhaul has reached Nightly

T3 Code now has over 400,000 users . Theo merged a large overhaul and warned that Nightly will be unstable for the next few days. A stable release was cut first, so you can switch back to it . Notable additions :

  • delegate_task, which lets an agent start child agents on any provider or model. Native subagents appear as child threads with their model, status and history.
  • An ACP Registry for adding other agents (Devin, Cline, Kimi, Droid). Cursor now runs through its official SDK instead of its CLI. Pi and OpenCode 2 are supported.
  • Switching provider or model mid-thread, thread forking, attaching another thread as context with @ or by dragging it in, server-side queueing and steering, scheduled tasks, and auto-resume when usage limits reset.
  • MCP tools that let agents create, message, wait on and interrupt threads, plus tools for worktree handoff.

A beta option, "Hide threads while working," hides threads that are running. Theo reports 18 threads going without the UI feeling cramped: "Things appear when they need your attention, and disappear when they don't" .

Dots and Codex: "delegate to my dot, collab with Codex"

OpenAI's Dan Kundel wrote a concrete setup guide :

  • Set up the dot in the ChatGPT desktop app, connect it to local Codex, and have it spin up the Codex tasks you would otherwise manage yourself.
  • Hand over a repetitive task where you know what a good result looks like. Show the dot how you do it, refine the results together, and spell out what it may do on its own.
  • Use custom rules to decide when it must ask you first. By default it researches but doesn't take action .

Kundel sends backend and quick copy changes to Codex Cloud, which keeps running with his laptop off. Major frontend work stays local, where the agent can use his machine's context . He says dots currently don't count against ChatGPT usage limits unless they delegate to Codex . One dots-team engineer lets his dot watch a feedback channel, investigate issues, start Codex fix tasks and work through CI failures, and only major issues come back to him .

Proactive use is already paying off. Charlie Marsh's dot warned him that GitHub will deliberately fail macos-14 jobs on Oct 5. That affects Astral's uv tests, its Ruff/ty release builds, setup-uv and ruff-action . Alexander Embiricos's reusable prompt: forward the Slack message to your dot and say "Stay on top of this" .

On cost, ThePrimeagen used Codex's "ultra fast" mode on the $500 plan. One task with no tests or review took 1m41s, produced 722 insertions, and used 2% of his weekly limit. He estimates that works out to 40–100 minutes of agent runtime per week . He also reports that GLM 5.3 Flash and Kimi K3 will be served natively in Codex .

Check how a "nerf" benchmark works before you believe it

BridgeMind's NerfBench put Opus 5.5 at 94.2%, though it conceded this was "still inside normal variance" . Theo listed what its write-up leaves out: the harnesses, the tasks, the number of runs, how daily variance is handled, which APIs are used, and why tokens are weighted equally with cost . He also says BridgeMind's April "nerf" claim about Opus 4.6 rested on 6 of 30 tests . Apply the same questions to any model-regression claim.

Airbnb by the numbers

Airbnb says 60% of its code is now AI-authored. It reports shipping nearly 80% more features year over year, and PR throughput per engineer is up about 1.6x . It uses the strongest frontier model for coding, because "every defect that a model produces... could easily cost us a lot more" . Agents triggered by monitoring alerts already do first-line on-call triage. They either propose a PR for an engineer to review or close the incident if the alert was flaky . Engineers must be able to explain any PR the AI generated .

Smaller items

  • Jev as a semantic if: Jason Zhou's guide follows the rule "Code runs the flow. A model answers the small questions." Code branches on Jev's confidence scores, and uncertain cases (his example scores 0.38) go to a human . Treg reports ~10x lower cost and 18x more speed than GPT6-Luna in its tests, and has open-sourced its implementation. The guide puts Luna at 30x slower, so treat the speed figures as rough . ThePrimeagen found worthwhile fixes with jevlint. Its author calls it rough and says not to use its Cloudflare support yet .
  • Pi Durable: checkpointed steps let agents resume after crashes, it supports pluggable storage, background compaction runs without pausing the agent, and tool code can be swapped while the agent runs .
  • Claude Projects: Riley Brown has a main agent spin off separate threads (Opus 5.5 Medium by default) so the main thread stays short. He uses local threads when work needs his computer or GitHub, and cloud threads from his phone . After about 30 hours on Opus 5.5 and 4 on Sonnet 5.5, he prefers Opus if you can afford it .
  • Cursor Rollouts now traces a regression to the offending PR, opens an issue, and lets you start a cloud agent to fix it in one click. Usage credits are included through Oct 3 .
  • LangSmith Engine v2 reproduces an issue in the same environment with the same inputs, then builds and tests a fix and iterates on it before handing you a PR . Managed Deep Agents now supports memory scoped to a single user or a whole team .
  • Codex has a new /experimental flag that keeps the machine from sleeping .
Huntley: let types and simulators check agent code so cheaper models can do the work
Summary
Coverage start
1 day ago
Coverage end
17 hours ago
Frequency
Daily
Published
15 hours ago
Reading time
6 min
Research time
4 hrs 34 min
Documents scanned
194
Documents used
26
Citations
40
Sources monitored
111 / 111
Insights
Skipped contexts
Source details
Source Docs Insights Status
LangChain Blog 0 0
Brent Traut 0 0
Lukas Möller 0 0
Jediah Katz 0 0
Aman Karmani 0 0
Jacob Jackson 0 0
Cursor Blog | RSS Feed 0 0
Nicholas Moy 0 0
Mike Krieger 0 0
Sualeh Asif 0 0
Michael Truell 0 0
Google Antigravity 0 0
Aman Sanger 0 0
cat 0 0
Mark Chen 0 0
Greg Brockman 0 0
Tongzhou Wang 0 0
fouad 0 0
Calvin French-Owen 0 0
Hanson Wang 2 0
Ed Bayes 0 0
Alexander Embiricos 10 4
Tibo 7 3
Romain Huet 4 0
DHH 18 0
Jane Street Blog 0 0
Miguel Grinberg's Blog: AI 0 0
xxchan's Blog 0 0
<antirez> 0 0
Brendan Long 0 0
The Pragmatic Engineer 0 0
David Heinemeier Hansson 0 0
Armin Ronacher ⇌ 8 1
Mitchell Hashimoto 0 0
Armin Ronacher's Thoughts and Writings 0 0
Peter Steinberger 0 0
Theo - t3.gg 59 16
Sourcegraph 0 0
Anthropic 0 0
Cursor 0 0
LangChain 0 0
Anthropic 0 0
LangChain 7 2
Cursor 2 1
Riley Brown 1 1
Riley Brown 19 3
Jason Zhou 10 3
Boris Cherny 0 0
Mckay Wrigley 0 0
geoff 11 4
Peter Steinberger 🦞 1 0
AI Jason 0 0
Alex Albert 0 0
Latent.Space 2 2
Logan Kilpatrick 0 0
Fireship 0 0
Fireship 0 0
Kent C. Dodds 🐨 7 2
Practical AI 0 0
Practical AI Clips 0 0
Stories by Steve Yegge on Medium 0 0
Kent C. Dodds Blog 0 0
ThePrimeTime 1 1
Theo - t3․gg 0 0
ThePrimeagen 5 1
Ben Tossell 7 0
swyx 3 1
AI For Developers 0 0
Geoffrey Huntley 1 1
Addy Osmani 2 1
Andrej Karpathy 2 0
Simon Willison 0 0
Matthew Berman 0 0
Changelog 0 0
Simon Willison’s Newsletter 0 0
Agentic Coding Newsletter 0 0
Latent Space 1 1
Simon Willison's Weblog 0 0
Elevate 0 0
Lukas Möller 0 0
Jediah Katz 0 0
Sualeh Asif 0 0
Mike Krieger 0 0
Michael Truell 0 0
Cat Wu 0 0
Kevin Hou 0 0
Aman Sanger 0 0
Nicholas Moy 0 0
Andrey Mishchenko 0 0
Jerry Tworek 0 0
Romain Huet 0 0
Thibault Sottiaux 0 0
Alexander Embiricos 0 0
xxchan 0 0
Salvatore Sanfilippo 2 1
Armin Ronacher 0 0
David Heinemeier Hansson (DHH) 0 0
Alex Albert 0 0
Logan Kilpatrick 0 0
Shawn "swyx" Wang 0 0
Jason Zhou 0 0
Riley Brown 1 1
McKay Wrigley 0 0
Boris Cherny 0 0
Ben Tossell 0 0
Geoffrey Huntley 1 1
Peter Steinberger 0 0
Addy Osmani 0 0
Simon Willison 0 0
Andrej Karpathy 0 0
Harrison Chase 0 0