ZeroNoise Logo zeronoise
Post
Huntley: let types and simulators check agent code so cheaper models can do the work
•
6 min read
• 194 docs
Geoffrey Huntley argues that verification built into the language and a simulator let cheaper models handle more coding. Also: T3 Code's large multi-agent overhaul, a practical setup for using a dot with Codex, Theo's critique of NerfBench, and Airbnb's numbers.

Pick tools that check the agent's work, then use cheaper models

Geoffrey Huntley's new essay argues that code no longer has to be easy to read, as long as a model can explain it. The practical advice is about verification and spend. Huntley says a Rust codebase is far more maintainable with agents than a Python one. Compiler errors act as "back pressure" that the LLM picks up and fixes on every loop. Because the language does the checking, he can use cheaper models: "you need less intelligence to stay on the rails" . His own setup:

  • Models: GPT 6.1 Sol on no or low reasoning, plus GLM/Kimi. Match the model to the task, then buy those tokens wherever they're cheapest, for example open-weights GLM on Baseten. For company work, check for zero data retention first .
  • Simulator first: he is having agents rebuild a source-control system, a distributed system written in Rust. He built the simulator before anything else and has the agent validate all its work through it, which he says "has kept the agent on the rails remarkably well" .
  • Reading unfamiliar code: paste the function into an LLM and ask it to "explain this function to me as if you were explaining it to my son or daughter but in Python as a reference" .

Two related notes from Huntley. He says he can leave sol6x unattended on a goal for days, but feels he has to watch Opus 5.5 . He also noticed new sol6x behavior: when a goal needs verification loops, it writes the whole check as one Python script that takes state as variables. The loop then passes state in and runs the script to validate . Separately, a post Huntley shared suggests pointing agents at SCXML statecharts for distributed-systems work and links a demo of an agent porting Pi Durable to statecharts .

T3 Code's large overhaul has reached Nightly

T3 Code now has over 400,000 users . Theo merged a large overhaul and warned that Nightly will be unstable for the next few days. A stable release was cut first, so you can switch back to it . Notable additions :

  • delegate_task, which lets an agent start child agents on any provider or model. Native subagents appear as child threads with their model, status and history.
  • An ACP Registry for adding other agents (Devin, Cline, Kimi, Droid). Cursor now runs through its official SDK instead of its CLI. Pi and OpenCode 2 are supported.
  • Switching provider or model mid-thread, thread forking, attaching another thread as context with @ or by dragging it in, server-side queueing and steering, scheduled tasks, and auto-resume when usage limits reset.
  • MCP tools that let agents create, message, wait on and interrupt threads, plus tools for worktree handoff.

A beta option, "Hide threads while working," hides threads that are running. Theo reports 18 threads going without the UI feeling cramped: "Things appear when they need your attention, and disappear when they don't" .

Dots and Codex: "delegate to my dot, collab with Codex"

OpenAI's Dan Kundel wrote a concrete setup guide :

  • Set up the dot in the ChatGPT desktop app, connect it to local Codex, and have it spin up the Codex tasks you would otherwise manage yourself.
  • Hand over a repetitive task where you know what a good result looks like. Show the dot how you do it, refine the results together, and spell out what it may do on its own.
  • Use custom rules to decide when it must ask you first. By default it researches but doesn't take action .

Kundel sends backend and quick copy changes to Codex Cloud, which keeps running with his laptop off. Major frontend work stays local, where the agent can use his machine's context . He says dots currently don't count against ChatGPT usage limits unless they delegate to Codex . One dots-team engineer lets his dot watch a feedback channel, investigate issues, start Codex fix tasks and work through CI failures, and only major issues come back to him .

Proactive use is already paying off. Charlie Marsh's dot warned him that GitHub will deliberately fail macos-14 jobs on Oct 5. That affects Astral's uv tests, its Ruff/ty release builds, setup-uv and ruff-action . Alexander Embiricos's reusable prompt: forward the Slack message to your dot and say "Stay on top of this" .

On cost, ThePrimeagen used Codex's "ultra fast" mode on the $500 plan. One task with no tests or review took 1m41s, produced 722 insertions, and used 2% of his weekly limit. He estimates that works out to 40–100 minutes of agent runtime per week . He also reports that GLM 5.3 Flash and Kimi K3 will be served natively in Codex .

Check how a "nerf" benchmark works before you believe it

BridgeMind's NerfBench put Opus 5.5 at 94.2%, though it conceded this was "still inside normal variance" . Theo listed what its write-up leaves out: the harnesses, the tasks, the number of runs, how daily variance is handled, which APIs are used, and why tokens are weighted equally with cost . He also says BridgeMind's April "nerf" claim about Opus 4.6 rested on 6 of 30 tests . Apply the same questions to any model-regression claim.

Airbnb by the numbers

Airbnb says 60% of its code is now AI-authored. It reports shipping nearly 80% more features year over year, and PR throughput per engineer is up about 1.6x . It uses the strongest frontier model for coding, because "every defect that a model produces... could easily cost us a lot more" . Agents triggered by monitoring alerts already do first-line on-call triage. They either propose a PR for an engineer to review or close the incident if the alert was flaky . Engineers must be able to explain any PR the AI generated .

Smaller items

  • Jev as a semantic if: Jason Zhou's guide follows the rule "Code runs the flow. A model answers the small questions." Code branches on Jev's confidence scores, and uncertain cases (his example scores 0.38) go to a human . Treg reports ~10x lower cost and 18x more speed than GPT6-Luna in its tests, and has open-sourced its implementation. The guide puts Luna at 30x slower, so treat the speed figures as rough . ThePrimeagen found worthwhile fixes with jevlint. Its author calls it rough and says not to use its Cloudflare support yet .
  • Pi Durable: checkpointed steps let agents resume after crashes, it supports pluggable storage, background compaction runs without pausing the agent, and tool code can be swapped while the agent runs .
  • Claude Projects: Riley Brown has a main agent spin off separate threads (Opus 5.5 Medium by default) so the main thread stays short. He uses local threads when work needs his computer or GitHub, and cloud threads from his phone . After about 30 hours on Opus 5.5 and 4 on Sonnet 5.5, he prefers Opus if you can afford it .
  • Cursor Rollouts now traces a regression to the offending PR, opens an issue, and lets you start a cloud agent to fix it in one click. Usage credits are included through Oct 3 .
  • LangSmith Engine v2 reproduces an issue in the same environment with the same inputs, then builds and tests a fix and iterates on it before handing you a PR . Managed Deep Agents now supports memory scoped to a single user or a whole team .
  • Codex has a new /experimental flag that keeps the machine from sleeping .
Huntley: let types and simulators check agent code so cheaper models can do the work
Salvatore Sanfilippo
Profile
  • Sanfilippo’s stated approach is to “check the ideas, not the code”; he says he has agents working nearly all the time during his workday. He also describes asking a model to plan a project, then delegating development to a local agent.
  • Dwarf Star currently supports three model families. Sanfilippo checks its outputs against the model provider’s API responses on 1,000 prompts, and says focusing on a few models enables optimization and keeps the codebase easier for agents to understand and modify.
  • With Dwarf Star, he can give an agent access to logged-in Chrome and ask it to retrieve his YouTube video comments via JavaScript and surface the meanest ones—a task he says a cloud model refused.
  • For choosing projects, he recommends experiments that teach something significant, or projects he expects to use for months; he favors work that can be released and may interest others. Since easier implementation removes some of the effort-based brake on scope, he stresses prioritizing important features and resisting feature creep.
IL VERO FUTURO DELL'AI? LLM LOCALI, STATI UNITI e CINA con Salvatore Sanfilippo @antirez e @enkk
Riley Brown
  • In Claude Projects, keep the main thread concise and delegate discrete tasks, such as creating a deck explaining an app, to separate threads; use local threads for work needing the computer, GitHub, Vercel, or local Claude Code, and cloud threads for other tasks—Riley says cloud sessions let him make app changes from his phone.
  • Riley’s firsthand comparison was about 30 hours using Opus 5.5 versus about 4 hours using Sonnet 5.5; he favored Opus if affordable, while recommending Sonnet when lower cost or more usage mattered. He said he had not hit limits on his $200 plan with his own workload of a couple concurrent sessions, while noting that large projects and many simultaneous sessions could hit limits.
  • In a Bluey voice-agent demo, Riley had the agent prepare email and Slack messages, then send them after his instruction; the agent caught an outdated project status and updated the messages before sending, and subsequently launched a local coding session to create three landing-page options. He described the agent as using a fast model before delegating coding work.
  • ChatGPT plugin extensions were shown controlling the design app Magic Path: Riley prompted it to create a light-mode version using his ChatGPT account.
  • On model choice, Riley said Gemini Argon was not yet publicly available for him to verify; his prior coding-related concern about Gemini was that Gemini 3 would refuse to use tools, and he also found Google’s choice of apps for using its models unclear.
ChatGPT Dots Is Insane & The Truth About Google's NEW Argon
ThePrimeTime
  • In a coding-agent demo, the presenter asked an agent to implement the next pending task in a V2 migration doc, without running tests or a review agent; the edit took about 1 minute 41 seconds and produced 722 insertions and 16 deletions. On the $500 plan's ultra-fast mode, weekly usage fell from 54% to 52%. This is a concrete raw-edit speed and usage-cost datapoint, not a test of review-ready output.
  • For agent environments, Namespace's sponsor pitch is to provide a development box with the codebase, test suite, database, and network, with egress filtering and SSH access; setup options include a Docker image, repository, operating system, or Mac.
  • The video reports planned native Codex access to GLM 5.3 Flash and Kimmy K3, with usage charged against an OpenAI commitment.
A total disaster
Geoffrey Huntley
Profile
  • Huntley builds an agent-driven source-control prototype with encrypted contents and claim-based materialized views, including subpath access controls; his validation workflow is to build a simulator first, then drive the agent against it to check its work and keep it on track.
  • For agent-maintained code, he favors Rust over Python because its types provide “back pressure” through compiler errors that the agent can fix; he says this makes Rust codebases more maintainable with agents.
  • He argues code need not be human-readable if it remains explainable, and suggests pasting an unfamiliar function into an LLM and asking for an accessible explanation.
  • He reports using GPT 55 with no or low reasoning and GLM, and recommends matching the model to the task and getting tokens as cheaply as possible; he says frontier intelligence is not needed for every task and companies should check for zero data retention.
YAAP | Your Code Doesn't Need to Be Readable Anymore
Jason Zhou
  • @treg_ai says it uses Typesafe’s Jev in production for web-traffic classification, smart onboarding, and search, and has open-sourced its implementation at GitHub. The team reports Jev is about 10x cheaper, 18x faster, and 30% more accurate than GPT6-Luna; the linked guide instead describes Luna as 30x slower, so the reported speedup differs across the team’s materials.
  • A reusable design pattern is to keep control flow in code and use the model for bounded semantic decisions; use its confidence scores to skip low-risk cases, act on high-confidence ones, and send uncertain cases to a person (the guide gives churn examples scored 0.06, 0.86, and 0.38). The guide also recommends having Claude Code read Typesafe’s advanced-primitives documentation when implementing Jev.
  • For agent tool discovery, keep keyword or hybrid search, retrieve a wider candidate set, then use Jev to rerank rather than replace search; Treg reports this improved accuracy 9x across its 3,800+ data endpoints and tools and reduced agents’ token use when finding the right tool.
  • For personalized onboarding, combine signup details with live company signals and product usage in a user profile; have Jev rank next actions, refresh the ranking as signals change, and optionally let a small LLM generate new candidate actions for Jev to sort. The guide describes one Jev call per refresh.
15min [@typesafe](https://x.com/typesafe) Jeveloper tutorial 🧠 My team at [@treg_ai](https://x.com/treg_ai) has been using Jev a lot in p… Build 'Smarter Software' with Jev (Step-by-step guide)
ThePrimeagen

ThePrimeagen tried jevlint in a project, called it very good, and found a few issues he considered worth fixing. The project author says it is still rough around the edges; it can be installed with go install github.com/codegirl-007/jevlint/cmd/jevlint@latest, but newly added Cloudflare support was not yet recommended.

this was very good. I gave it a try in the project, found a few little issues that were real worth while fixes [https://x.com/codegirl007… [@ThePrimeagen](https://x.com/ThePrimeagen) It's still a little rough around the edges. [https://github.com/codegirl-007/jevlint](https:/…
Riley Brown
Profile
  • Brown uses Claude Projects’ main agent to delegate app-design and deck work into separate threads, keeping the main conversation short. He uses cloud sessions from his phone and local sessions when work needs his computer or GitHub/Vercel access; results appear as artifacts in the project library.
  • In his own testing—about four hours with Sonnet 5.5 and 30 with Opus 5.5—Brown preferred Opus for capability and value, especially with Projects, while describing Sonnet as a more efficient option for simpler tasks or when cost matters. This is his hands-on assessment, not a benchmark.
  • In a demo, Brown asked his voice agent to launch a local Codex session for a landing-page redesign; it created three options with a toggle and site preview, and started the coding task after confirming the email and Slack sends.
ChatGPT Dots Is Insane & The Truth About Google's NEW Argon
Geoffrey Huntley
  • Huntley argues that agent-maintained Rust is more maintainable than Python because types give agents compiler-error feedback as verification (“back pressure”) on each loop; that stronger verification can let developers use less capable, cheaper models. He recommends choosing models by task and price, and cites GLM on Baseten as a lower-cost option; for company use, he advises checking for zero data retention.
  • For an agent-driven rebuild of a distributed source-control system, he is building in Rust and developing the simulator first, then using it to validate the agent’s work—a workflow he says has kept the agent on track.
  • To make unfamiliar or obfuscated code accessible, he suggests asking an LLM to explain it in beginner-friendly terms using Python as a reference; his example turns an obscure Haskell function into an explanation of its behavior.
software doesn't need to be readable anymore. it needs to be explainable.
Latent.Space
  • Airbnb reports that AI authors 60% of its code, feature and improvement shipments are nearly 80% higher year over year, and average engineer pull-request throughput is about 1.6× higher. Its teams also moved product, design, and engineering work directly into prototypes, using code and prototypes rather than excessive documents as the artifacts they reason about.
  • Airbnb’s internal agent AirChat includes MCP organizational context. Its Everest context graph uses LLMs, embeddings, and AI retrieval across the organization; learnings from the grocery-delivery project helped the airport-pickup team build a similar service in about six weeks, versus eight to nine months for groceries.
  • Airbnb is starting to run asynchronous agents in containers in response to monitoring events: agents can perform initial on-call triage and propose a PR for human review, or close an incident if an alert appears flaky.
  • Airbnb evaluates models for each use case against relevant sampled production queries, including edge cases. It favors the strongest available frontier model for coding because defects are costly, while using smaller specialized models for latency-sensitive search; narrow post-trained models can sometimes outperform frontier models.
  • To preserve engineering judgment and craft, Airbnb expects engineers to explain work done by AI, including AI-generated PRs.
Inside-Out AI: Rebuilding Airbnb Behind the Scenes and Across the Guest Experience
swyx

Swyx cites what he describes as an “exploding number” of rogue agents, more breaches, and AI-driven attacks as reasons security should be central to AI discussions.

It's time to get serious about Security x AI. One of the joys of doing AIE is helping others start their own high quality conferences in …
Latent Space

SemiAnalysis’s ClusterMax team moved from manually running benchmark scripts to a repo of industry-standard and custom benchmarks, including reliability and burn-in tests; they now use coding agents to launch tests and autonomously resolve many issues, improving evaluation coverage and depth.

Which GPU Clouds Are Actually Good? | ClusterMAX 3.0
Addy Osmani

Claude Code mods are small JavaScript/TypeScript files that run in a session and can watch events, change behavior, or draw UI. You can write one in a few lines or ask Claude to build it, then share it as a plugin and install it with /plugin in the CLI or desktop app.

Introducing Claude Code mods! Customize how it looks and works just by prompting it. [https://claude.dev/blog/getting-started-with-claude… You can now mod Claude Code: - Change how it behaves - Customize the UI - Swap in your own features Write one with a few lines of TypeScr…
Kent C. Dodds 🐨

Kent C. Dodds endorsed rewarding feature deletion . The linked argument warns that in the agentic era, low-cost feature creation can encourage feature bloat; teams should counter this by rewarding removal rather than treating added features as the only visible progress .

Reward deleting features [https://x.com/poteto/status/2106202416853262408](https://x.com/poteto/status/2106202416853262408) deleting the product it's funny how you can sometimes tell if a team culture is dysfunctional just by looking at how their app is laid ou…
Jason Zhou

Knowix’s SEO research workflow uses Claude Code with Treg’s MCP server (also compatible with Codex and Hermes) to access keyword data, Search Console, live Google results, and AI-answer citations; Treg shows the price before each call. The agent checks whether a page already exists, analyzes search results and AI citations, then recommends improve/create/skip and produces a one-page brief; the prompt caps the batch at $0.15 and says not to write the article, leaving the decision and writing to the person. The author reports a first run improved an existing page for about $0.08, and describes the agent’s role as gathering evidence while the human supplies judgment.

I Replaced My $250 SEO Tool With Treg
Latent.Space
  • Pi 1.0 adds Codemode with native MCP, Jev, and image-model support, plus virtual-model extensions, deferred tool loading, Anthropic cache warming, transcript-aware mid-conversation system-message and tool changes, and a full-screen-by-default TUI.
  • Pi Durable ports Pi to TypeScript and externalizes state: checkpointed steps let agents and subagents resume after failures, while one harness can run parallel, branching conversations. Developers can package prompts, tools, hooks, and durable tasks as extensions; background compaction maintains token limits, shared documents let multiple users or UIs steer an agent, and tool or extension code can be hot-swapped while it runs.
[AINews] Pi 1.0, Pi Durable, and AIE NYC
geoff

Geoffrey Huntley shared a recommendation to use Statecharts for agent-assisted distributed-systems work: the post argues that their local reasoning suits agents and suggests pointing an agent at SCXML; it also links an interactive demo of an agent porting Pi Durable to Statecharts.

Revised: 1986 [https://www.state-machine.com/doc/Harel87.pdf](https://www.state-machine.com/doc/Harel87.pdf) ![](https://pbs.twimg.com/me… Do yourself a favor and read about Statecharts. They take all the pain out of distributed systems, are free (yay, ideas - point your agen…
Jason Zhou

Josh Jackson says he took Gimme Gimme from idea to 450+ users in 48 hours by using the best coding agent available, relying on tools and platforms to unblock progress (he calls Treg AI a key one), then shipping before it felt ready, finding user pain, and fixing it quickly . Jason Zhou highlighted Treg AI’s support for the Grok hackathon’s #1 winner .

How did we scale Gimme Gimme from idea to 450+ users in 48 hours? Had this question a lot this week so dropped a quick walkthrough of how… Proud that [@treg_ai](https://x.com/treg_ai) support grok hackthon [#1](https://x.com/hashtag/1) winner! From idea to 450+ users in 48 hr…
geoff

OpenAI added an experimental /experimental flag in Codex that stops the computer from going to sleep.

a tale of two companies in 2026: a. Microsoft: "We are committed to reducing greenhouse gas emissions" labels in Windows Update. b. OpenA…
Riley Brown

Riley says he built Bluey with Opus 5.5 and that frontier models let him accomplish a lot in an hour; he does not specify Bluey’s build time. Bluey is an iOS and Mac computer companion that can point, comment, and control his computer. He tentatively identifies GPT Realtime 2 as what powers it, adding that Opus chose what to use, so the model attribution is uncertain. Bluey’s GitHub repository

Every day i'm blown away at what i can do in an hour with frontier models like Opus 5.5. I didn't know what to do with an old iPhone, so … [https://github.com/rbrown101010/bluey-by-riley](https://github.com/rbrown101010/bluey-by-riley)
Kent C. Dodds 🐨

After reaching alignment with an agent, Kent C. Dodds uses “Make it happen” as his go-ahead phrase.

What's your "go" phrase for agents once you've hit alignment? Mine is "Make it happen."