ZeroNoise Logo zeronoise
Post
Anthropic's Claude speed sprint: give agents a benchmark that can only improve, then fan out
•
5 min read
• 71 docs
Theo breaks down how Anthropic used Claude to make Claude.ai about 3x faster, and the benchmark guardrail pattern he reuses in T3 Code. Also: DHH on choosing Rust when agents write the code, Opus 5.5 vs GPT-6 Astra, and always-on "AI sheds".

How to run agent-driven performance work

Theo went through Anthropic's write-up of a two-week sprint that made the core Claude.ai and desktop experience about 3x faster. The sprint covered four journeys that make up 95% of user activity. At p75, a fresh load went from 3.1s to 0.55s, and starting a Claude Code session went from 0.8s to 0.3s . The method carries over to other codebases:

  1. Measure from the user's side first. Claude analyzed usage through the Datadog MCP server and picked the highest-impact journeys. Each measurement started at a user interaction, ended when the result was rendered, and separated client work from server work .
  2. Use deterministic proxies, but check them. Wall-clock time is noisy, so the team used instruction counts under Valgrind for pure JS paths. For the browser they used React commits, V8 call counts, style recalculations and DOM mutations . Every benchmark had two jobs: give Claude a number to lower in the lab, and act as a CI guardrail that could only go down. Benchmarks that were flaky or didn't track user latency were thrown out .
  3. Ratchet. After Claude cut instructions on two hot paths by 48% and 31%, any PR that raised those counts failed CI, and a daily job lowered the ceiling whenever the count dropped .
  4. Loop in a shared Slack channel. Someone opens a thread about a slow step. Claude traces it, builds a benchmark, and sends PRs sized for risk, with anything user-visible behind a flag. It then reads field data and either ratchets the benchmark or turns the flag off . About 200 flags were added over the sprint, and more than half were cleaned up by the end .

Theo's caveat: optimizing for counts alone pushes agents to cut network requests that don't matter. He suspects this caused a bug he found live, where deleted threads stayed in Claude's cached sidebar after a refresh . He still keeps performance work in the loop. He asks the agent for theories, has it build demos of each, then tests which feels faster himself .

His own version: an agent-written GitHub Action benchmarks T3 Code's request formats against stored thread data and comments on PRs when results drift from the baseline, so coding and review agents can fix regressions themselves . With that in place, he fanned out agents and 38 of 40 performance PRs auto-merged . For exploratory work he says to explicitly ask for "extreme" options with several proposals ("Don't be afraid to boil the ocean"). That prompt led him to fork Rusty V8 into a custom runtime, which he reports gave a 2–4x improvement .

DHH: if you don't read the code, reconsider the language

DHH had agents implement and optimize Basecamp's Campfire in Elixir, Go and Rust. He asks whether Ruby's slowness still matters "if you're no longer reading the code" . Agents have also finished Laravel and Django ports. He frames the exercise as a test of how frontier agents handle each environment out of the box . The Rust version is more than 10x the lines of code of Rails, which he says mattered far more when humans wrote the code . He also says agents write much faster Rust than Elixir or Go . Caveats: Rust compiles slowly, and he says most apps don't need Rust-level performance . His sharper claim is that any language agents can't write well out of the box "is going to have a hard time" . This echoes Huntley's argument from the last brief that compiler back-pressure keeps agents on the rails.

Opus 5.5 vs GPT-6 Astra: daily driver vs reviewer

Theo now says preferring OpenAI models for code "makes almost no sense." In his view Opus 5.5 is fast and surprisingly cheap, while OpenAI models are slow without fast mode and burn through a $200 plan in hours . His vibe rankings: Opus 5.5 scores 9/10 on code and 9/10 on understanding intent; Astra scores 7.5/10 and 3/10 .

A narrower pattern has support from others. Robin Ebers uses Opus day to day but has Astra run an adversarial review of plans Opus wrote. About 30% of Astra's feedback is over-engineering, and he says 70% points to real holes . Riley Brown keeps Astra for complex, high-stakes tasks and uses Opus for frontend, documents and most coding, which feels unlimited on the $200 plan .

Where agents run

DHH recommends an "AI shed": an always-on machine on your Tailscale network that runs most of your agents. Compiles happen there, so a modest laptop is enough and closing the lid doesn't stop the agents . His main shed is a Minisforum MS-A2, and he says a Beelink box with an 8745HS or similar handles most work . Theo announced t3os on the same theme: a headless, mostly Ubuntu-based OS for "personal servers," built from tech choices agents prefer, with no ISO and not meant for a computer you use . OpenAI's Tibo said in a podcast that specialist dots get extra guardrails and run on their own hardware, sometimes Mac minis, and that one dot's harness can control many devices .

Smaller items

  • Verified hot-reload: Geoffrey Huntley had an OpenAI model inspect a running SBCL process and add reverse-string through a tool call. The change was saved only after it passed 21 fixed cases and 1,000 generated ones. A fresh SBCL process passed the same checks, and rolling back and restoring the revision both worked .
  • Codex: Tibo says that for the next 28 days, Codex/Work will ship one clear improvement relevant to most users each day, or do a full usage reset .
  • Agent count: Tibo now runs fewer agents in parallel because ultrafast mode lets him stay in flow. He grows his team of agents on frontier work and shrinks it again after model jumps .
  • Durability: Armin Ronacher says making an agent "somewhat durable" is easy, but a harness that is durable by design without becoming "a crazy mess" took many iterations .
  • Outputs as artifacts: Building on Karpathy's writing tip, Omar Sar has his agent produce artifacts suited to the task, on Notion-like pages, instead of plain text. He reviews them through comments and suggestions, including for code review .
Anthropic's Claude speed sprint: give agents a benchmark that can only improve, then fan out
Theo - t3․gg
  • Anthropic reported making core Claude.ai and desktop journeys about 3× faster in a two-week sprint, targeting four flows that accounted for 95% of activity. Its team used Claude in a shared Slack channel and Datadog MCP to identify and work through bottlenecks: trace a focused user journey, benchmark it, submit risk-sized PRs with user-visible changes behind flags, then check deploy data and ratchet benchmarks when changes won.
  • Theo’s measurement lesson: define timings from a user interaction through the rendered result, separating client and server work; validate that deterministic proxies such as instruction counts or React/layout/DOM activity track wall-clock improvement. Treat a useful benchmark as both an optimization target and a ratcheting CI guardrail, and discard flaky or non-correlating metrics.
  • In T3 Code, Theo used an agent to build a GitHub Action that benchmarks existing request formats against stored data for a representative thread, compares results to a baseline, and comments on PRs; he says this caught regressions and gave coding and review agents a chance to fix them. After establishing tests, he fanned out performance work and reports that 38 of 40 PRs auto-merged; one animation change was disliked and another broke the marketing-site ticker. He cautions against optimizing proxy counts at the expense of real user flows: while inspecting Claude, he found deleted threads lingering in its cached sidebar after refresh.
  • For exploratory work, Theo recommends explicitly inviting ambitious or extreme options. He says a prompt asking for deep, even extreme performance improvements and multiple proposals led to a persistent V8 execution service/custom runtime built by forking Rusty V8, with a reported 2–4× improvement after substantial agent iteration and performance/behavior verification.
How Anthropic made Claude 3x faster
Thibault Sottiaux
Profile
  • Codex writes most of Sottiaux’s code for analysis—such as understanding trends, the business, upcoming features, and how previous launches are doing. He merges code occasionally; his only hands-on coding is occasional weekend LeetCode.
  • He varies agent parallelism with model capability: faster agents have helped him regain flow, while frontier work leads him to build larger agent teams; after model breakthroughs, he can consolidate work into a larger agent that handles more while keeping context in memory and learning. He also says voice control and dictation with faster agents help restore a creative flow, though it differs from the flow of hands-on coding.
  • Rather than manually tuning fixed loops, he favors an agent that works continuously, understands goals and preferences, and learns from feedback.
  • A production-monitoring example: his agent alerted him that ChatGPT production was down five minutes before a live demo, connected the outage to the imminent event, and offered to try a fix; he declined and contacted engineering. He says agents can access some production systems with guardrails; specialist agents get additional guardrails and monitoring and run on their own hardware, sometimes Mac minis, while the harness can connect to multiple devices.
OpenAI’s Head of ChatGPT: We’re entering a new era of AI (again) | Tibo Sottiaux
Andrej Karpathy
  • Karpathy recommends choosing output formats that make model work easier to understand: request ASD-STE100 explanations (or “80%” of the way to it), diagrams, interactive HTML, or bespoke explainer videos—such as a 3b1b-style video with ElevenLabs narration. He expects more work to shift toward oversight and understanding, with custom, disposable software artifacts becoming more practical as code and intelligence become abundant.
  • Omar Sar describes a flexible, Notion-like workspace for human-agent collaboration, embedding artifacts, text, videos, and visual explainers; he asks agents to build missing features and produce task-appropriate artifacts instead of default text. He uses comments and suggestions for code review and other work, and says the setup helps him move, consume, and review faster.
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks: Writing. Something … 50K likes on this!? Most of what Karpthy has shared is things we have been doing for over a year. No surprises, as many replies are highl…
geoff

Geoffrey Huntley’s live SBCL experiment had OpenAI introspect a running Lisp process and add reverse-string via a tool call; the caller verified 21 fixed cases and 1,000 generated cases, then a fresh SBCL process passed the same checks. Rollback removed the function while preserving the demo, and restoring the successful revision brought it back; this demonstrates the described LLM-driven kernel approach of self-modifying and hot-reloading only after properties pass.

The live experiment succeeded. OpenAI introspected SBCL, added \`reverse-string\` through a tool call whilst SBCL was running, the caller… behold, an evolutionary LISP kernel which uses llms to self-modify/hot-reload running application without downtime once all properties pa…
Riley Brown

Riley Brown said he spent the weekend building, plans to open-source all @agentnative_ projects, and will publish five videos the following week.

Spent the entire weekend building. Will be open sourcing all [@agentnative_](https://x.com/agentnative_) projects. 5 videos next week as …
Theo - t3.gg

Theo’s comparison shifted from July 2026, when he saw Anthropic’s coding edge as slight and OpenAI as faster, cheaper, more pleasant to use, and less restrictive on a $200 subscription, to September, when he judged Opus 5.5 substantially better, fast, surprisingly cheap, and pleasant to read, while describing OpenAI as less capable and reliable, slow without fast mode, and easy to exhaust even on paid plans. In his subsequent subjective “vibe rankings,” Opus 5.5 scored 9/10 for coding, 8.5 for prose, 8 for price, and 9 for understanding intent; GPT-6 Astra scored 7.5, 8, 3.5, and 3 respectively.

July 2026: Anthropic has the best code models. Gap isn’t very big though. They’re slow, expensive, and the “claudeisms” are at an all tim… More simply put as some vibe rankings Fable 5: - code capability: 7/10 - prose/readability: 6/10 - price: 2/10 - understanding intent: 8/…
Riley Brown
  • Riley recommends matching models to task stakes: use Astra for high-complexity, high-importance work, and Opus for most other tasks—especially frontend, documents, and routine coding; Opus feels unlimited on the $200/month plan.
  • Riley agrees with using Astra for an adversarial review of an Opus-generated plan: the quoted practitioner says Astra’s feedback can over-engineer, but often surfaces genuine gaps and improvements.
I agree - when a task is of high complexity and importance I’ll use Astra. But for most other stuff esp frontent / documents and most cod… serious hot take: astra > opus hear me out before you start crying in my comments i'd never pick astra over opus for daily work, but ever…
DHH
  • DHH says he had agents implement and optimize the Campfire web app in Elixir, Go, and Rust. Although Ruby was slowest, he had previously valued its productivity and developer-joy payoff; he frames whether developers still read the code as part of reassessing that tradeoff.
  • He argues that agents change the economics of choosing a language: faster languages without framework overhead are more verbose, and the Rust version had more than 10× the lines of code of the Rails version; verbosity mattered more when humans had to write the code.
I've had agents implement and optimize the Campfire web app in Elixir, Go, and Rust. Yes, Ruby is the slowest. I didn't care about that w… But of course faster languages without any framework overhead are also vastly more verbose! Rust is more than 10x the LOC of the Rails ve…
Theo - t3.gg

Theo announced t3os, an operating-system project he describes as built around technical choices he dislikes but agents prefer. He says it will be headless, Ubuntu-based, intended for personal servers, and have no ISO; he warns against installing it on a computer you use.

My team failed to talk me out of this. t3os is happening. It will be the worst OS ever and I'm hyped for it. [https://x.com/shivamhwp/sta… I'll give some hints: - t3os is not for humans - t3os is entirely composed of tech decisions I hate but agents prefer - t3os is headless …
geoff

Geoffrey Huntley described an evolutionary LISP kernel that uses LLMs to self-modify and hot-reload a running application without downtime, once all properties pass.

behold, an evolutionary LISP kernel which uses llms to self-modify/hot-reload running application without downtime once all properties pa…
Riley Brown

Riley Brown asked when Anthropic would publicly roll out Claude Code Projects, indicating he and people contacting him did not know how to get access; he said he had made two videos about it and received hundreds of access questions.

Can anyone from [@Anthropic](https://x.com/Anthropic) plz tell me when the new Claude Code Projects will be rolled out publicly. I’ve mad…
Armin Ronacher ⇌

A refresh-handling change landed on main as a possible fix for an unclear issue; sharing auth.json across multiple computers was raised as a possible cause, not a confirmed diagnosis.

[@CodeAkram](https://x.com/CodeAkram) [@pidotdev](https://x.com/pidotdev) [@badlogicgames](https://x.com/badlogicgames) I cannot say for …
DHH

DHH recommends an always-on “AI shed” on the Tailscale network to run most coding agents; because compilation happens there, he says a less powerful laptop is sufficient and its lid can be closed without interrupting the agents. His main shed is an MS-A2, though he says a Beelink box with an 8745HS or similar will work well for most tasks.

Every developer needs an AI shed: An always-on machine running on their tailscale network where the majority of their herdr agents are ru… Get a good shed, and you don't even need your laptop to be all that powerful. I've been running on the Wildcat XPS13 all week, and it's b… My main shed is a [@Hi_MINISFORUM](https://x.com/Hi_MINISFORUM) MS-A2, but you don't even have to spend that much. A [@Beelinkofficial](h…
DHH

DHH says he had agents implement and optimize Campfire in Elixir, Go, and Rust; Ruby was slowest, but he had accepted that tradeoff for productivity and developer joy. He argues that if agent acceleration means developers no longer read every generated line, Rust’s performance becomes a stronger reason to choose it—and languages frontier agents cannot write well out of the box may be disadvantaged.

I've had agents implement and optimize the Campfire web app in Elixir, Go, and Rust. Yes, Ruby is the slowest. I didn't care about that w… Again, I'm not saying that all apps need Rust levels of performance. Far, far from it! But I am saying that if agents are really good at … If your answer here is "frontier agents aren't good at writing Elixir or Go", then it only makes the argument worse. Any programming lang…
DHH

DHH says frontier agents’ coding speed is prompting him to reconsider Rust: he contrasts Rust’s runtime-performance advantage over Ruby with agents writing Rust “much faster” than Elixir or Go, and says his interest in agent-led Rust is scientific despite disliking writing Rust himself .

It's all fun and games when Rust moggs Ruby on performance, but show that frontier agents also write much faster Rust than Elixir or Go!? 😡🤬 Also hilarious if anyone thinks I've become a big fan of agent-led Rust out of any other consideration than purely scientific inquiry. Do…
DHH

DHH had coding agents implement and optimize the Campfire web app in Elixir, Go, and Rust; although Ruby was slower, he had accepted that tradeoff for productivity and developer joy. He now questions that choice if developers no longer read the generated code, arguing that agent-written and agent-validated code changes the tradeoffs developers must consider. Rust still carries a practical cost: slow compile times, and speed is not the only factor in choosing a codebase.

I've had agents implement and optimize the Campfire web app in Elixir, Go, and Rust. Yes, Ruby is the slowest. I didn't care about that w… Speed is not the only consideration for a code base. Rust has other drawbacks, like slow compile times. But you really have to bury your …
Romain Huet

Codex/Work is entering a 28-day shipping push, with a stated goal of delivering each day either a clear improvement relevant to most users or “a full reset”; Romain Huet framed DevDay as day one of the run.

Over the next 28 days, each day we’ll either ship one thing that is a clear improvement and relevant for most codex/work users or ship a … DevDay was just day one. 28 days of shipping ahead! 🚢 [https://x.com/thsottiaux/status/2106845241357824205](https://x.com/thsottiaux/stat…
Tibo

Tibo announced a 28-day push in which each day would bring either a broadly relevant improvement for Codex/Work users or a “full reset.” The stated focus was simplifying the products, improving efficiency to support more usage, and delivering groundbreaking features or new models, in response to user feedback favoring simplicity.

Over the next 28 days, each day we’ll either ship one thing that is a clear improvement and relevant for most codex/work users or ship a … All right, we’re locking in. Only things being worked on are simplifications, more efficiency for more usage, groundbreaking features or …
DHH

Frontier agents completed ports of Campfire to Laravel and Django; DHH framed the work as a way to gauge how agents perform across different environments out of the box, while noting the ports may still have optimization opportunities .

The agents finished porting Campfire to Laravel and Django too. I'm sure there are optimizations that could be found there too, but the w…