ZeroNoise Logo zeronoise
Post
Anthropic's Claude speed sprint: give agents a benchmark that can only improve, then fan out
•
5 min read
• 71 docs
Theo breaks down how Anthropic used Claude to make Claude.ai about 3x faster, and the benchmark guardrail pattern he reuses in T3 Code. Also: DHH on choosing Rust when agents write the code, Opus 5.5 vs GPT-6 Astra, and always-on "AI sheds".

How to run agent-driven performance work

Theo went through Anthropic's write-up of a two-week sprint that made the core Claude.ai and desktop experience about 3x faster. The sprint covered four journeys that make up 95% of user activity. At p75, a fresh load went from 3.1s to 0.55s, and starting a Claude Code session went from 0.8s to 0.3s . The method carries over to other codebases:

  1. Measure from the user's side first. Claude analyzed usage through the Datadog MCP server and picked the highest-impact journeys. Each measurement started at a user interaction, ended when the result was rendered, and separated client work from server work .
  2. Use deterministic proxies, but check them. Wall-clock time is noisy, so the team used instruction counts under Valgrind for pure JS paths. For the browser they used React commits, V8 call counts, style recalculations and DOM mutations . Every benchmark had two jobs: give Claude a number to lower in the lab, and act as a CI guardrail that could only go down. Benchmarks that were flaky or didn't track user latency were thrown out .
  3. Ratchet. After Claude cut instructions on two hot paths by 48% and 31%, any PR that raised those counts failed CI, and a daily job lowered the ceiling whenever the count dropped .
  4. Loop in a shared Slack channel. Someone opens a thread about a slow step. Claude traces it, builds a benchmark, and sends PRs sized for risk, with anything user-visible behind a flag. It then reads field data and either ratchets the benchmark or turns the flag off . About 200 flags were added over the sprint, and more than half were cleaned up by the end .

Theo's caveat: optimizing for counts alone pushes agents to cut network requests that don't matter. He suspects this caused a bug he found live, where deleted threads stayed in Claude's cached sidebar after a refresh . He still keeps performance work in the loop. He asks the agent for theories, has it build demos of each, then tests which feels faster himself .

His own version: an agent-written GitHub Action benchmarks T3 Code's request formats against stored thread data and comments on PRs when results drift from the baseline, so coding and review agents can fix regressions themselves . With that in place, he fanned out agents and 38 of 40 performance PRs auto-merged . For exploratory work he says to explicitly ask for "extreme" options with several proposals ("Don't be afraid to boil the ocean"). That prompt led him to fork Rusty V8 into a custom runtime, which he reports gave a 2–4x improvement .

DHH: if you don't read the code, reconsider the language

DHH had agents implement and optimize Basecamp's Campfire in Elixir, Go and Rust. He asks whether Ruby's slowness still matters "if you're no longer reading the code" . Agents have also finished Laravel and Django ports. He frames the exercise as a test of how frontier agents handle each environment out of the box . The Rust version is more than 10x the lines of code of Rails, which he says mattered far more when humans wrote the code . He also says agents write much faster Rust than Elixir or Go . Caveats: Rust compiles slowly, and he says most apps don't need Rust-level performance . His sharper claim is that any language agents can't write well out of the box "is going to have a hard time" . This echoes Huntley's argument from the last brief that compiler back-pressure keeps agents on the rails.

Opus 5.5 vs GPT-6 Astra: daily driver vs reviewer

Theo now says preferring OpenAI models for code "makes almost no sense." In his view Opus 5.5 is fast and surprisingly cheap, while OpenAI models are slow without fast mode and burn through a $200 plan in hours . His vibe rankings: Opus 5.5 scores 9/10 on code and 9/10 on understanding intent; Astra scores 7.5/10 and 3/10 .

A narrower pattern has support from others. Robin Ebers uses Opus day to day but has Astra run an adversarial review of plans Opus wrote. About 30% of Astra's feedback is over-engineering, and he says 70% points to real holes . Riley Brown keeps Astra for complex, high-stakes tasks and uses Opus for frontend, documents and most coding, which feels unlimited on the $200 plan .

Where agents run

DHH recommends an "AI shed": an always-on machine on your Tailscale network that runs most of your agents. Compiles happen there, so a modest laptop is enough and closing the lid doesn't stop the agents . His main shed is a Minisforum MS-A2, and he says a Beelink box with an 8745HS or similar handles most work . Theo announced t3os on the same theme: a headless, mostly Ubuntu-based OS for "personal servers," built from tech choices agents prefer, with no ISO and not meant for a computer you use . OpenAI's Tibo said in a podcast that specialist dots get extra guardrails and run on their own hardware, sometimes Mac minis, and that one dot's harness can control many devices .

Smaller items

  • Verified hot-reload: Geoffrey Huntley had an OpenAI model inspect a running SBCL process and add reverse-string through a tool call. The change was saved only after it passed 21 fixed cases and 1,000 generated ones. A fresh SBCL process passed the same checks, and rolling back and restoring the revision both worked .
  • Codex: Tibo says that for the next 28 days, Codex/Work will ship one clear improvement relevant to most users each day, or do a full usage reset .
  • Agent count: Tibo now runs fewer agents in parallel because ultrafast mode lets him stay in flow. He grows his team of agents on frontier work and shrinks it again after model jumps .
  • Durability: Armin Ronacher says making an agent "somewhat durable" is easy, but a harness that is durable by design without becoming "a crazy mess" took many iterations .
  • Outputs as artifacts: Building on Karpathy's writing tip, Omar Sar has his agent produce artifacts suited to the task, on Notion-like pages, instead of plain text. He reviews them through comments and suggestions, including for code review .
Anthropic's Claude speed sprint: give agents a benchmark that can only improve, then fan out
Summary
Coverage start
1 day ago
Coverage end
17 hours ago
Frequency
Daily
Published
16 hours ago
Reading time
5 min
Research time
2 hrs 1 min
Documents scanned
71
Documents used
21
Citations
35
Sources monitored
111 / 111
Insights
Skipped contexts
Source details
Source Docs Insights Status
LangChain Blog 0 0
Brent Traut 0 0
Lukas Möller 0 0
Jediah Katz 0 0
Aman Karmani 0 0
Jacob Jackson 0 0
Cursor Blog | RSS Feed 0 0
Nicholas Moy 0 0
Mike Krieger 0 0
Sualeh Asif 0 0
Michael Truell 0 0
Google Antigravity 0 0
Aman Sanger 0 0
cat 0 0
Mark Chen 0 0
Greg Brockman 0 0
Tongzhou Wang 0 0
fouad 0 0
Calvin French-Owen 0 0
Hanson Wang 4 0
Ed Bayes 0 0
Alexander Embiricos 0 0
Tibo 2 1
Romain Huet 2 1
DHH 17 6
Jane Street Blog 0 0
Miguel Grinberg's Blog: AI 0 0
xxchan's Blog 0 0
<antirez> 0 0
Brendan Long 0 0
The Pragmatic Engineer 0 0
David Heinemeier Hansson 0 0
Armin Ronacher ⇌ 5 2
Mitchell Hashimoto 0 0
Armin Ronacher's Thoughts and Writings 0 0
Peter Steinberger 0 0
Theo - t3.gg 12 3
Sourcegraph 0 0
Anthropic 0 0
Cursor 0 0
LangChain 0 0
Anthropic 0 0
LangChain 0 0
Cursor 0 0
Riley Brown 0 0
Riley Brown 4 3
Jason Zhou 0 0
Boris Cherny 0 0
Mckay Wrigley 0 0
geoff 9 2
Peter Steinberger 🦞 2 0
AI Jason 0 0
Alex Albert 0 0
Latent.Space 0 0
Logan Kilpatrick 0 0
Fireship 0 0
Fireship 0 0
Kent C. Dodds 🐨 4 0
Practical AI 0 0
Practical AI Clips 0 0
Stories by Steve Yegge on Medium 0 0
Kent C. Dodds Blog 0 0
ThePrimeTime 0 0
Theo - t3․gg 1 1
ThePrimeagen 3 0
Ben Tossell 1 0
swyx 0 0
AI For Developers 0 0
Geoffrey Huntley 0 0
Addy Osmani 0 0
Andrej Karpathy 3 1
Simon Willison 0 0
Matthew Berman 0 0
Changelog 0 0
Simon Willison’s Newsletter 0 0
Agentic Coding Newsletter 0 0
Latent Space 0 0
Simon Willison's Weblog 0 0
Elevate 0 0
Lukas Möller 0 0
Jediah Katz 0 0
Sualeh Asif 0 0
Mike Krieger 0 0
Michael Truell 0 0
Cat Wu 0 0
Kevin Hou 0 0
Aman Sanger 0 0
Nicholas Moy 0 0
Andrey Mishchenko 0 0
Jerry Tworek 0 0
Romain Huet 0 0
Thibault Sottiaux 1 1
Alexander Embiricos 0 0
xxchan 0 0
Salvatore Sanfilippo 1 0
Armin Ronacher 0 0
David Heinemeier Hansson (DHH) 0 0
Alex Albert 0 0
Logan Kilpatrick 0 0
Shawn "swyx" Wang 0 0
Jason Zhou 0 0
Riley Brown 0 0
McKay Wrigley 0 0
Boris Cherny 0 0
Ben Tossell 0 0
Geoffrey Huntley 0 0
Peter Steinberger 0 0
Addy Osmani 0 0
Simon Willison 0 0
Andrej Karpathy 0 0
Harrison Chase 0 0