ZeroNoise Logo zeronoise
Post
CI Is Now the Bottleneck for Agent-Heavy Teams: Codex-Picked Tests vs. Nix Remote Builders
•
6 min read
• 92 docs
Practitioners say CI now slows agent-driven teams down, and Steinberger and Huntley propose two different fixes. Also covered: the math behind Opus 5.5's roomier limits, Antigravity's /teamwork and sidecars, and DHH and Matz on treating agent PRs as suggestions.

🔥 TOP SIGNAL

CI, not inference, is becoming the thing that slows agent-heavy teams down. Altimor (Lindy) says CI is "the top bottleneck of every engineering team I talk to" and calls Lindy's own CI spend "stratospheric" . Two responses are worth comparing. Peter Steinberger plans to "let codex decide which tests actually need to run and drastically nix CI and run tests hourly" . Geoffrey Huntley's answer is nix-build on big Hetzner bare-metal boxes set up as trusted remote builders for your agent, with traditional CI kept only for complex integration and pre-release regression tests. He points to Hydra as the state of the art . Neither has published results yet: Steinberger's is a plan and Huntley's is a recommendation.

⚡ TRY THIS

  • Stop being the "meat proxy" (DHH, Rails World panel). If you're copying error messages between agents, "consider just letting the agent read the error messages directly" . He treats agent PRs and issues as suggestions: if he likes one, "my agent will pick the slop apart and return with a beautiful implementation" . He says agent-opened PRs and issues are better than the average human ones on any repo he has managed . On how much autonomy to give, his answer is "as much as their competence warrants." Measure that by whether the last PR was handled correctly, and check in less as trust builds . Matz's version: 20 issues and PRs came in within an hour before the keynote. He asked the Claude app on his phone to address them, and by the end of the talk it had resolved 12 .

  • Make your internal tools agent-native (Riley Brown). He listed the features he actually used across Notion, My Mind, Typefully and Google Drive, then asked Claude for "an app that does all those things that looks really good." It took about 20 prompts, and he tested it with several video-editor agencies . The useful part: a Claude Code skill called "native note" writes directly to the app's database. "Add it to native note" creates a linked script thread, so the agent uses the same software the team does .

  • Property-fuzz your agent's output all the way down to I/O (Huntley). He "vibed up" a setup that uses bombadil to autonomously explore web and terminal apps with property-based testing, asserting properties on side effects in databases, the filesystem and systemd units. It's packaged as deterministic NixOS machine tests that take a seed, so you can fan verification out across a fleet with Hydra .

  • Use a constrained machine as the performance target (DHH). Keep a ten-year-old "potato PC," put your agents to work inside it, "and watch the weight drop" .

📡 WHAT SHIPPED

  • Why Opus 5.5's limits feel much bigger (Theo's math). Claude Code subscriptions allow Fable to use only 50% of the usage, and Opus 5.5 High costs less than half as much as Fable 5.1 High. Theo puts Fable 5.1 High → Opus 5.5 High at about 4.3× more usable limits, and his own switch from Fable 5.1 xhigh at about 6.6× . Addy Osmani's official-side numbers: limits go about 25% further than on Opus 5, cache reads cost 60% less, and output is 30% faster (cost post) . Riley Brown hasn't hit a limit in five days. He grants Astra may be better "at a good amount of things," but says it's slower and much more expensive .

  • Antigravity 2.0 agent teams (Kevin Hou, Google DeepMind). The agent manager is now a separate app from the IDE . /teamwork (public preview) starts a lead agent that asks clarifying questions, then creates specialized subagents. The lead can pick a different model for each subagent . Hou says Gemini 3.5 Flash is "really good at leading teams," and faster and cheaper . Sidecars are long-running plugin processes that listen for SMS, webhooks, cron or GitHub PRs. Hou said the spec would be released "later this summer" . Their showcase run built an OS kernel that runs Doom: 93 subagents, 12 hours, 15,000 requests, 2B tokens, under $1,000. Hou says it's not something they'd do every day .

  • Computer use: people are preferring Codex. @jtaby: "insane how much worse computer use is in claude vs codex" . Riley Brown: "Codex computer use is absurd. Especially on Mac" . These are impressions, with no tasks or benchmarks given.

  • jevgrep (repo). A research-agent CLI from @dzhng built on typesafe's jev. It claims a 40% cut in coding-agent cost, "verified on SWE-bench." Install its built-in skill so your agent uses jg for context collection . The claim comes from the tool's author and hasn't been independently checked.

  • ttfx sequel: people rewrite the agent's code too. Daniel Collin has a WIP Rust rewrite of Opus's assembly port with AVX-512/AVX2/SSE2/Neon paths. He reports it's about 10.7× faster than the original Rust and 1.25× faster than the asm on average (PR #44) . DHH: "I have zero allegiance to whatever code my agents wrote" .

  • Security caution. @_mattata's project to audit every line of the top 1,000 GitHub repos in under 30 days for $100 is "nearing completion" . Huntley's conclusion is that with LLM red-teaming, putting source on GitHub now reduces its security . No findings have been published.

  • A personal-agent failure worth studying. Simon Willison quoted a Muse agent admitting that its auto-reply told a buyer "Yep I'm here!" when its user wasn't there. The buyer left a negative rating, and the agent then offered to stop making claims it can't verify . Lesson: don't let an agent assert real-world state it can't check.

🎬 GO DEEPER

  • Kevin Hou on automating side-by-side evals with subagents. Researchers load a skills file and ask about an eval in plain language. The agent computes the control/experiment delta, a research subagent proposes about 100 hypotheses, one subagent investigates each in parallel, and the result is a report with a generated UI. Hou says 90% of the workflow is automated and it takes minutes instead of Jupyter "elbow grease" .
  • DHH and Matz on agent PRs, "meat proxies," and calibrating trust. A short stretch covering the whole "suggestions, not implementations" stance and Matz's 12 issues fixed from his phone.
  • Riley Brown's agent-native internal app demo. The skill that writes to the app's database is the part to copy.
  • Huntley, "the eighteen-month recap". Agents are "just a while loop, and Ralph is a while loop on a while loop" . His free 300-line workshop has you build your own Cursor/Copilot/Codex . Separately, he says he doesn't pre-install skills or cram markdown into context: "Stop chasing LLM prompt fairies and instead focus on the craft" .

Editorial take: Opus 5.5's cheaper limits make generating code close to free, so the costs now sit in CI, test selection, and deciding how far to trust agent output.

CI Is Now the Bottleneck for Agent-Heavy Teams: Codex-Picked Tests vs. Nix Remote Builders
Summary
Coverage start
2 days ago
Coverage end
1 day ago
Frequency
Daily
Published
1 day ago
Reading time
6 min
Research time
3 hrs 13 min
Documents scanned
92
Documents used
20
Citations
37
Sources monitored
111 / 111
Insights
Skipped contexts
Source details
Source Docs Insights Status
LangChain Blog 0 0
Brent Traut 0 0
Lukas Möller 0 0
Jediah Katz 0 0
Aman Karmani 2 0
Jacob Jackson 0 0
Cursor Blog | RSS Feed 0 0
Nicholas Moy 0 0
Mike Krieger 0 0
Sualeh Asif 0 0
Michael Truell 0 0
Google Antigravity 0 0
Aman Sanger 0 0
cat 0 0
Mark Chen 0 0
Greg Brockman 0 0
Tongzhou Wang 0 0
fouad 0 0
Calvin French-Owen 0 0
Hanson Wang 0 0
Ed Bayes 0 0
Alexander Embiricos 0 0
Tibo 2 0
Romain Huet 0 0
DHH 20 3
Jane Street Blog 0 0
Miguel Grinberg's Blog: AI 0 0
xxchan's Blog 0 0
<antirez> 0 0
Brendan Long 0 0
The Pragmatic Engineer 0 0
David Heinemeier Hansson 0 0
Armin Ronacher ⇌ 3 0
Mitchell Hashimoto 0 0
Armin Ronacher's Thoughts and Writings 0 0
Peter Steinberger 0 0
Theo - t3.gg 15 4
Sourcegraph 0 0
Anthropic 0 0
Cursor 0 0
LangChain 0 0
Anthropic 0 0
LangChain 0 0
Cursor 0 0
Riley Brown 1 1
Riley Brown 7 1
Jason Zhou 0 0
Boris Cherny 0 0
Mckay Wrigley 0 0
geoff 13 5
Peter Steinberger 🦞 2 1
AI Jason 0 0
Alex Albert 0 0
Latent.Space 0 0
Logan Kilpatrick 1 0
Fireship 0 0
Fireship 0 0
Kent C. Dodds 🐨 10 1
Practical AI 0 0
Practical AI Clips 0 0
Stories by Steve Yegge on Medium 0 0
Kent C. Dodds Blog 0 0
ThePrimeTime 1 1
Theo - t3․gg 1 1
ThePrimeagen 6 1
Ben Tossell 0 0
swyx 0 0
AI For Developers 0 0
Geoffrey Huntley 1 1
Addy Osmani 2 1
Andrej Karpathy 0 0
Simon Willison 0 0
Matthew Berman 0 0
Changelog 0 0
Simon Willison’s Newsletter 0 0
Agentic Coding Newsletter 0 0
Latent Space 0 0
Simon Willison's Weblog 1 0
Elevate 0 0
Lukas Möller 0 0
Jediah Katz 0 0
Sualeh Asif 0 0
Mike Krieger 0 0
Michael Truell 0 0
Cat Wu 0 0
Kevin Hou 1 1
Aman Sanger 0 0
Nicholas Moy 0 0
Andrey Mishchenko 0 0
Jerry Tworek 0 0
Romain Huet 0 0
Thibault Sottiaux 0 0
Alexander Embiricos 0 0
xxchan 0 0
Salvatore Sanfilippo 1 0
Armin Ronacher 0 0
David Heinemeier Hansson (DHH) 1 1
Alex Albert 0 0
Logan Kilpatrick 0 0
Shawn "swyx" Wang 0 0
Jason Zhou 0 0
Riley Brown 1 1
McKay Wrigley 0 0
Boris Cherny 0 0
Ben Tossell 0 0
Geoffrey Huntley 0 0
Peter Steinberger 0 0
Addy Osmani 0 0
Simon Willison 0 0
Andrej Karpathy 0 0
Harrison Chase 0 0