ZeroNoise Logo zeronoise
Post
CI Is Now the Bottleneck for Agent-Heavy Teams: Codex-Picked Tests vs. Nix Remote Builders
•
6 min read
• 92 docs
Practitioners say CI now slows agent-driven teams down, and Steinberger and Huntley propose two different fixes. Also covered: the math behind Opus 5.5's roomier limits, Antigravity's /teamwork and sidecars, and DHH and Matz on treating agent PRs as suggestions.

🔥 TOP SIGNAL

CI, not inference, is becoming the thing that slows agent-heavy teams down. Altimor (Lindy) says CI is "the top bottleneck of every engineering team I talk to" and calls Lindy's own CI spend "stratospheric" . Two responses are worth comparing. Peter Steinberger plans to "let codex decide which tests actually need to run and drastically nix CI and run tests hourly" . Geoffrey Huntley's answer is nix-build on big Hetzner bare-metal boxes set up as trusted remote builders for your agent, with traditional CI kept only for complex integration and pre-release regression tests. He points to Hydra as the state of the art . Neither has published results yet: Steinberger's is a plan and Huntley's is a recommendation.

⚡ TRY THIS

  • Stop being the "meat proxy" (DHH, Rails World panel). If you're copying error messages between agents, "consider just letting the agent read the error messages directly" . He treats agent PRs and issues as suggestions: if he likes one, "my agent will pick the slop apart and return with a beautiful implementation" . He says agent-opened PRs and issues are better than the average human ones on any repo he has managed . On how much autonomy to give, his answer is "as much as their competence warrants." Measure that by whether the last PR was handled correctly, and check in less as trust builds . Matz's version: 20 issues and PRs came in within an hour before the keynote. He asked the Claude app on his phone to address them, and by the end of the talk it had resolved 12 .

  • Make your internal tools agent-native (Riley Brown). He listed the features he actually used across Notion, My Mind, Typefully and Google Drive, then asked Claude for "an app that does all those things that looks really good." It took about 20 prompts, and he tested it with several video-editor agencies . The useful part: a Claude Code skill called "native note" writes directly to the app's database. "Add it to native note" creates a linked script thread, so the agent uses the same software the team does .

  • Property-fuzz your agent's output all the way down to I/O (Huntley). He "vibed up" a setup that uses bombadil to autonomously explore web and terminal apps with property-based testing, asserting properties on side effects in databases, the filesystem and systemd units. It's packaged as deterministic NixOS machine tests that take a seed, so you can fan verification out across a fleet with Hydra .

  • Use a constrained machine as the performance target (DHH). Keep a ten-year-old "potato PC," put your agents to work inside it, "and watch the weight drop" .

📡 WHAT SHIPPED

  • Why Opus 5.5's limits feel much bigger (Theo's math). Claude Code subscriptions allow Fable to use only 50% of the usage, and Opus 5.5 High costs less than half as much as Fable 5.1 High. Theo puts Fable 5.1 High → Opus 5.5 High at about 4.3× more usable limits, and his own switch from Fable 5.1 xhigh at about 6.6× . Addy Osmani's official-side numbers: limits go about 25% further than on Opus 5, cache reads cost 60% less, and output is 30% faster (cost post) . Riley Brown hasn't hit a limit in five days. He grants Astra may be better "at a good amount of things," but says it's slower and much more expensive .

  • Antigravity 2.0 agent teams (Kevin Hou, Google DeepMind). The agent manager is now a separate app from the IDE . /teamwork (public preview) starts a lead agent that asks clarifying questions, then creates specialized subagents. The lead can pick a different model for each subagent . Hou says Gemini 3.5 Flash is "really good at leading teams," and faster and cheaper . Sidecars are long-running plugin processes that listen for SMS, webhooks, cron or GitHub PRs. Hou said the spec would be released "later this summer" . Their showcase run built an OS kernel that runs Doom: 93 subagents, 12 hours, 15,000 requests, 2B tokens, under $1,000. Hou says it's not something they'd do every day .

  • Computer use: people are preferring Codex. @jtaby: "insane how much worse computer use is in claude vs codex" . Riley Brown: "Codex computer use is absurd. Especially on Mac" . These are impressions, with no tasks or benchmarks given.

  • jevgrep (repo). A research-agent CLI from @dzhng built on typesafe's jev. It claims a 40% cut in coding-agent cost, "verified on SWE-bench." Install its built-in skill so your agent uses jg for context collection . The claim comes from the tool's author and hasn't been independently checked.

  • ttfx sequel: people rewrite the agent's code too. Daniel Collin has a WIP Rust rewrite of Opus's assembly port with AVX-512/AVX2/SSE2/Neon paths. He reports it's about 10.7× faster than the original Rust and 1.25× faster than the asm on average (PR #44) . DHH: "I have zero allegiance to whatever code my agents wrote" .

  • Security caution. @_mattata's project to audit every line of the top 1,000 GitHub repos in under 30 days for $100 is "nearing completion" . Huntley's conclusion is that with LLM red-teaming, putting source on GitHub now reduces its security . No findings have been published.

  • A personal-agent failure worth studying. Simon Willison quoted a Muse agent admitting that its auto-reply told a buyer "Yep I'm here!" when its user wasn't there. The buyer left a negative rating, and the agent then offered to stop making claims it can't verify . Lesson: don't let an agent assert real-world state it can't check.

🎬 GO DEEPER

  • Kevin Hou on automating side-by-side evals with subagents. Researchers load a skills file and ask about an eval in plain language. The agent computes the control/experiment delta, a research subagent proposes about 100 hypotheses, one subagent investigates each in parallel, and the result is a report with a generated UI. Hou says 90% of the workflow is automated and it takes minutes instead of Jupyter "elbow grease" .
  • DHH and Matz on agent PRs, "meat proxies," and calibrating trust. A short stretch covering the whole "suggestions, not implementations" stance and Matz's 12 issues fixed from his phone.
  • Riley Brown's agent-native internal app demo. The skill that writes to the app's database is the part to copy.
  • Huntley, "the eighteen-month recap". Agents are "just a while loop, and Ralph is a while loop on a while loop" . His free 300-line workshop has you build your own Cursor/Copilot/Codex . Separately, he says he doesn't pre-install skills or cram markdown into context: "Stop chasing LLM prompt fairies and instead focus on the craft" .

Editorial take: Opus 5.5's cheaper limits make generating code close to free, so the costs now sit in CI, test selection, and deciding how far to trust agent output.

CI Is Now the Bottleneck for Agent-Heavy Teams: Codex-Picked Tests vs. Nix Remote Builders
ThePrimeTime
  • Trash’s firsthand test of newly released Astra found it good at modeling but weaker at code . His PartyKit multiplayer prototype was too laggy to play, and he could not identify the cause, so this is an anecdotal result rather than an isolated comparison of the tools .
  • TJ’s Gotcha prototype used a research-first approach: he said he asked the model to research material from their podcast, which helped it use podcast references without many prompts; the exchange does not establish that this came from training-data recall .
  • Prime said his game was made entirely with “Opus 55” in under two hours, starting that morning . Although its networking synchronized the song and arrows for players, the core game had broken behavior—including advancing by repeatedly pressing down—and Prime said he had not tested the build that day .
Casey Muratori Judges Our Digging Games | TheStandup
Kevin Hou
Profile
  • Firsthand internal workflow: Kevin Hou, who leads part of Antigravity engineering, said Google researchers automated 90% of side-by-side model-evaluation analysis. They load a skills file and ask about an evaluation in natural language; the agent compares control and experiment results, calculates deltas, proposes about 100 hypotheses, dispatches subagents to investigate them in parallel, and returns a report with an interactive UI for filtering and segmenting results. Hou said work that involved manual Jupyter analysis and assembling agent pools, judges, and data pipelines now takes minutes.
  • Agent-team workflow: Antigravity 2.0 separates the IDE from a standalone agent manager, and its agent-teams mode is in public preview via /teamwork. Give the lead agent a specific task; it can ask for clarification, generate specialized subagents that work independently, and select different models for them. Hou described Gemini 3.5 Flash as faster and cheaper, with improved ability to lead teams.
  • Integration pattern: Antigravity sidecars are long-lived plug-in processes that listen for triggers such as SMS, webhooks, cron jobs, or GitHub PRs. Hou said the sidecar spec was planned for release later that summer.
  • Scale demonstration, not an everyday benchmark: Hou said his team built an OS kernel from scratch and ran Doom using 93 subagents over 12 hours, with 15,000 requests and two billion tokens, for under $1,000; he also noted this was not something they would do every day.
Get Out of the Model's Way — Kevin Hou, Google DeepMind
David Heinemeier Hansson (DHH)
Profile
  • David Heinemeier Hansson (DHH) recommends removing the human relay from agent workflows: let agents read error messages directly instead of copying them between agents. Treat agent-created issues and PRs as proposals; have an agent develop suggestions worth pursuing, without becoming attached to the submitted implementation. DHH says agent-authored issues and PRs have exceeded average human submissions on repositories he has managed.
  • Matz described a hands-on issue workflow: after seeing 20 issues appear within an hour, he used the Claude app on his phone to ask it to address the issues and a pull request; by the end of the keynote, it had resolved 12 issues.
  • DHH recommends calibrating agent autonomy to demonstrated competence: use the correctness of prior PRs as a trust signal, check in less as trust grows, and allow for occasional mistakes rather than expecting perfection.
Ruby, Rails, and the future of programming - Matz, DHH & Jeremy Daer
Riley Brown
Profile
  • Riley Brown’s firsthand report: after five days using Opus 5.5 in Claude Code, he found it fast and inexpensive in Claude at the time and said he had not hit a limit; this is a usage impression, not a benchmark.
  • His practical build workflow was to identify the small set of features he actually used across tools such as Notion, My Mind, Typefully, and Google Drive, then ask Claude to combine them into one good-looking app. He iterated through about 20 prompts and tested it with multiple video-editor agencies; the resulting team tool included shareable scripts with image/video comments and downloadable assets.
  • Agent-native pattern: Brown created a Claude Code skill that writes directly to the app’s database, so he can ask the agent to add a video idea and have it create a linked script/thread in the same tool.
My Thoughts on Opus 5.5 .... And Dev Day Predictions
Riley Brown
  • Build workflow: Riley Brown reports using Opus 5.5 in Claude Code and Claude’s Projects feature to build an internal app for his team in a couple of hours. He scoped it to the features he used across Notion, My Mind, Typefully, and Google Drive, gave Claude the combined-app brief, iterated through about 20 prompts, and tested the result with multiple video-editor agencies.
  • Practical product pattern: The app supports script-line comments with image/video attachments, share links for editors, asset downloads and ZIP export; comments are searchable through their linked script lines.
  • Agent integration: Brown added a Claude Code skill called “native note” that writes directly to the app’s database; he asks the agent to add a video idea, and it creates a linked script/thread in the same app.
  • Model experience: In a five-day firsthand report, Brown says Opus 5.5 felt fast and inexpensive and that he had not hit a usage limit. He says Astra may be better at some tasks but is slower and much more expensive; these are his subjective impressions, not a quantified comparison.
My Thoughts on Opus 5.5 .... And Dev Day Predictions
geoff

Geoffrey Huntley argues that LLM-based automated red-teaming makes publishing source code on GitHub reduce its security . He points to @_mattata’s project, which was nearing completion of an audit of every line in the top 1,000 GitHub repositories for $100, with a target of under 30 days . The posts provide no audit method, model, or findings, so this is a security warning and an in-progress scale-and-cost target, not demonstrated vulnerability results .

in the era of automated red-team with LLMs putting your source code on GitHub reduces the security of your source code. [https://x.com/_m… My "Audit every line of the Top 1000 Github repos in less than 30 days for $100" project is nearing completion. I'm currently 7 seconds o…
Geoffrey Huntley

Geoffrey Huntley, an independent AI researcher behind the Ralph loop and former Canva AI tech lead, recommends building an agent rather than only using one: his free GitHub workshop is about 300 lines and teaches the fundamentals by having engineers build a Cursor-, Copilot-, or Codex-like agent. He describes LLM agents as loops—and Ralph as a loop around a loop—and says achieving useful outcomes also depends on context engineering. For engineering organizations, he argues that removing process waste can be a bigger accelerator than AI itself, and recommends reassessing processes such as agile in light of agents.

the eighteen-month recap: AI Engineer, Singapore, May 2026
geoff

For per-request code generation, Geoffrey Huntley suggests creating an HTTP server that also acts as a coding agent and serves outcomes to requests. He says the approach is currently slow, a waste of resources, and that he has not figured out why it matters.

“code may be generated on a per request basis” i’ve been saying, i’ve been saying. if you wanna taste this try promoting the creation of …
Addy Osmani

Addy Osmani says Opus 5.5’s lower price makes Claude limits go about 25% further than Opus 5; cache reads cost 60% less and output is generated 30% faster—a useful cost-and-quota signal for Claude users.

Glad Claude limits feel different! The lower Opus 5.5 price goes straight into your limits. They go about 25% further than Opus 5. Cache …
Theo - t3․gg

In a sponsored Clerk segment, Theo says he has had success asking an agent, “Go set up Clerk on whatever project,” without using Clerk’s copy-and-paste agent skill; he reports using Clerk in T3 Code across web, mobile, and desktop projects.

So much for "Pacing" the Frontier
ThePrimeagen

ThePrimeagen says, “I still read the code,” replying to @trq212’s concern that developers may squander agents’ productivity gains by becoming lazier—a reminder to keep reading code, though the post gives no specific agent or review workflow.

I still read the code ![](https://pbs.twimg.com/media/HTQWK50akAAzwZK.png) [https://x.com/trq212/status/2104273243599405395](https://x.co… I am most afraid of us eating the productivity gains of agents by just becoming lazier.
Peter Steinberger 🦞

Peter Steinberger’s proposed CI strategy is to let Codex decide which tests need to run, reduce CI runs, and run tests hourly; he presents this as a plan, not a measured result. He agrees with Altimor’s concern that CI had become a major engineering bottleneck with very high spend, and says Blacksmith had been an “amazing sponsor” but the load needed to be distributed.

Same. [@useblacksmith](https://x.com/useblacksmith) has been an amazing sponsor but we need to distribute the load. My plan is to let cod… CI has become the top bottleneck of every engineering team I talk to (including Lindy). Our CI spend has become stratospheric. [https://x…
geoff

Geoffrey Huntley recommends using nix-build with large Hetzner bare-metal machines configured as trusted remote builders for coding agents, reserving traditional CI for complex integration and pre-release regression tests; he points to Hydra as a state-of-the-art reference. He offers this as a response to Altimor’s claim that CI had become engineering teams’ top bottleneck and spending had become “stratospheric,” including at Lindy. This is a recommendation, not a report of Huntley’s own implementation or measured results.

nix-build fixes this. grab yourself some large bare metal machines from hetzner, configure em as trusted remote builders for your agent a… CI has become the top bottleneck of every engineering team I talk to (including Lindy). Our CI spend has become stratospheric. [https://x…
Theo - t3.gg

@theo says he has been tracking which models he swears at . @argofowl proposes scoring agents by how often a user has to stop them and ask what they are doing . The post text provides no repeatable scoring method or named model results .

Funny enough, I've been tracking which models I swear at for a while now. The results probably won't surprise you all too much. ![](https… can we score agents by how often you have to stop them and ask 'what the fuck are you doing' [https://x.com/theo/status/21039752081054435…
Kent C. Dodds 🐨

Kent C. Dodds reports firsthand that, over the prior couple of weeks, Cursor made two native Android apps for personal software he wanted; he describes the process as easy, but gives no workflow or implementation details.

In the last couple weeks I've had Cursor make me two native Android apps for personal software I wanted because it's that easy
geoff

Geoffrey Huntley describes a tool he says he “vibed up”: bombadil autonomously explores web and terminal apps with property-based fuzzing, checking that specified properties hold and inspecting I/O effects in underlying layers such as databases, filesystems, and systemd units. He packages the tests as deterministic NixOS machine tests that can take a seed; connecting them to Hydra would let verification fan out across a fleet of machines.

So I vibed up this thing, and it's kinda strange and kinda legit. Imagine you have the ability to do property-based test fuzzing/autonomo…
Theo - t3.gg

Theo reports going through about 3.5 $200 Claude Code accounts in five days, calls Opus 5.5 “incredible,” and describes the $200 Max subscription as a steal . He says he would likely switch to models better than Opus 5.5, but would miss the subscription’s “unlimited” feel .

With much, much effort, I have been able to kill \~3.5 $200 Claude Code accounts in the past 5 days. Opus 5.5 is incredible, and the $200… When models better than Opus 5.5 drop, I will probably move to them, but I will miss this "unlimited" feeling a lot :(
Riley Brown

Riley Brown rates Codex’s computer-use capability highly, especially on Mac . The quoted @jtaby post says Claude’s computer use is much worse than Codex’s, but gives no task details or benchmark .

Codex computer use is absurd. Especially on Mac. [https://x.com/jtaby/status/2104280530573513068](https://x.com/jtaby/status/210428053057… insane how much worse computer use is in claude vs codex
geoff

Geoffrey Huntley says he does not pre-install skills or rely on maximizing markdown in the context window; he favors applying software-engineering craft over searching for magic prompts or a perfect prompt, and says that craft becomes recursive quickly.

There is no silver bullet or perfect prompt but if you do hard core engineering (not related to LLMs themselves but instead how you apply…
Theo - t3.gg

@theo says Claude Code subscriptions let Fable use only half of the available usage, while Opus 5.5 High costs over 2× less than Fable 5.1 High; together, he estimates that switching between those High settings provides roughly 4.3× more usage limits. He says his own switch from Fable 5.1 xhigh to Opus 5.5 High increased his limits by roughly 6.6×.

Why does Opus 5.5 feel practically unlimited when Fable 5.1 was so heavily limited? It's a combination of two things: Opus's efficiency, …