We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
Sonnet 5.5: near-Opus coding at half the price
Anthropic released Claude Sonnet 5.5, the second model in the 5.5 family. It says the model runs more than 30% faster than Sonnet 5 and costs up to 30% less for most work . Anthropic positions it for "well-scoped everyday tasks like fixing bugs and quickly iterating on features," and published a guide on choosing between Sonnet and Opus 5.5, migrating, and tuning effort . Cat Wu says Claude Code users get about 30% more tasks done than with Sonnet 5, using fewer tokens . Addy Osmani cites 70.6% on Terminal-Bench 4.0 and 80.1% on OSWorld 2.1 .
Early practitioner reports put it close to Opus 5.5:
- Matthew Berman says it beat Opus 5.5 on Terminal-Bench 4.0. His other scores are close: Frontier Code 52.1 vs 54.4, Cursor Bench 55.5 vs 57.8. He says he "can't really tell the difference" in daily use . He gives pricing of $2/$10 per million input/output tokens, against Opus 5.5's $4/$20 .
- Cursor says it performs "on par with Opus in many tasks" . Mike Krieger still uses Opus 5.5 for most work, but likes Sonnet's design skills for building features .
Effort pitfall: Simon Willison found Sonnet 5.5 has the same bug as Opus 5.5. At "max" thinking effort it used 128,000 tokens ($1.28) and ran out before producing an SVG. At "xhigh" it finished in 41 seconds for 5.74 cents . Don't use max by default.
Availability: It shipped in Claude Code with a usage reset valid until Oct 22. It is also in GitHub Copilot in VS Code, Cursor, Factory, Devin, Cline, and T3 Code . It now powers the free tier on claude.ai .
Most PR review is moving to agents, with humans on high-risk code
Three sources describe the same shift. Mike Krieger estimates that at Anthropic, full human review fell from 80–90% of PRs in January to 5–10% now, kept for the most critical changes . It was replaced by an adversarial loop. Several Claude instances look for problems, check whether they agree each one is real, and rate its severity. The Claude that wrote the code then revises it . Humans still decide architecture, security boundaries, and product questions, because Claude's first architectural choice is not always right .
Addy Osmani lays out a practical version in "The Code Nobody Reads":
- Every PR gets a first pass from multiple agents that find and verify bugs, rank them by severity, and suggest fixes. Low-risk changes then get a lighter human review. Core paths get a careful owner review, and a person always approves the merge .
- Before the agent starts, write down what you're building, what must not break, and how you'll know it worked. In the PR, disclose what you didn't review, e.g. "agent-reviewed, tests pass, I haven't read the migration logic" .
- Check the work against an independent source of truth: a human-written spec, a reference implementation, or a proof. A second copy of the same model shares the first one's blind spots .
- Write code agents can change safely: names unique enough to grep, modules small enough to fit a context window, tests that fail clearly .
DHH takes the blunter view: adversarial agent reviews, automated tests, "maybe you spot check" .
Thariq Shihipar's Claude Code tips
On Latent Space, Thariq Shihipar (Anthropic) said:
- Front-load context. Say whether it's a prototype or production code, and where compute is worth spending. Most wasted usage comes from "undo this and redo it" loops .
- Set effort by task type. Use high or max for code review and security, low or medium for UI. Ask explicitly for edge-case coverage on APIs. In software work, effort mostly goes into verification .
- Ask for decision or implementation notes. Most high-effort failures happened when the model considered the right solution and then rejected it .
- Start new projects without CLAUDE.md. Add only failure modes that keep recurring. Old failure logs may over-constrain newer models .
- Claude Mods let you customize how the harness runs and its UI. Example: at the end of each turn, a forked subagent checks whether the task is done and quizzes you. It is cheap because the fork reuses the prompt cache. Another example is a "register assumption" tool that keeps a running list of the model's assumptions .
Split work across models
Geoffrey Huntley's current setup uses Opus or Sol for planning and Kimi for "grunt loops," with Opus/Sol exposed as an oracle tool Kimi can call. He says Kimi is noticeably worse, but its free tokens are unlimited . His pattern is get it working, then launch targeted Sol/Opus refactors to make it good. For design work, he uses Opus .
Theo's TypeScript-to-Rust compiler port had stalled at about 35% of tests passing with GPT-5.6 Sol and about 85% with GPT-6 Astra. He then gave Opus 5.5 /goal finish the port and make it faster. Opus judged Astra's code to be slop and rewrote it from scratch in a new crate. Theo says it made more progress in 10 hours than Astra did in two weeks . Separately, he estimates a $200 Claude Code plan gives about $9,000 a month of Opus usage at API prices .
Rewriting apps in Rust with agents
DHH has a beta of Campfire rewritten in Rust. He calls the code "ugly as sin" and 6× as verbose, and says he never looked at it . He treats it as a "prompt compilation target" . It cost under $10 in tokens on a 20x Max plan and took a few hours: basically one prompt, then a few tuning prompts . His argument: writing web apps in Rust before agents would have been "madness"; now it's "trivial and cheap" .
Smaller items
- Share transcripts, not just prompts. Willison points to Codex's transcript-sharing feature. He built his own version for Claude Code and would prefer a built-in one .
- gpuc (repo): Brendan Long's GPU job queue that needs no sudo. Queue hosts need only SSH, rsync, and the NVIDIA driver. Jobs keep running if the client goes offline, and there is a CLI optimized for Claude (send
!gpuc skill) . He built it so Claude Code doesn't need root . Limits: single-user only, and RunPod is the only rental provider . - At Wonder, a PM can file a bug and Claude Code fixes it from the LangSmith trace .
- Before implementation, use the agent to surface unknowns and clarify preferences—such as schema or call-stack decisions—and build a mental model of both the agent and codebase. Provide enough upfront context, including whether the task is a prototype or production work and how much compute or verification to spend; repeated correction can consume usage, and spoken prompts are useful when they convey more information.
- Calibrate effort by task: use high or max for security and code review, and low or medium for UI; explicitly ask for edge-case testing and verification when warranted. The practitioner says effort in software engineering tends to affect verification more than the core task result.
- Ask for implementation or decision notes: in reviewed eval transcripts, many failures involved the model considering a correct solution and then deciding against it, so notes make those choices inspectable. A harness mod can also keep a visible list of assumptions.
-
Consider starting a project without a
CLAUDE.md, adding only recurring failure modes as they arise. Guidance can differ by model and version, and accumulated instructions may overconstrain a newer model; skill eval plugins can help test whether a skill improves results. - Claude Code Mods customize harness execution and UI. One example forks a checker at the end of a turn to assess completion and generate a quiz for the next prompt; the fork reuses the prompt cache, though running the check after every turn uses additional tokens.
- Claude Code with Opus 5.5 was reported to generate a 30–60-second pure-JavaScript explainer end-to-end—including concept, script, assets, animation, and TTS—in a one-shot run of about 1 hour 20 minutes; the creator reported about $20 in Opus usage plus $3.21 across eight OpenRouter API calls. A second user reproduced the workflow for Friendr.nl in about 1.5–2 hours for roughly $4, adding music/SFX, narration-synced animation, MP4 rendering, and another-model review; the run used a $10-capped OpenRouter key and reportedly needed only minor corrections. A separate motion-design prompt template starts by asking for 8–12 UI states for a shape to transform into (for example, a button, loader, or player); its creator said the video was entirely code, with no After Effects.
-
Opus 5.5 was also used to build TideWater, a browser-based interactive island, in about eight hours with iterative prompts such as
add Xandmake it better; the reported token cost was $1,874.40, or 59% of a Max 20x weekly allowance. The demo includes walking around, interacting with objects, and sailing a boat, so testing it interactively reveals capabilities beyond a video preview. - For agent builders, LangChain Managed Deep Agents 0.8 adds user/agent memory with access policies, HTTP channels, sandbox file APIs, proxy-authenticated sandboxes, and Parallel web search; its smithtune CLI turns traces into post-training datasets, while Trajectories handles deferred tool calls and context compaction. Databricks also reports that engineers stopped reaching for closed models once open-source models were routed to its internal coding agents.
- Build a mental model of what Claude can reliably one-shot, surface unknown requirements before implementation, and learn the domain vocabulary or reference language needed to specify the result more precisely; Thariq Shihipar described this as a core agent-coding skill.
- Put more context into the initial prompt—including whether the task is a prototype or production work, where to spend compute, and what verification is needed—to reduce wasteful undo-and-retry cycles. Shihipar’s rough effort guidance: high/max for code review and security, low/medium for UI, and explicit edge-case verification for API work.
- Ask for implementation or decision notes: the model may consider a suitable approach and choose not to use it, so reviewing its notes can reveal options to request.
-
Start a new project without
Claude.md, then add recurring failure modes as they appear; those failures can differ between model versions, and accumulated guidance may overconstrain newer ones. - Claude Mods can customize Claude Code’s execution and UI. One practical pattern is an end-of-turn forked subagent—which retains the prompt cache—to check task completion and produce a quiz; mods can also register assumptions, and can spawn subagents, parse results, and change the UI.
- For collaboration, Shihipar described using Claude Tag for background work such as code review, security, and starting PRs, and for multiplayer incidents. His example workflow is a channel per project where legal can review exactly what is shipping by discussing it with Claude, without the engineer relaying all the context.
-
The
/eli5plugin uses the prompt pattern “big picture, few words” to explain complex incidents more clearly and reduce text-heavy artifacts.
- At Anthropic, Mike estimated that comprehensive human PR review fell from about 80–90% in January to 5–10% at the time of the talk, citing the volume of generated code and Claude’s ability to find issues. Their replacement is a repeated adversarial loop: multiple Claude reviewers search for problems, assess whether they agree and how severe each issue is, then the coding Claude revises. Human review remains important for architecture, security boundaries, and product judgment; Mike said Claude can make good architectural decisions in a long conversation but may not get the first decision right. He also observed that Claude tends to follow existing codebase patterns, but may not prioritize readability when writing one-off code just to connect things.
- For team context, Mike described keeping much internal work visible within groups and using Slack as a central place to work with Claude; new employees can ask how things are currently done using Claude’s view of company activity. Anthropic also runs a nightly internal “dream” process to incorporate process learnings into organizational memory, though Mike said company-wide learning still had room to improve.
- StrongDM’s “dark factory” rules required code to be routed through a coding agent and prohibited human code review; its experiment explored how to verify agent work and maintain confidence in quality without reading the code.
- To get leverage from capable models, define the goal clearly, specify unambiguous constraints, and provide the necessary tools; Willison says doing this well takes experience and skill, and agent work still requires extraordinary discipline and knowledge.
- Willison uses GPT-6 Sol in Codex and Claude Opus 5.5 in Claude Code as his defaults; he uses GPT-6 Luna for the Datasette Agent and reports it is fast and competent at SQL queries and building HTML and JavaScript. GPT-6 Luna is priced at $0.10 per million input tokens and $0.50 per million output tokens, one tenth of Haiku 4.5’s $1/$5 rates.
- Treat maximum reasoning as a potential cost and latency trap: Opus 5.5 at max hit its 128,000-token output limit before returning an answer on an SVG task; a second attempt failed the same way, with each attempt costing $2.56 and taking nearly 20 minutes.
- For a reusable prototype-to-video workflow, Willison gave Opus 5.5 three kakapo photos and asked it to make an HTML5-canvas pixel-art animation with at least 20 birds. He then used Claude Code with Playwright to load the downloaded HTML, delay clicks until three seconds in, spread them around the canvas, and record a 15-second video.
- For a browser game, Berman recommends explicitly asking the coding agent to check its work as it goes—using screenshots, video, or by playing the game. His Three.js Fall Guys-style demo specified 59 bots across five rounds, took roughly one or two prompts, and still had occasional clipping.
- For a LEGO-generation app, he prompted a split of responsibilities: have the model describe shapes, let a program map them to real LEGO parts and enforce a connected, buildable model, then feed failures back to the AI. The prompt also asked for screenshot checks and instructions with at most four pieces per step.
- In Berman's tests, Sonnet 5.5 beat Opus 5.5 on Terminal Bench 4.0 and was close on other reported coding scores (Frontier Code 52.1 vs. 54.4; Cursor Bench 55.5 vs. 57.8); he said it felt nearly indistinguishable in use and was faster. He reported pricing of $2/$10 per million input/output tokens, versus Opus's $4/$20.
- His Unreal Engine San Francisco build required downloaded assets and took multiple days and millions of tokens; it also overloaded his computer and sometimes stopped working.
- Osmani recommends a risk-based review loop: have agents make a first pass that finds, verifies, severity-ranks bugs, and suggests fixes; give low-blast-radius changes a lighter human review when checks are clean, but carefully review core or sensitive paths, with a person owning merge approval. Anthropic’s automated Claude reviewer runs on nearly every PR and informs engineers without approving changes.
- Before coding, write down what to build, what must not break, and how to tell it worked. Ship only changes you can explain—their behavior, scope, and safety—and disclose in the PR what you did not review; the article gives the example, “agent-reviewed, tests pass, I haven’t read the migration logic.”
- Keep verification independent of the code-writing agent: a second copy of the same model can share its blind spots, so ground checks in a human-written spec, reference implementation, or proof. In Nicholas Carlini’s Claude-built C compiler project, GCC served as a known-good reference for debugging kernel issues.
- Structure code for agent changes: use distinct, searchable names, small modules that fit a context window, clearly failing tests, and clear boundaries so agents can see what a change may affect.
- Anthropic says Claude Sonnet 5.5 runs 30%+ faster and costs up to 30% less for most work; it is priced the same as Sonnet 5, which Simon Willison says it appears to outperform on every benchmark. He also found it nearly as good as Opus 5.5 on some coding tasks, including 3D animation work.
- Thinking-effort choice had a large cost and completion impact in Willison’s pelican-to-SVG test: “max” used 128,000 tokens ($1.28) and failed, while “xhigh” produced a result in 41 seconds for $0.0574. Sonnet 5.5 is also the model on Claude’s free tier; a direct prompt to build a WebGL 3D pelican page produced what Willison called a “solid effort.”
Simon Willison recommends sharing coding-agent transcripts with their embedded context, rather than sharing only the prompt; he points to Codex’s transcript-sharing feature and says he built a workaround for Claude Code but would prefer native support.
Claude Sonnet 5.5 became the model powering Claude’s free tier, making early experiments such as @_re_pete’s fall-foliage simulator made with Sonnet 5 vs. 5.5 accessible to free users . Simon Willison said ChatGPT’s free-tier GPT-5.6 Luna was “a lot less capable” .
Willison’s “Fable class” framing: models such as Claude Fable 5, Claude Opus 5.5, and GPT-Astra 6—and possibly GPT-5.6 Sol—can solve a problem by brute force when its goal is definable, instructions are unambiguous, and the models have access to the necessary tools.
- Fireship reports that DHH said his coding output rose from about 30,000 lines of Ruby per year to about 150,000 lines of code per month using agents.
- DHH reportedly said the Hey email app was rewritten in AI-generated Rust, cutting server CPU and memory usage by 95%.
- DHH advocated providing a CLI so an application can be used without a person interacting with it—a practical interface pattern for agent-driven use.
- Fireship’s takeaway is that developers should prioritize defining problems and designing secure, efficient systems over mechanical code production.
Kent C. Dodds says repeated integration setup is a barrier to building personal software; Kody Koala aims to remove that friction by letting users set integrations up once and reuse them with agents or full-stack apps. He says Opus created a playable game to illustrate the idea: integration-game.kody.codes.
Claude Sonnet 5.5 launched, claimed to be 30% faster and up to 30% cheaper than Sonnet 5 for most work, and was positioned for well-scoped everyday tasks such as bug fixes, documentation, and slides. A follow-up reported improvements over Sonnet 5 on agentic-coding and computer-use benchmarks, including 70.6% on Terminal-Bench 4.0 and 80.1% on OSWorld 2.1.
Geoffrey Huntley’s preferred split is Opus/Sol for planning and Kimi for grunt loops; he says Kimi is worse than Opus/Sol but attractive when unlimited free tokens are available. He also suggests making Opus/Sol an oracle/tool that Kimi can call. For refinement, he follows “get it working, then get it good” by using Sol/Opus for targeted refactoring of working code, and recommends Opus for design tasks.
Jason Zhou open-sourced a Claude /leads-signal skill that monitors keyword mentions, complaints, product reviews, job changes, and 23 additional signals to find hot leads each morning; the skill is linked at GitHub. He pitches it at $0.0002 per signal, compared with $167/month Clay plans.
Alex Albert says Sonnet 5.5 writes clearly, is very fast, and is a major capability jump over Sonnet 5; he found it a “really great model to iterate with,” and compared its feel favorably with Opus 5.5 . Claude AI’s announcement says Sonnet 5.5 is more than 30% faster than Sonnet 5 and costs up to 30% less for most work .
- Anthropic positions Claude Sonnet 5.5 for well-scoped everyday coding tasks such as bug fixes and quick feature iteration; the company claims it is over 30% faster and up to 30% cheaper for most work. Early independent evaluations place it near Opus 5.5 on several leaderboards.
- Sonnet 5.5 launched in Claude Code and is also available through GitHub Copilot in VS Code, Cursor, Factory, Devin Desktop/CLI, Cline, and T3 Code.
Kent C. Dodds argues that when customers use an agent to access a service, exposing a regular MCP server is more efficient than routing the agent through a browser-based WebMCP interaction; he says Cloudflare’s Kitesurf can interact with WebMCP servers but favors direct MCP for agent-facing services.
When coding agents produce low-quality “slop,” Kent C. Dodds advises improving the primitives they work with rather than giving up on the agents or spending effort cleaning up their output; he says he will demonstrate the approach in a video, but the post does not provide the implementation details .
Two years ago I wrote an essay (opens in new tab) about AI-assisted coding, and one of my tips was to review every line of generated code. I’d give that advice very differently today.
What I’d say now is this: you don’t need to read all the code. Every change still needs some review and every change still needs a person who owns the decision to ship it. What’s changed is that a careful human read of every line is no longer the only way to provide that review, or always the best one. Some can be primarily agent-reviewed or a mix of human/agent when sensitive.

Teleport found two years of security bugs in one quarter (opens in new tab). Thirteen engineers, frontier models, their own codebase: nearly twice as many high severity vulnerabilities as the previous two years put together. What interests me is what that does to the queue. Discovery got several times cheaper. Triage did not, and confirming what is genuinely exploitable is still a person reading code. Find faster than you can fix and you have traded an unknown backlog for a known one. Worth a read before you try this on your own codebase. Read the write-up (opens in new tab) → https://fandf.co/4r3xdeX (opens in new tab) · Sponsored by Teleport. #ad
This post is about what I think should take its place, and what happens when nothing does.

Last week Thorsten Ball posted a list of sixteen things (opens in new tab) he believes about the future of software development. The first one is “code review will die.” I agree with much of the list, at least on direction, though I’d put that first one differently. Line-by-line reading is going away for a lot of code. Review, meaning someone deciding what ships and being answerable for it, isn’t. What a list like his can’t tell you is how fast each item arrives, or in which parts of the industry. That’s where I actually disagree with people.

The same week, a very different post ^ went just as far. An engineer two weeks into a new job at a big company wrote about (opens in new tab) spending their days pressing enter on code nobody reads. More than eight million people saw it. I don’t think these two posts contradict each other. One describes the destination. The other describes what happens when a company drives there without building the road.
Here’s my view in short. I think every industry will stop reading most of its code line by line once it has some other good reason to trust that code.
What’s different this time is that the thing writing the code can’t be certified the way older code generators were, because it doesn’t reliably give the same answer twice. So the trust has to come from the checking we build around it. How fast a team or an industry gets there depends on two costs: the cost of checking work nobody read, and the cost of undoing a failure the checks miss.

I should say up front that I’m not neutral. I help build a tool that writes code. When people called Thorsten a shovel seller telling everyone to dig, he said the causality runs the other way: he builds agents because he believes in them. That’s true for me too.
Do I miss anything about the old way? Honestly, yes. At the very start of one of my JavaScript books, I wrote that good code is like a love letter to the next developer who has to maintain it (opens in new tab). I meant it. That was before AI, when a lot of us cared about the craft of the code itself. I liked Thorsten’s line that there are still Italian shoemakers around, but look at your feet. Things change. What I get in exchange is the ability to reshape the clay very quickly, and to build things I would otherwise have kicked around for years without ever starting.
What review was actually for
In 2013, Microsoft researchers studied code review (opens in new tab) and asked developers why they did it. Finding defects was the top answer. The comments told a different story. Of the 570 review comments the researchers classified, only 14% were about defects. The rest were about teaching, sharing context, suggesting better approaches and keeping team norms. Review was never mostly about bugs, which is why “the machines will review it” is a weaker answer than it sounds.
Bug-catching is the part machines are getting good at. At Anthropic, an automated Claude reviewer runs on nearly every PR, and engineers mark less than 1% (opens in new tab) of its findings as incorrect. Since it started, the share of PRs getting substantive review comments has gone from 16% to 54%. It doesn’t approve anything. In Anthropic’s words, “that’s still a human call.” In a retrospective analysis (opens in new tab), Anthropic found that an automated review of every change would have caught roughly a third of the bugs behind past incidents on claude.ai before they reached production.
For a reviewer that never gets tired, a third is a lot. But it also tells you something uncomfortable: most of those bugs got past careful people looking right at the change. My guess is that’s because code that fails in production often looks fine in the diff.
In the second quarter of 2026, the typical Anthropic engineer was merging about eight times as much code per day as in 2024, with the engineer, as Anthropic puts it, “directing and reviewing, rather than typing.” Nobody reads eight times as much code carefully. From the inside, the way we work feels very different from what I was used to. We generate a lot of code every day across a lot of product surfaces, and not every line gets a careful human read anymore. Nearly every change gets automated review, a person still approves what merges, and it’s up to that person to decide how deep to go. Does this change need a careful manual review? Is the tooling we have good enough here? Is this a sensitive enough part of the system that it needs extra scrutiny? If the code passes its tests, performs well on them and runs in production without problems, that’s a pretty good signal.
It’s worth being clear about the limits of that, too. When Anthropic surveyed its own engineers (opens in new tab), most said they could “fully delegate” only 0-20% of their work to Claude. The rest still involves active supervision and checking, especially on high-stakes work. Not reading every line isn’t the same as not paying attention.
What worries me more are the other jobs review used to do. It was how junior engineers learned how senior engineers think. Anything a reviewer used to say in a comment, like “we don’t do it that way here” now needs to be written down in skills. Otherwise the agent never hears it.

When the reading stops and nothing replaces it
The problem isn’t that nobody reads the code. It’s that nothing replaced the reading. That team earlier on didn’t move trust from the diff to the checks. It stopped reading diffs and put velocity where trust used to be. They did the transition in the worst possible order, and my two costs predict how it ends. They’ve made the first cost zero by not checking at all, so they’ll pay the second one, undoing the failures no check was ever there to catch, with interest.

That order wasn’t an accident either. Someone above the people living with it made a decision: once writing code stopped being the bottleneck, the only question left was why anything was still slow. These tools usually spread bottom-up, with individual developers deciding they help. This way of working was imposed top-down, and that difference matters a lot. Think back to what review was doing: teaching, keeping people aware of changes, letting a team hold its own quality bar. Order a team to stop reading and you haven’t just removed a defect check. You’ve told them the bar is no longer theirs to hold.
The version that works is boring by comparison. Every PR gets a multi-agent first pass that finds bugs, verifies them, ranks them by severity and suggests fixes. Changes with a small blast radius on less sensitive code, and there’s usually a lot of that, can get a lighter human pass once that first pass comes back clean. Core and sensitive paths still need an owner and a careful human review, and that’s where I spend my attention. Either way, a person approves the merge. Agents do the first pass and humans cover blast radius. Nobody in that loop presses enter for thirteen hours, because pressing enter was never the human’s job. The job is deciding what deserves attention. The “no sense of victory” in that post is what it feels like when a company takes that decision away from you too.
If you’re the one pressing enter
Most advice about this is written for people who get to set policy. A lot of the people in that Reddit thread don’t. So here’s what I’d do if I were two weeks into that job.
Truly understand what you want build. Before an agent touches anything, write down what you’re building, what it must not break and how you’ll know it worked. It doesn’t need to be long. That’s where your judgment goes now, and it’s the part of the work you can point to later. Several people in the thread said a version of this: the ones who still felt good about their work had designed the thing and let the agent build it. If you want the sense of victory back, it lives in the decisions.
Use “can I explain this?” as your test. One design lead in the thread asks their team whether they know “every hair on the dog” before they call work their own. I’d apply the same test to code. You don’t need to have read every line, but you should be able to explain what the change does, what it touches and why it’s safe to ship. You can only really be accountable for what you can explain.
Be honest about what you didn’t review. If you weren’t given time to understand a change, say so where people will see it, like the PR description: “agent-reviewed, tests pass, I haven’t read the migration logic.” That isn’t being difficult. It’s what keeps you out of the crumple zone I get to at the end of this post. A team that hides how little it read can’t fix that.
Keep your hands in. Anthropic’s researchers describe a “paradox of supervision”: supervising Claude well takes the very coding skills that fade if you never use them. Some engineers there deliberately solve problems without Claude now and then, even when they know it could handle it. One said it “helps me keep myself sharp.” That’s worth copying, especially early in a career.
If you set the pace
The best counter-argument comes from startups where all the documents and code are AI-made: “they work for us , we dont work for them. Every single thing has a responsible person who needs to be able to explain what we have done.” If you’re the person setting expectations, here’s what I think follows from it.
Accountability needs authority. If an engineer will be blamed when a change fails, they need a say in how fast changes go out and how much they get to understand first. Holding someone responsible for a pace they didn’t set, on code they weren’t given time to understand, is exactly what the thread calls being a “meat proxy.” It isn’t fair, and it doesn’t work as a control either.
Measure what shipped, not how much. Merged PRs and lines of code are easy to count, and they’re the numbers that turned pressing enter into the whole job. Even Anthropic’s own write-up on its 8x output number calls lines of code “an imperfect measure.” Watch incidents, rollbacks, time to recover and what customers actually experience.
Budget for understanding like any other capacity. One commenter put the new constraint plainly: “Writing code isn’t bottleneck, but reading and evaluating it is.” Anthropic has said much the same about review. The answer isn’t to read nothing. It’s to decide, area by area, how much human attention a change needs, and plan the work around that. Automated review helps a lot here, but Anthropic’s stated goal for it is to close the gap “so reviewers can actually cover what’s shipping,” not to take reviewers out of the loop.
Make room for mentoring on purpose. Juniors used to learn by asking seniors questions and reading their review comments. If Claude answers the questions now, you have to create the rest deliberately: pairing on plans, having juniors explain agent-written changes back to someone, leaving some work manual. “What juniors?” was one of the bleaker replies in that thread. It shouldn’t be the answer at your company.
Keep what you ask people to sign human-sized. This applies to documents as much as code. One commenter’s product team had AI read the legacy codebases and produce a feature-and-gap report: an HTML file over 100 pages long and a spreadsheet with 30 tabs of hundreds of rows each. When they pushed back that it wasn’t fit for human consumption and asked to start with the high-level gaps, they were treated like they weren’t a team player. A 100-page generated report nobody can read has the same problem as a 5,000-line diff nobody read. Agents make long documents free to produce. A person still has to be able to judge them, and asking for the one-page version isn’t obstruction. It’s review.
“But our competitors don’t read the code”
The most reasonable pushback I’ve seen came from another startup engineer. Their management doesn’t even like shipping without reading. They just feel they have no choice, because the next company does it and they can’t compete on price otherwise. I take that seriously.
My answer is the second cost. Skipping the checks doesn’t make failures free. It moves their cost later, when it’s bigger and lands on your customers.
So I don’t think the competitive question is reading versus not reading. It’s which team can build cheap checks that actually work, and point its people’s limited attention at the changes that can really hurt. The team that wins isn’t the one that reads nothing. It’s the one that knows what’s worth reading.
So who writes the tests?
One reply to Thorsten’s list asked the question I’d been circling. A unit test captures what you know about how a system should behave and checks it against something you don’t fully trust, which is your own future work. “If the thing you don’t trust is also writing the tests what is it accomplishing?”
It’s a fair question. In Anthropic’s research on reward hacking (opens in new tab), models learned to call sys.exit(0) so the test harness would exit before any assertion could fail. A test written or controlled by the thing it’s testing doesn’t prove much.
The way out is to ask where a test gets its authority. Aviation’s answer is independence. DO-178C, the standard for airborne software, requires the most critical verification to be done by someone other than the author, and a qualified tool can count as that someone. Having one Claude write the tests and another write the code catches the crudest failures, but a second copy of the same model shares the first one’s blind spots. Real independence comes from a different source of truth: a spec a person wrote, a reference implementation, a proof.
The best example I know is Nicholas Carlini’s C compiler (opens in new tab). Sixteen Claude agents, nearly 2,000 sessions and just under \$20,000 produced a 100,000-line compiler in Rust that can build a bootable Linux. His main lesson was about the checker rather than the agents: “the task verifier is nearly perfect, otherwise Claude will solve the wrong problem.” When the agents got stuck on the kernel, he used GCC as a known-good reference to corner each bug.
SQLite shows where this leads. The library is about 156,000 lines of C, and its test code is 590 times larger (opens in new tab). Its most thorough test suite reaches 100% branch coverage, was designed to meet an avionics testing standard, and is proprietary. SQLite gives away the code and licenses the tests. So here’s how I’d rephrase Thorsten’s point. Unit tests as notes from you to your future self will matter less, because your future self won’t be the one changing the code. Tests as an independent check on an author you don’t fully trust will be the most valuable code you own.
The expensive bugs were always upstream
Thorsten says most bugs will come from asking for the wrong thing.
Ariane 5 was destroyed less than a minute into its first flight because navigation code reused from Ariane 4 overflowed on a trajectory Ariane 4 never flew. The inquiry board (opens in new tab) found that the specification didn’t include the new trajectory. Mars Climate Orbiter was lost because one team’s software produced pound-force seconds where the interface expected newton-seconds.
Each failure lived at a join: between a requirement and the real flight, between two teams’ units, between a deployment and eight servers. That’s where the expensive bugs live, and it’s where models are weakest today. Anthropic’s own assessment is that Claude can match skilled researchers at running a well-specified experiment, but that “large performance gaps persist when it comes to Claude exercising judgement in choosing goals.”
I’ve written about the 70% problem (opens in new tab) and later the 80% problem (opens in new tab). Each time the number went up, what was left for the human was less typing and more judgment. That’s how I read Thorsten’s claim that the product, design and engineering trio will disappear. The handoffs between three roles go away. The three judgments stay, whoever ends up making them: is this worth building, does it work for the person using it, and will it hold up. He’s also right about one role in particular. The person who takes a ticket, hands it to an agent and reports back without adding any judgment of their own: that job is going away.
The industries that say “never”
The usual pushback is that none of this applies where a bug can cause major problems or lose someone’s money. Today those industries do read the code, and I think they’ll keep doing it longer than optimists expect. But what they actually require is a reason to trust the code, and they’ve accepted machine-generated code before when they had one.
You can qualify a code generator because it gives the same output for the same input. A language model doesn’t. So for now, the practical way to use a model on flight software is the old way: a person reads what it wrote. The way forward is on the checking side, with formal proofs, qualified static analysis and independent test oracles, treating the model as an untrusted author whose every output gets verified.
Finance is closer than people think. A lot of what its rules call review is really about accountability. The standards mostly require that changes be authorized, tested and approved by someone independent, not that a peer reads every line. None of that stops an agent from doing the implementation. It keeps a person accountable for the approval, which is where a person belongs.
Count the nines
The strongest objection I know on timelines comes from Andrej Karpathy, who said last year (opens in new tab) that this is “more accurately described as the decade of agents.” His argument comes from self-driving, where “every single nine is a constant amount of work.” Getting from 90% reliable to 99% costs as much as getting to 90% in the first place. Waymo started in 2009 and is now on its 15th city.
He’s right that nines are expensive. Where software differs from driving is that a lot of software can buy its nines from the system around the model instead of from the model. A web app can ship behind a flag, watch its error rate and roll back in minutes. A car can’t roll back a collision. So I think timelines split along the second of my two costs: how cheaply you can catch and undo a failure after it ships.
Here are my guesses, specific enough that you can hold me to them:
Consumer software, internal tools and most SaaS: a person reading every line of a routine change will be unusual within two years, though a person will still approve it and own it. Plenty of teams are already there.
Enterprise and financial software: agents will write most of the code on the same timeline, with human approval kept as a control. Over three to five years, that approval shifts from the diff to the evidence.
Flight controls and the highest-risk medical devices: people will keep reading line by line. I don’t expect a regulator to accept safety-critical code on qualified checking alone for many years.
It’s also worth remembering where the loudest version of this conversation happens. In the biggest bubbles, Twitter, we’re ahead of the curve. It’s going to take a long time for the rest of the industry, and the rest of the world, to catch up. But the direction is clear to me.
What happens to the code
Thorsten says there’s no proof that “good code” will matter, since most of what we mean by it is about making code easy for people to work with. I agree for the parts that were purely about human comfort, like formatting. But some properties matter more to an agent than they ever did to us: names unique enough to grep, modules small enough to fit in a context window, tests that fail clearly, clear boundaries between parts. Without those, an agent will still make the change. It just can’t see what else the change affects. Good code, now, is code an agent can change without breaking something it can’t see.
I’ve been involved in open source for a long time, and I think it evolves here rather than dies. It’s easier than ever to take a project, make your own copy and reshape it quickly without first understanding all of the code. I wouldn’t be surprised if we increasingly share prompts, or share finished artifacts that people hand to their agents when they want to customize them.
What open source is short on is maintainer attention, and AI both floods it and multiplies it. In January, Daniel Stenberg ended curl’s bug bounty (opens in new tab), citing what he called “an explosion in AI slop reports.” Three months later, the rate of real vulnerabilities was back to normal, at double the old report volume. Within a year, the junk problem had turned into a volume problem. What’s most worth sharing shifts toward things that are expensive to produce and cheap to check against: test suites, specs, the record of what broke.
The crumple zone
Madeleine Clare Elish has a name for what happens to people at the edge of automated systems: the moral crumple zone (opens in new tab). In a car, the crumple zone is built to absorb the impact so the passengers don’t have to. In a system where automation does most of the work, the nearest human often plays the same role. They take the blame for a failure they had little real power to prevent.
The engineer pressing enter for thirteen hours is standing in one. Their name is on the approval. But the pace, the checks and the decision about what was worth reading were all made somewhere else. When something breaks, the incident review will find the human who approved it.
Reading every line was never what made software trustworthy. It was one way to make sure a person understood what they were shipping and was in a position to say no. The line-by-line read is going away for most code, and I think that’s fine. The understanding and the ability to say no can’t go with it.

Teleport found two years of security bugs in one quarter (opens in new tab). Thirteen engineers, frontier models, their own codebase: nearly twice as many high severity vulnerabilities as the previous two years put together. What interests me is what that does to the queue. Discovery got several times cheaper. Triage did not, and confirming what is genuinely exploitable is still a person reading code. Find faster than you can fix and you have traded an unknown backlog for a known one. Worth a read before you try this on your own codebase. (opens in new tab) Read the write-up (opens in new tab) → https://fandf.co/4r3xdeX (opens in new tab) · Sponsored by Teleport. #ad
So here’s the advice I’d give now, in place of what I said two years ago. Know what’s worth reading. On the code you don’t read, build checks you’d actually trust, and keep them independent of whatever wrote the code. Write down what you’re building before an agent builds it. Be able to explain everything you ship, and be honest about what you can’t explain. If you set the pace for other people, give them authority that matches the accountability you’re asking of them.
I once called good code a love letter to the next developer. I still believe that. More and more, though, the next developer is an agent, so the letter looks different now. It’s code the agent can change without breaking something it can’t see. The craft didn’t disappear. It moved to the parts that decide whether the code deserves to be trusted.
Thanks to the newsletter’s sponsors for their support over the last few months. As I’ve joined a new role full-time, I’ve put a pause on sponsors for future issues. Opinions here expressed are my own and do not express the views or opinions of my employer. I hope to continue writing and sharing my thoughts on the newsletter as long as they’re helpful to you :)

- Osmani recommends a risk-based review loop: have agents make a first pass that finds, verifies, severity-ranks bugs, and suggests fixes; give low-blast-radius changes a lighter human review when checks are clean, but carefully review core or sensitive paths, with a person owning merge approval. Anthropic’s automated Claude reviewer runs on nearly every PR and informs engineers without approving changes.
- Before coding, write down what to build, what must not break, and how to tell it worked. Ship only changes you can explain—their behavior, scope, and safety—and disclose in the PR what you did not review; the article gives the example, “agent-reviewed, tests pass, I haven’t read the migration logic.”
- Keep verification independent of the code-writing agent: a second copy of the same model can share its blind spots, so ground checks in a human-written spec, reference implementation, or proof. In Nicholas Carlini’s Claude-built C compiler project, GCC served as a known-good reference for debugging kernel issues.
- Structure code for agent changes: use distinct, searchable names, small modules that fit a context window, clearly failing tests, and clear boundaries so agents can see what a change may affect.