We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
Sonnet 5.5: near-Opus coding at half the price
Anthropic released Claude Sonnet 5.5, the second model in the 5.5 family. It says the model runs more than 30% faster than Sonnet 5 and costs up to 30% less for most work . Anthropic positions it for "well-scoped everyday tasks like fixing bugs and quickly iterating on features," and published a guide on choosing between Sonnet and Opus 5.5, migrating, and tuning effort . Cat Wu says Claude Code users get about 30% more tasks done than with Sonnet 5, using fewer tokens . Addy Osmani cites 70.6% on Terminal-Bench 4.0 and 80.1% on OSWorld 2.1 .
Early practitioner reports put it close to Opus 5.5:
- Matthew Berman says it beat Opus 5.5 on Terminal-Bench 4.0. His other scores are close: Frontier Code 52.1 vs 54.4, Cursor Bench 55.5 vs 57.8. He says he "can't really tell the difference" in daily use . He gives pricing of $2/$10 per million input/output tokens, against Opus 5.5's $4/$20 .
- Cursor says it performs "on par with Opus in many tasks" . Mike Krieger still uses Opus 5.5 for most work, but likes Sonnet's design skills for building features .
Effort pitfall: Simon Willison found Sonnet 5.5 has the same bug as Opus 5.5. At "max" thinking effort it used 128,000 tokens ($1.28) and ran out before producing an SVG. At "xhigh" it finished in 41 seconds for 5.74 cents . Don't use max by default.
Availability: It shipped in Claude Code with a usage reset valid until Oct 22. It is also in GitHub Copilot in VS Code, Cursor, Factory, Devin, Cline, and T3 Code . It now powers the free tier on claude.ai .
Most PR review is moving to agents, with humans on high-risk code
Three sources describe the same shift. Mike Krieger estimates that at Anthropic, full human review fell from 80–90% of PRs in January to 5–10% now, kept for the most critical changes . It was replaced by an adversarial loop. Several Claude instances look for problems, check whether they agree each one is real, and rate its severity. The Claude that wrote the code then revises it . Humans still decide architecture, security boundaries, and product questions, because Claude's first architectural choice is not always right .
Addy Osmani lays out a practical version in "The Code Nobody Reads":
- Every PR gets a first pass from multiple agents that find and verify bugs, rank them by severity, and suggest fixes. Low-risk changes then get a lighter human review. Core paths get a careful owner review, and a person always approves the merge .
- Before the agent starts, write down what you're building, what must not break, and how you'll know it worked. In the PR, disclose what you didn't review, e.g. "agent-reviewed, tests pass, I haven't read the migration logic" .
- Check the work against an independent source of truth: a human-written spec, a reference implementation, or a proof. A second copy of the same model shares the first one's blind spots .
- Write code agents can change safely: names unique enough to grep, modules small enough to fit a context window, tests that fail clearly .
DHH takes the blunter view: adversarial agent reviews, automated tests, "maybe you spot check" .
Thariq Shihipar's Claude Code tips
On Latent Space, Thariq Shihipar (Anthropic) said:
- Front-load context. Say whether it's a prototype or production code, and where compute is worth spending. Most wasted usage comes from "undo this and redo it" loops .
- Set effort by task type. Use high or max for code review and security, low or medium for UI. Ask explicitly for edge-case coverage on APIs. In software work, effort mostly goes into verification .
- Ask for decision or implementation notes. Most high-effort failures happened when the model considered the right solution and then rejected it .
- Start new projects without CLAUDE.md. Add only failure modes that keep recurring. Old failure logs may over-constrain newer models .
- Claude Mods let you customize how the harness runs and its UI. Example: at the end of each turn, a forked subagent checks whether the task is done and quizzes you. It is cheap because the fork reuses the prompt cache. Another example is a "register assumption" tool that keeps a running list of the model's assumptions .
Split work across models
Geoffrey Huntley's current setup uses Opus or Sol for planning and Kimi for "grunt loops," with Opus/Sol exposed as an oracle tool Kimi can call. He says Kimi is noticeably worse, but its free tokens are unlimited . His pattern is get it working, then launch targeted Sol/Opus refactors to make it good. For design work, he uses Opus .
Theo's TypeScript-to-Rust compiler port had stalled at about 35% of tests passing with GPT-5.6 Sol and about 85% with GPT-6 Astra. He then gave Opus 5.5 /goal finish the port and make it faster. Opus judged Astra's code to be slop and rewrote it from scratch in a new crate. Theo says it made more progress in 10 hours than Astra did in two weeks . Separately, he estimates a $200 Claude Code plan gives about $9,000 a month of Opus usage at API prices .
Rewriting apps in Rust with agents
DHH has a beta of Campfire rewritten in Rust. He calls the code "ugly as sin" and 6× as verbose, and says he never looked at it . He treats it as a "prompt compilation target" . It cost under $10 in tokens on a 20x Max plan and took a few hours: basically one prompt, then a few tuning prompts . His argument: writing web apps in Rust before agents would have been "madness"; now it's "trivial and cheap" .
Smaller items
- Share transcripts, not just prompts. Willison points to Codex's transcript-sharing feature. He built his own version for Claude Code and would prefer a built-in one .
- gpuc (repo): Brendan Long's GPU job queue that needs no sudo. Queue hosts need only SSH, rsync, and the NVIDIA driver. Jobs keep running if the client goes offline, and there is a CLI optimized for Claude (send
!gpuc skill) . He built it so Claude Code doesn't need root . Limits: single-user only, and RunPod is the only rental provider . - At Wonder, a PM can file a bug and Claude Code fixes it from the LangSmith trace .
- Before implementation, use the agent to surface unknowns and clarify preferences—such as schema or call-stack decisions—and build a mental model of both the agent and codebase. Provide enough upfront context, including whether the task is a prototype or production work and how much compute or verification to spend; repeated correction can consume usage, and spoken prompts are useful when they convey more information.
- Calibrate effort by task: use high or max for security and code review, and low or medium for UI; explicitly ask for edge-case testing and verification when warranted. The practitioner says effort in software engineering tends to affect verification more than the core task result.
- Ask for implementation or decision notes: in reviewed eval transcripts, many failures involved the model considering a correct solution and then deciding against it, so notes make those choices inspectable. A harness mod can also keep a visible list of assumptions.
-
Consider starting a project without a
CLAUDE.md, adding only recurring failure modes as they arise. Guidance can differ by model and version, and accumulated instructions may overconstrain a newer model; skill eval plugins can help test whether a skill improves results. - Claude Code Mods customize harness execution and UI. One example forks a checker at the end of a turn to assess completion and generate a quiz for the next prompt; the fork reuses the prompt cache, though running the check after every turn uses additional tokens.
- Claude Code with Opus 5.5 was reported to generate a 30–60-second pure-JavaScript explainer end-to-end—including concept, script, assets, animation, and TTS—in a one-shot run of about 1 hour 20 minutes; the creator reported about $20 in Opus usage plus $3.21 across eight OpenRouter API calls. A second user reproduced the workflow for Friendr.nl in about 1.5–2 hours for roughly $4, adding music/SFX, narration-synced animation, MP4 rendering, and another-model review; the run used a $10-capped OpenRouter key and reportedly needed only minor corrections. A separate motion-design prompt template starts by asking for 8–12 UI states for a shape to transform into (for example, a button, loader, or player); its creator said the video was entirely code, with no After Effects.
-
Opus 5.5 was also used to build TideWater, a browser-based interactive island, in about eight hours with iterative prompts such as
add Xandmake it better; the reported token cost was $1,874.40, or 59% of a Max 20x weekly allowance. The demo includes walking around, interacting with objects, and sailing a boat, so testing it interactively reveals capabilities beyond a video preview. - For agent builders, LangChain Managed Deep Agents 0.8 adds user/agent memory with access policies, HTTP channels, sandbox file APIs, proxy-authenticated sandboxes, and Parallel web search; its smithtune CLI turns traces into post-training datasets, while Trajectories handles deferred tool calls and context compaction. Databricks also reports that engineers stopped reaching for closed models once open-source models were routed to its internal coding agents.
- Build a mental model of what Claude can reliably one-shot, surface unknown requirements before implementation, and learn the domain vocabulary or reference language needed to specify the result more precisely; Thariq Shihipar described this as a core agent-coding skill.
- Put more context into the initial prompt—including whether the task is a prototype or production work, where to spend compute, and what verification is needed—to reduce wasteful undo-and-retry cycles. Shihipar’s rough effort guidance: high/max for code review and security, low/medium for UI, and explicit edge-case verification for API work.
- Ask for implementation or decision notes: the model may consider a suitable approach and choose not to use it, so reviewing its notes can reveal options to request.
-
Start a new project without
Claude.md, then add recurring failure modes as they appear; those failures can differ between model versions, and accumulated guidance may overconstrain newer ones. - Claude Mods can customize Claude Code’s execution and UI. One practical pattern is an end-of-turn forked subagent—which retains the prompt cache—to check task completion and produce a quiz; mods can also register assumptions, and can spawn subagents, parse results, and change the UI.
- For collaboration, Shihipar described using Claude Tag for background work such as code review, security, and starting PRs, and for multiplayer incidents. His example workflow is a channel per project where legal can review exactly what is shipping by discussing it with Claude, without the engineer relaying all the context.
-
The
/eli5plugin uses the prompt pattern “big picture, few words” to explain complex incidents more clearly and reduce text-heavy artifacts.
- At Anthropic, Mike estimated that comprehensive human PR review fell from about 80–90% in January to 5–10% at the time of the talk, citing the volume of generated code and Claude’s ability to find issues. Their replacement is a repeated adversarial loop: multiple Claude reviewers search for problems, assess whether they agree and how severe each issue is, then the coding Claude revises. Human review remains important for architecture, security boundaries, and product judgment; Mike said Claude can make good architectural decisions in a long conversation but may not get the first decision right. He also observed that Claude tends to follow existing codebase patterns, but may not prioritize readability when writing one-off code just to connect things.
- For team context, Mike described keeping much internal work visible within groups and using Slack as a central place to work with Claude; new employees can ask how things are currently done using Claude’s view of company activity. Anthropic also runs a nightly internal “dream” process to incorporate process learnings into organizational memory, though Mike said company-wide learning still had room to improve.
- StrongDM’s “dark factory” rules required code to be routed through a coding agent and prohibited human code review; its experiment explored how to verify agent work and maintain confidence in quality without reading the code.
- To get leverage from capable models, define the goal clearly, specify unambiguous constraints, and provide the necessary tools; Willison says doing this well takes experience and skill, and agent work still requires extraordinary discipline and knowledge.
- Willison uses GPT-6 Sol in Codex and Claude Opus 5.5 in Claude Code as his defaults; he uses GPT-6 Luna for the Datasette Agent and reports it is fast and competent at SQL queries and building HTML and JavaScript. GPT-6 Luna is priced at $0.10 per million input tokens and $0.50 per million output tokens, one tenth of Haiku 4.5’s $1/$5 rates.
- Treat maximum reasoning as a potential cost and latency trap: Opus 5.5 at max hit its 128,000-token output limit before returning an answer on an SVG task; a second attempt failed the same way, with each attempt costing $2.56 and taking nearly 20 minutes.
- For a reusable prototype-to-video workflow, Willison gave Opus 5.5 three kakapo photos and asked it to make an HTML5-canvas pixel-art animation with at least 20 birds. He then used Claude Code with Playwright to load the downloaded HTML, delay clicks until three seconds in, spread them around the canvas, and record a 15-second video.
- For a browser game, Berman recommends explicitly asking the coding agent to check its work as it goes—using screenshots, video, or by playing the game. His Three.js Fall Guys-style demo specified 59 bots across five rounds, took roughly one or two prompts, and still had occasional clipping.
- For a LEGO-generation app, he prompted a split of responsibilities: have the model describe shapes, let a program map them to real LEGO parts and enforce a connected, buildable model, then feed failures back to the AI. The prompt also asked for screenshot checks and instructions with at most four pieces per step.
- In Berman's tests, Sonnet 5.5 beat Opus 5.5 on Terminal Bench 4.0 and was close on other reported coding scores (Frontier Code 52.1 vs. 54.4; Cursor Bench 55.5 vs. 57.8); he said it felt nearly indistinguishable in use and was faster. He reported pricing of $2/$10 per million input/output tokens, versus Opus's $4/$20.
- His Unreal Engine San Francisco build required downloaded assets and took multiple days and millions of tokens; it also overloaded his computer and sometimes stopped working.
- Osmani recommends a risk-based review loop: have agents make a first pass that finds, verifies, severity-ranks bugs, and suggests fixes; give low-blast-radius changes a lighter human review when checks are clean, but carefully review core or sensitive paths, with a person owning merge approval. Anthropic’s automated Claude reviewer runs on nearly every PR and informs engineers without approving changes.
- Before coding, write down what to build, what must not break, and how to tell it worked. Ship only changes you can explain—their behavior, scope, and safety—and disclose in the PR what you did not review; the article gives the example, “agent-reviewed, tests pass, I haven’t read the migration logic.”
- Keep verification independent of the code-writing agent: a second copy of the same model can share its blind spots, so ground checks in a human-written spec, reference implementation, or proof. In Nicholas Carlini’s Claude-built C compiler project, GCC served as a known-good reference for debugging kernel issues.
- Structure code for agent changes: use distinct, searchable names, small modules that fit a context window, clearly failing tests, and clear boundaries so agents can see what a change may affect.
- Anthropic says Claude Sonnet 5.5 runs 30%+ faster and costs up to 30% less for most work; it is priced the same as Sonnet 5, which Simon Willison says it appears to outperform on every benchmark. He also found it nearly as good as Opus 5.5 on some coding tasks, including 3D animation work.
- Thinking-effort choice had a large cost and completion impact in Willison’s pelican-to-SVG test: “max” used 128,000 tokens ($1.28) and failed, while “xhigh” produced a result in 41 seconds for $0.0574. Sonnet 5.5 is also the model on Claude’s free tier; a direct prompt to build a WebGL 3D pelican page produced what Willison called a “solid effort.”
Simon Willison recommends sharing coding-agent transcripts with their embedded context, rather than sharing only the prompt; he points to Codex’s transcript-sharing feature and says he built a workaround for Claude Code but would prefer native support.
Claude Sonnet 5.5 became the model powering Claude’s free tier, making early experiments such as @_re_pete’s fall-foliage simulator made with Sonnet 5 vs. 5.5 accessible to free users . Simon Willison said ChatGPT’s free-tier GPT-5.6 Luna was “a lot less capable” .
Willison’s “Fable class” framing: models such as Claude Fable 5, Claude Opus 5.5, and GPT-Astra 6—and possibly GPT-5.6 Sol—can solve a problem by brute force when its goal is definable, instructions are unambiguous, and the models have access to the necessary tools.
- Fireship reports that DHH said his coding output rose from about 30,000 lines of Ruby per year to about 150,000 lines of code per month using agents.
- DHH reportedly said the Hey email app was rewritten in AI-generated Rust, cutting server CPU and memory usage by 95%.
- DHH advocated providing a CLI so an application can be used without a person interacting with it—a practical interface pattern for agent-driven use.
- Fireship’s takeaway is that developers should prioritize defining problems and designing secure, efficient systems over mechanical code production.
Kent C. Dodds says repeated integration setup is a barrier to building personal software; Kody Koala aims to remove that friction by letting users set integrations up once and reuse them with agents or full-stack apps. He says Opus created a playable game to illustrate the idea: integration-game.kody.codes.
Claude Sonnet 5.5 launched, claimed to be 30% faster and up to 30% cheaper than Sonnet 5 for most work, and was positioned for well-scoped everyday tasks such as bug fixes, documentation, and slides. A follow-up reported improvements over Sonnet 5 on agentic-coding and computer-use benchmarks, including 70.6% on Terminal-Bench 4.0 and 80.1% on OSWorld 2.1.
Geoffrey Huntley’s preferred split is Opus/Sol for planning and Kimi for grunt loops; he says Kimi is worse than Opus/Sol but attractive when unlimited free tokens are available. He also suggests making Opus/Sol an oracle/tool that Kimi can call. For refinement, he follows “get it working, then get it good” by using Sol/Opus for targeted refactoring of working code, and recommends Opus for design tasks.
Jason Zhou open-sourced a Claude /leads-signal skill that monitors keyword mentions, complaints, product reviews, job changes, and 23 additional signals to find hot leads each morning; the skill is linked at GitHub. He pitches it at $0.0002 per signal, compared with $167/month Clay plans.
Alex Albert says Sonnet 5.5 writes clearly, is very fast, and is a major capability jump over Sonnet 5; he found it a “really great model to iterate with,” and compared its feel favorably with Opus 5.5 . Claude AI’s announcement says Sonnet 5.5 is more than 30% faster than Sonnet 5 and costs up to 30% less for most work .
- Anthropic positions Claude Sonnet 5.5 for well-scoped everyday coding tasks such as bug fixes and quick feature iteration; the company claims it is over 30% faster and up to 30% cheaper for most work. Early independent evaluations place it near Opus 5.5 on several leaderboards.
- Sonnet 5.5 launched in Claude Code and is also available through GitHub Copilot in VS Code, Cursor, Factory, Devin Desktop/CLI, Cline, and T3 Code.
Kent C. Dodds argues that when customers use an agent to access a service, exposing a regular MCP server is more efficient than routing the agent through a browser-based WebMCP interaction; he says Cloudflare’s Kitesurf can interact with WebMCP servers but favors direct MCP for agent-facing services.
When coding agents produce low-quality “slop,” Kent C. Dodds advises improving the primitives they work with rather than giving up on the agents or spending effort cleaning up their output; he says he will demonstrate the approach in a video, but the post does not provide the implementation details .
Sonnet 5.5 Is Here. Look What It Can Build.
- For a browser game, Berman recommends explicitly asking the coding agent to check its work as it goes—using screenshots, video, or by playing the game. His Three.js Fall Guys-style demo specified 59 bots across five rounds, took roughly one or two prompts, and still had occasional clipping.
- For a LEGO-generation app, he prompted a split of responsibilities: have the model describe shapes, let a program map them to real LEGO parts and enforce a connected, buildable model, then feed failures back to the AI. The prompt also asked for screenshot checks and instructions with at most four pieces per step.
- In Berman's tests, Sonnet 5.5 beat Opus 5.5 on Terminal Bench 4.0 and was close on other reported coding scores (Frontier Code 52.1 vs. 54.4; Cursor Bench 55.5 vs. 57.8); he said it felt nearly indistinguishable in use and was faster. He reported pricing of $2/$10 per million input/output tokens, versus Opus's $4/$20.
- His Unreal Engine San Francisco build required downloaded assets and took multiple days and millions of tokens; it also overloaded his computer and sometimes stopped working.