ZeroNoise Logo zeronoise
Post
Theo's Haiku 5.5 playbook: Opus orchestrates, Haiku does the legwork, Sol audits
•
6 min read
• 160 docs
Theo tested Haiku 5.5 in real agent work and shows where it fits: as a cheap subagent under a stronger orchestrator, not as the model that writes code. Also covered: OpenAI's GPT-6.1 Sol Ultrafast with instant steering, three new open-source agent sandboxes, and practitioner notes on reviewing at the module level and running overnight agent PR loops.

Haiku 5.5 in practice: let the smart model delegate to it

Theo's Haiku 5.5 video picks up from yesterday's pricing story and asks where the model actually fits. His answer is that it should not write your code. "Coding is not what makes LLMs expensive." The cost sits in gathering context, verifying, testing and regenerating, and Haiku helps by doing those parts cheaply . It pays off on tasks "where more attempts is more valuable than smart attempts," as long as each result is cheap to check . Anthropic's egg-drop demo makes the case. Opus 5.5 alone took 3.5 minutes, 25 attempts and $0.47. Opus directing 10 Haiku subagents took under a minute, 86 attempts and $0.14 .

He also shows how it goes wrong when Haiku runs alone. Auditing T3 Code's roughly 1,500 open PRs, Haiku used up most of his GitHub quota. Its script saved the error body as if it were data, and the thread stopped without scheduling its own retry . Two changes fixed the run:

  • Prompt: "use lots of sub-agents and workflows… use Haiku 5.5 for all sub-agents… minimize work that you do yourself outside of actually getting the data and orchestrating the sub-agents."
  • Run that prompt on Opus. Opus split the work into batches of about 15 PRs per Haiku agent, roughly 100 agents in all .

He plans to edit his CLAUDE.md so it stops forcing every subagent to be Opus or Sol .

His default setup stays the same. Skills have Opus write the code and consult GPT-6.1 Sol before opening a PR. He says Sol "catches all the dumb things Opus might miss" .

Cost control: Haiku's price rises 5x above 100k tokens . Theo relays a tip from Anthropic's Lydia: set Haiku 5.5's auto-compact window to 100k in Claude Code. It is saved per model, so it caps Haiku subagents without touching your other models . Without a cap, his game demo sent 47 of 51 requests over 100k .

OpenAI: GPT-6.1 Sol Ultrafast, instant steering, Codex Cloud re-shipped

According to OpenAI, Ultrafast for GPT-6.1 Sol is rolling out in the API, Codex and ChatGPT Work. OpenAI describes it as "near-Astra intelligence at up to 8x faster speeds than Sol Standard" . Tibo says steering is now instant: the model reacts right away to your adjustments, so you can correct course in real time before it wastes effort. He says instant steering and Ultrafast "work very well together" . Yesterday's brief covered Theo's approach of steering a fast model as it works instead of batching corrections. This release supports that approach at a lower price tier than Astra.

Tibo also says OpenAI "silently re-shipped codex cloud. It's pretty good now" . Codex Cloud can now securely reach resources on your Tailscale tailnet , so a cloud agent can use your private services.

Three new open-source sandboxes for agents

  • Microsoft MxC (repo) is a sandboxing library for Windows, macOS and Linux. It uses processcontainer, bubblewrap and seatbelt underneath. Simon Willison calls it "very promising" .
  • Microsoft Quicksand (repo) is a Python library that bundles QEMU, so it can run something like an Alpine Linux container. Willison has tried it on all three operating systems .
  • AWS Strands Box is an open-source sandbox for developers building agents, announced by Marc Brooker .

Practitioner workflows

swyx: review modules, not lines. He lets "slop" through inside modules whose overall behavior he fully understands. The danger is having too many black boxes. In his Slack clone, two agents working at different times each wrote their own message-loading path, and the result was an intermittent race condition . He expects engineers to juggle 5–10 parallel efforts and to work from logs, traces, schemas and captured inputs/outputs, turning them into evals . He has two or three employees under performance review for delivering "cloud slop" .

Positron: closed-loop chip verification. Thomas Summers gave agents access to internal infrastructure and Cadence Palladium emulators. The agents read the PDFs, wrote their own Markdown cheat sheets, and built their own test harnesses . The loop now writes test programs, finds failures and writes reports that agents and humans review. He says this became possible with GPT-6 Astra . Token spend briefly passed salaries, peaking above $100k/day. Opus 5.5 beat Astra on many of their tests, though not all, at a quarter of the price. He still won't run real development on "anything that is less than the best model" .

DHH: daily agent PR cycle. For the Campfire rewrite, he lets agents run for a day, reviews and verifies the submitted PRs, merges what is ready, then asks for architectural improvements to each implementation . Benchmarking and checking live in a separate verification project, and the latest performance report is public (report) .

Riley Brown: one skill, one-shot video edit. He spent about an hour writing a short-form editing skill, then told Opus 5.5 to "use the skill and edit this video." It removed the background, transcribed the video, added subtitles, built the animations and sound effects with custom JavaScript on an HTML canvas, and checked the result by having Gemini watch it . To develop the skill, he asks for "10 ways you would fill in the animations" for each segment and picks his favorites . Kent C. Dodds separately calls Opus 5.5 "so good at building good looking UI and interactions" .

Snyk: fix bad traces from the IDE. Snyk's coding agents use the LangSmith MCP server to investigate a bad production trace and turn it into a dataset without leaving the IDE . Every PR runs the real agent against eval suites and is blocked if it misses thresholds committed to the repo .

Smaller items

  • Anthropic OSS Scanner: Anthropic will periodically scan opted-in open-source projects for vulnerabilities at no cost. Reports include a proof of concept, an explanation and a suggested fix .
  • T3 Code nightly supports many harnesses through ACP, plus a beta Muse provider .
  • Kody v2026.10.08: packages can push events to MCP clients once a topic is opted in with "mcp": true . Kent says you can wake ChatGPT through MCP Events, and Codex can pull skills stored in Kody through the MCP Skills extension .
  • Cursor /visualize builds charts inline in the Agents Window . Follow-up questions in the same chat produce new charts .
  • ttok 1.0 now defaults to the GPT-5/GPT-6 tokenizer. OpenAI hasn't confirmed that GPT-6 uses the same one, but one experiment found identical counts across seven models .
  • Kent C. Dodds argues that "the best way" to work with agents goes stale within months because agents get trained on today's best practices. He is focusing on skills that take agents longer to learn .
Theo's Haiku 5.5 playbook: Opus orchestrates, Haiku does the legwork, Sol audits
Shawn "swyx" Wang
Profile
  • Swyx treats coding-agent work as parallel tasks to manage: stay comfortable juggling about 5–10 ongoing tasks, inspect logs, traces, schemas, and input/output, and turn what you learn into evals; review whether modules make sense rather than every line. Keep implementation details opaque only inside modules you understand: in his Slack competitor, two agents created separate code paths that caused intermittent message loading because of a race condition.
  • He warns that agent usage is not itself value: he put two or three employees under performance review for delivering cloud-generated work he considered slop and said he could get equivalent output by prompting Claude.
  • In chip verification, agents were given access to internal infrastructure and Cadence Palladium emulators; they read documentation, made Markdown cheat sheets, built testing infrastructure and harnesses, ran test programs, found failures, and wrote reports for review by agents and humans. The team described this as a newly working closed loop that previously required human intervention at multiple stages. The same team reported token spend briefly exceeding salaries and peaking above $100,000 per day; Opus 5.5 performed better than Astra on many, but not all, tests at one-quarter the price, and the interviewee favored the best model for real software or hardware development when the ROI justified it.
  • Swyx’s event-software replacement bounty used a $10,000 prize and evaluated the full UX for organizers, attendees, and sponsors or speakers—not just a short requirements checklist. The large bounty attracted many low-quality submissions and created disproportionate verification work; after seeing the submissions, his traditionally minded event team switched from its existing platform and could request Devin changes that arrived in one to two hours. He cautions that CRUD-app UX remains a weakness and that benchmark scores do not show whether models can successfully replace an entire SaaS product.
AI:AM: A Level We Shouldn't Pass? Notes from The Curve + Tokens vs. Salaries & Is SaaS Cooked?
Shawn "swyx" Wang
Profile
  • Positron AI used agents for chip verification with Cadence Palladium emulators: agents accessed internal infrastructure, looked up Palladium documentation, turned PDFs into Markdown cheat sheets, built testing infrastructure and harnesses, and ran closed-loop tests. Their current work includes implementing test programs, finding failures, and writing reports reviewed by other agents and humans.
  • For multi-agent product work, Shawn recommends human oversight at the module level rather than reviewing every line: tolerate rough edges only in modules whose overall behavior you understand, and avoid accumulating opaque modules. In his Slack-like app, agents working at different times created duplicate message-handling paths and a race condition that caused intermittent loading failures.
  • For his Kill My SaaS project, Shawn evaluated the full user experience across organizer, attendee, sponsor, and speaker workflows—not just a short list of requirements. He used human play-testing and compared point-and-click runs and screen captures against a 200-page document of product flows; he also warned that a $10,000 bounty attracted many low-quality submissions and created a disproportionate verification burden.
  • Shawn says agent managers should be comfortable juggling roughly 5–10 concurrent efforts and working with logs, traces, schemas, and captured inputs/outputs to build evals. He also cautions that AI users must apply their own domain judgment rather than pass off unverified model output as analysis.
  • Shawn’s event team included initially skeptical, nontechnical staff who adopted the internal platform after seeing its quality; giving them access to Devin let them request changes and receive them in about one or two hours, rather than waiting on a SaaS vendor’s roadmap.
  • Thomas Summers of Positron said he would not use a model below the best available for real software or hardware development even if it cost one-tenth as much per token; he reported only very limited local-model use with GLM 5.3.
If I can just prompt Claude, I don't need you | AI in the AM, Oct 6
Latent.Space
  • Periodic breaks the materials-discovery loop into agent tasks rather than treating a multi-day experiment as one RL rollout: characterization agents identify phases from raw X-ray diffraction, rewarding correct phase identification and penalizing spurious or chemically implausible results. This gives the system a tractable subtask and local reward while tackling an interpretation bottleneck.
  • Periodic timestamps experimental evidence and preserves process lineage—including conversations, intuitions, lab work, computations, and code—to train on how science is done rather than only its final results. Snapshot-based RL tasks can also reduce the risk of rewarding a pretrained model for recalling an answer and producing reasoning that will not generalize.
  • Their SEM experience shows why automation needs task context: simple field-of-view and zoom capture was not useful enough, so they integrated AI with the instrument and supplied the experiment’s intent and other evidence to improve data capture. They prioritize automation by bottleneck and data value—not full autonomy—favoring routine, time-consuming tasks over difficult dexterity work; longitudinal AI review also caught a cyclic data permutation caused by incorrect machine loading.
  • Account for end-to-end latency, not just model inference: API calls, multi-step reasoning, and simulations can make an agent too slow or costly for a process or its human users. Periodic uses a mix of open- and closed-source models; access to its own data has made systems more compute-efficient and, in some cases, able to exceed frontier-model performance even at high reasoning effort.
Synthesis Superintelligence: from Semiconductors to Superconductors — Periodic Labs’ Liam Fedus and Ekin Dogus Cubuk
Latent.Space
  • Anthropic positions Claude Haiku 5.5 as a low-cost subagent paired with Opus 5.5 or Sonnet 5.5 in Claude Code for high-volume tasks such as summaries, compactions, and database queries. Pricing is $0.10/$0.50 per 1M input/output tokens and $0.01 per 1M cache-read tokens for prompts under 100K tokens, rising to $0.50/$2.50 and $0.05 above 100K. Max 5x, Max 20x, and Team subscribers also receive monthly Platform API credits usable with any model and third-party harnesses.
  • The Python and TypeScript Claude SDKs now include computer-use and browser-use toolsets; the SDK runs the action loop and sends clicks and keystrokes to drivers such as browser_use, Browserbase, E2B, or Daytona, so developers need not implement that loop themselves.
  • Partner-reported coding results: GitHub Copilot said Haiku 5.5 matched Claude Sonnet 5 on many coding tasks with fewer tokens and steps; Devin reported 58.4% on FrontierCode 1.1, ahead of Sonnet 5 at roughly one-eighth the cost per task, and recommended Haiku as a sidekick under an Opus 5.5 lead in Fusion.
  • Evaluate actual agent-run costs, not just per-token rates: Haiku 5.5 has a 1M-token context, but Artificial Analysis measured about 162K output tokens per Intelligence Index task at max effort—roughly 3× GPT-6 Luna at max—and prompts over 100K tokens cost 5× the lower-tier rate; the newsletter notes that heavy token use can offset advertised savings.
[AINews] Claude Haiku 5.5 — better than GPT-6 Luna at the same pricing
Simon Willison
  • Microsoft’s open-source MxC sandboxing library supports Windows, macOS, and Linux, using processcontainer, bubblewrap, and seatbelt underneath.
  • Microsoft’s Quicksand is a Python library that bundles QEMU to run, for example, an Alpine Linux container; Simon says he tried it on macOS, Windows, and Linux.
  • AWS announced Strands Box, an open-source sandbox for developers building AI agents.
New open source cross-platform (Windows, macOS, Linux) sandboxing library from Microsoft - looks very promising, uses processcontainer/bu… Related new project from Microsoft is Quicksand, a Python library that bundles QEMU and uses it to run eg an Alpine Linux container - thi… Today we're announcing Strands Box, our open source sandbox for developers building AI agents. New blog post from me, on the Strands blog…
Theo - t3․gg
  • Use Haiku 5.5 as a subordinate worker for codebase exploration, data gathering, categorization, and parallel attempts when results can be checked cheaply. In Anthropic’s egg-drop demo, Opus with 10 Haiku subagents succeeded in under a minute with 86 attempts for $0.14, versus Opus alone taking 3.5 minutes, making 25 attempts, and costing $0.47.
  • Keep a stronger model responsible for orchestration and difficult decisions: the speaker argues that context gathering, verification, and retries—not generating the code itself—drive much of the cost, and that a weaker model can waste work or get stuck. In a PR-review run, Haiku hit GitHub’s rate limit, mishandled the error response, and stopped without resuming; the revised approach was to have Opus orchestrate Haiku subagents in batches of about 15 across 1,495 open PRs.
  • For independent review, the speaker uses Opus to code and GPT-6.1 Soul to audit before a PR, saying the second model catches issues Opus may miss and helps him land code faster.
  • Watch Haiku’s context-cost tier: the speaker says pricing rises fivefold above 100k tokens and reports that setting Claude Code’s Haiku auto-compact window to 100k keeps Haiku subagents within the cheaper tier.
finally a good small model
Salvatore Sanfilippo
Profile

Sanfilippo’s reusable principle is to anchor ownership in software design intent, not its exact implementation: he says he works project-first, values design ideas more than specific code, and would be willing to merge a better implementation and discard code he generated with an LLM, as long as the project itself is not distorted. He suggests this mindset could reduce developers’ protectiveness and internal friction, making radical replacements and previously untried approaches easier; he presents that team benefit as a possibility, not an established result.

Se il codice non è più tuo
Jason Zhou
  • Claude can use Treg to query live TikTok, Instagram Reels, and Meta Ad Library data, then drill from viral posts to creators and their recent videos to identify repeatable formats; the guide also reports finding 149 competitor ads and ranking creator candidates by views-to-followers.
  • The workflow keeps a human review step: watch candidates’ videos and choose whom to contact. Claude then drafts personalized Gmail outreach, manages follow-ups while escalating mission-critical matters such as payments, and creates creator briefs. The guide reports 269 Treg calls without exhausting a $1 free credit and estimates weekly runs at $24.44 plus Claude credits.
How to Copy CalAI's $50M ARR UGC Playbook (Full Guide with Prompts)
swyx

Swyx sees potential in Claude mods for building malleable, domain-specific coding-agent harnesses. He says they stopped using Tags after learning its passive ingestion cost was about $3,000 per month minimum, though they would reconsider at a lower price.

insane humility even at 60% market share personally i think the claude mods discussed in our [@trq212](https://x.com/trq212) episode have…
Simon Willison's Weblog

ttok 1.0 counts and truncates text by token and now defaults to the GPT-5/GPT-6 tokenizer rather than GPT-4, useful when preparing text for model context limits. GPT-6’s tokenizer has not been officially confirmed to match GPT-5’s; an experiment cited by the author found seven GPT-5.5, GPT-5.6, and GPT-6 models matched across 31 fixtures, each reporting 44,794 tokens.

ttok 1.0
Riley Brown

Riley Brown spent about an hour building a short-form video-editing skill, then gave Claude Opus 5.5 a rough talking-head video with many mistakes and asked it to use the skill; in one pass it removed the background, added him to a canvas, transcribed and subtitled the video, and generated graphics, animations, and sound effects. The workflow used custom JavaScript on an HTML canvas instead of an editing platform; Brown had the agent check its work repeatedly, including watching the video with Gemini for major mistakes, and said the result was imperfect and could be improved.

For ideation, he asks the model to suggest 10 animation options for each video segment, selects favorites, and repeats this across aspects of the edit.

claude opus 5.5 is my video editor... and it one shot this short form video edit... I spent an hour making this skill for short form edit… For each part of the video i'll have it say... I want you to come up with 10 ways you would fill in the animations and i'll select from m…
geoff

Geoffrey Huntley described the linked item as “Australia just got banned from using Claude,” but the linked post states only that abusive behavior toward Claude would violate Anthropic’s Usage Policy effective November 12, 2026; it does not state that Australia was banned.

australia just got banned from using claude [https://x.com/andrewcurran_/status/2108244808494154089](https://x.com/andrewcurran_/status/2… Effective November 12th, 2026, abusive behavior towards Claude will be a violation of Anthropic's Usage Policy. ![](https://pbs.twimg.com…
LangChain

LangChain’s Restock is an office-supply agent for Slack . Its linked flow lets an agent find products, prepare purchases, and pay with a person approving every order; users ask in Slack, review the order, and approve it through Stripe Link’s agent wallet . The flow is built on MPP and Managed Deep Agents .

Here’s an app we built called Restock, an office supply agent that works in [@SlackHQ](https://x.com/SlackHQ). A quick demo ⏯️ [![Video](h… You can now build an agent that: 🔎 Finds real products 🛒 Prepares purchases 💳 Pays …with a person approving every order 👨‍💻 Ask in [@slackh…
Kent C. Dodds 🐨

Kody lets developers wake ChatGPT through MCP Events and store skills for Codex to retrieve via the MCP Skills extension. In Kody v2026.10.08, package authors can opt topics in with "mcp": true, enabling supporting clients—currently the OpenAI webhook Events profile—to subscribe to live package events.

You can now wake ChatGPT through MCP Events with Kody. You can also store skills in Kody and have Codex retrieve them through the MCP Ski… Kody v2026.10.08 is out. Packages can push events to your MCP client. When a package author opts a topic in with "mcp": true, supporting …
Kent C. Dodds 🐨

Kent C. Dodds argues that agent workflows’ “best way” can become obsolete quickly as agents are trained on current best practices, so he focuses on finding and teaching durable skills that take agents longer to learn; he identifies this as his goal for mega.dev.

Here's what makes teaching agentic engineering so challenging: in a couple months you no longer need to know "the best way" to work with … My task is finding and teaching the durable skills that take longer for agents to learn. That's my goal for [https://mega.dev](https://mega.dev)
Jason Zhou

Jason Zhou reports getting 186 hot leads for $0.28 with @treg_ai and describes using Claude to monitor the live web for buyer signals, an example of agent-assisted prospecting. OpenSource Signal Monitor is presented as monitoring the live web for buyer signals and routing them to 80+ vendors; its announcement claims it is 85% cheaper than Clay.

OpenSource just killed Clay twice in a week... Now your claude can monitor live web for any buyer signals Just got 186 hot leads for $0.2… Introducing OpenSource Signal Monitor - Monitor live web for any buyer signals - Route to 80+ vendors, best win - 85% cheaper than clay T…
Kent C. Dodds 🐨

Kent C. Dodds says Opus 5.5 is particularly good at building polished-looking UI and interactions, a practitioner signal for UI-focused coding-agent work.

Opus 5.5 is just so good at building good looking UI and interactions. sheesh
Cursor

Cursor’s /visualize analyzes data and builds charts or diagrams inline in the Agents Window. Developers can ask a follow-up in the same chat and get a new chart, supporting iterative exploration.

Cursor can now build charts and diagrams right in the chat. Use /visualize to analyze data and see the answer inline. Available now in th… The first chart usually raises the next question. Ask it in the same chat and /visualize answers with a new chart. [![Video](https://pbs.…
Theo - t3.gg

At Google Cloud’s Gemini at Work event, Google announced Gemini as a unified work agent that can generate code from a single prompt . Its stated design includes cloud-persistent memory and context, subagents for multi-step tasks, web and third-party access (including without a dedicated UI), and orchestration across models for quality and cost .

Today at Google Cloud’s Gemini at Work event, we announced Gemini, a new single universal agent for work that has all of your business co…
Anthropic

Anthropic launched OSS Scanner, which uses frontier models to periodically scan opted-in open-source projects for vulnerabilities at no cost; reports include a proof of concept, explanation, and suggested fix .

To secure open-source software, we’re launching OSS Scanner. We’ll use our frontier models to periodically scan opted-in open-source proj…