We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
tsc-rs: an agent-built TypeScript compiler, and what the model choice cost
Theo released tsc-rs (also called ts-rust), a full Rust rewrite of the TypeScript compiler, type checker and LSP. He describes it as an open-source drop-in replacement for tsc (repo) . Agents had been working on it for five months. In his words: "I burned ~$400k of Codex tokens and got nowhere. Burned ~$20k of Opus and got there in 2 weeks." He also says he has not read a single line of the code . He adds that the Opus usage equals about 10 weeks on the $200 Claude subscription .
The port matched its test oracle, upstream TypeScript, closely. Of the first five issues filed, he says four were real upstream TypeScript behaviour that the port reproduced faithfully . He is open about its limits. bun check shipped the same day and is "probably the right choice for most apps using Bun," and tsc-rs is only faster when you use Effect TS checks. He prefers tsc-rs for its drop-in compatibility, WASM readiness and built-in Effect mods .
Mario Zechner and Armin Ronacher explained in a separate conversation why projects like this work. If you have "an Oracle that can tell the LLM if what it did is correct or not, then you basically won." Bun's test suite is enough for agents to drive a port with occasional steering . Ronacher had agents implement every format in a serialization library from sample TOML, CBOR, MessagePack and JSON files, then told them to "fuzz the hell out of it." That found many bugs, and an LLM can reason about cases a fuzz generator misses .
Haiku 5.5: Anthropic's pitch is a cheap subagent
Anthropic says Haiku 5.5 is in Claude Code and the Claude Platform, costs about 75% less to run than Haiku 4.5, and suits use as a subagent under Opus 5.5 or Sonnet 5.5. It suggests "summaries, compactions, or database queries" . Addy Osmani says his team "loves it as a subagent alongside Opus 5.5," and notes it has an adjustable effort setting . Cursor has it under Settings > Models. Pricing is $0.10/$0.50 per million input/output tokens, rising to $0.50/$2.50 above 100k input tokens. Sonnet 5.5 cache reads dropped from $0.20/M to $0.10/M .
Simon Willison found catches that change the routing math:
- Up to 100k tokens it costs exactly the same as GPT-6 Luna. Above that, its price rises 5x, while Luna's goes up only at 272k and only to $0.20/$0.75. For long-context jobs, Luna looks like the better deal .
- The new tokenizer used about 1.25x as many tokens as Haiku 4.5 on the same long prompt, which amounts to a hidden price increase .
-
You can't turn reasoning off. The default is
medium. To try it:llm install -U llm-anthropic,llm anthropic refresh, thenllm -m claude-haiku-5.5 ... -o thinking_effort low. - Max and Team subscribers now get monthly API credits: $100 for Max 5x, $200 for Max 20x, and up to $500 pooled for Team. They don't roll over. You can turn off auto-reload so requests stop when the balance runs out .
ThePrimeagen offers an early counterpoint. Dropping Haiku in for Luna, both with reasoning off, took a one-line config change. Haiku "performed significantly worse" and was slower on both passes and failures, and he doesn't yet know why . Separately, he found OpenAI's new decision model much faster and more accurate than his previous setup at finding click targets in Omarchy QA screenshots . Benchmark your own automation before switching.
Cross-model adversarial review takes one sentence
DHH's technique: when you're driving from Codex, say "Review this with claude", and when you're in Claude, say "Review this with codex". The models know how to start a review through the CLI, take turns and settle an argument. "No magic" . He describes the rest of his setup as "any harness, multiple agents concurrently, barely any skills, and using adversarial reviews" . Theo runs a version of this. Claude spins up Sol subagents and calls Codex through T3 Code to review its work, because he finds OpenAI models "a bit more thorough with their analysis." He still calls the Claude plan "absolutely mandatory" for shipping serious engineering work . On the "nerfed" $200 Codex plan, he argues that cost per task matters more than tokens per dollar. In his own Terminal Bench run, 6.1 Sol at x-high roughly tied Opus 5.5 at about one-thirteenth the cost per task .
Audit your agents for logic they keep reinventing
Theo says his agents wrote over 200 bad "watch PR" scripts. His audit prompt: "audit my history with Claude Code, Codex and other agents on this machine. Look for every time I asked for a PR to be watched or babysat." Then ask how many times the logic was reinvented, how many versions had visible flaws, and roughly how many tokens and dollars were wasted . The production fix in T3 Code took about 30 tries. It replaced CLI calls with direct API calls and switches between GraphQL and REST depending on which uses rate limits more efficiently, with fallbacks. That cut rate-limit usage by over 75% .
Huntley: Nix as the shared environment for agent sandboxes
Geoffrey Huntley argues for a single devenv.nix as the one source of truth for laptops, CI and ephemeral agent sandboxes. His example stanza provides Rust, Postgres, a rustfmt hook and prek "for agent backpressure" . He develops on NixOS and explicitly tells agents to use sudo, counting on rollback. With runNixOSTest, he puts the whole OS, including multi-machine networking and firewall rules, under test . He also uses an overlay to strip force-push out of the Git binary inside agent sandboxes (nix-demo) . The tradeoff he names is that incremental caching is weaker than in Bazel or Buck2 . Separately, he says he is dropping Opus 5.5 because Claude Code's classifier "is too paternalistic and breaks my flow/development loops" .
Smaller items
-
Deep Agents skills: you can now bind tools to a skill with
metadata.include_tools. A bound tool stays out of context until the agent reads that skill . Apps can passpinned_skillsso a skill like/meeting-prepis loaded before the first model call . Settingskills_metadata=Nonemakes the next run rescan the skill library . - Sourcegraph Deep Search lets Claude Code search code you haven't checked out. Their demo finds deprecated packages across Kubernetes .
- OpenAI integrated its desktop app with chat.openai.com in 30 days—an effort Sottiaux said would traditionally take 6–12 months—and he said Astra wrote most of the code.
- OpenAI engineers used Astra to develop the next-generation inference stack and create a version eight times faster in a short time; Sottiaux described engineers collaborating with the model and said improved infrastructure, including Codex, lets teams build faster. This is a model-assisted infrastructure loop that can speed up subsequent development.
- Before a live keynote demo, Sottiaux’s Dot alerted him that production was down, connected the issue to the demo script, and offered to look into a fix; they did not resolve it in time.
- Coding agents work best when correctness has an objective pass/fail check: a large test suite can guide a port even if coverage is incomplete, and humans can steer at checkpoints. For serialization work, giving an agent sample files for formats including CBOR, MessagePack, and JSON, then asking it to fuzz the implementation, exposed bugs; an LLM can also suggest cases a fuzzer may miss.
- Keep humans responsible for issue triage, architecture, and system-level performance: Mario said agents can handle small fixes, but he still reviews reported issues and sees design or architectural work as requiring human judgment. A practical boundary described for Pi was to establish a minimal API first, then have an agent build components against it with rendering tests; agents can help with local performance fixes but may miss broader architectural causes.
- Pi Durable is designed to let long-running agents resume after a process crash without manual continuation or losing tool results, while preserving application state and conversation history. Its task pattern persists an effect’s intent before execution and its result afterward; external effects need idempotency support to recover safely if a crash occurs during execution. Application state can be stored in versioned documents associated with transcript positions, so an agent can recover the state appropriate to a point in the conversation.
- Treat coding agents as parallel workers, with human oversight at module boundaries: Wang says practitioners should be comfortable juggling 5–10 ongoing tasks and using logs, traces, schemas, and input/output data to build evals; he allows implementation slop only in modules whose overall behavior he understands. In his Slack-clone project, agents working at different times created duplicate message paths and an intermittent race condition when context was lost, illustrating the risk of accumulating opaque modules.
- Don’t let AI-generated analysis replace firsthand judgment: Wang favors depth of insight over breadth and describes an employee reporting 14.7% week-on-week video growth without knowing why, after relying on Claude rather than watching the videos; he says people must use their own thinking and domain expertise to assess AI output.
- Evaluate generated software through complete user workflows, not by whether it looks finished: his replacement-app evaluation considered organizer, attendee, sponsor, and speaker perspectives, role-specific logins, and application flows; he says frontier functionality still needs human playtesting. He also describes comparing implementations with point-and-click tests and screen captures against a 200-page flow document.
- Wang gave nontechnical event staff access to modify code with Devon; staff could request changes and get them in roughly one or two hours, and initially skeptical, spreadsheet-oriented staff switched after seeing the submission quality. He contrasts this with a SaaS vendor’s uncertain roadmap timing.
- As a personal, unmeasured estimate, Wang said AI assistance now lets him do the work of roughly four or five copies of himself, noting that work he did with 12 prompts the prior night would otherwise have taken weeks.
- Periodic describes software work moving from GitHub Copilot and early ChatGPT as assistants to Codex as automation improved; it says few of its engineers now write code as before and expects a similar autonomy progression in research tools, potentially shifting from assistance toward pricing outcomes.
- For deciding what to automate, use humans and automation together to find bottlenecks, then prioritize routine tasks that consume substantial staff time and are easy to automate rather than spending months automating tasks where people have a fine-motor-skill advantage. The goal is high-quality, varied data—not full automation for its own sake.
- Periodic builds RL environments from accumulated experimental or computational histories, and can create additional environments around individual tools or tool subsets. It retains conversations, intuitions, lab actions, computations, and code, aiming to train on the process of doing science rather than only its final outputs.
- Periodic uses both open- and closed-source models; it says access to data unavailable to others can improve compute efficiency and, in some cases, let its systems go beyond frontier models. It also cautions that spending substantial compute on noisy data will not produce good results.
- Treat coding-agent work as an iterative design, coding, and verification loop—not a one-shot prompt. Stay engaged in specifying features and UX, checking results, and requesting small changes; distinguish your active steering time from time the agent runs on its own.
- For a CAD prototype, he first used a Fable 5.1 session to explore SDF feasibility and questions such as real-time editing, face extrusion, and holes. He says LLMs write C better, asks for no dependencies, builds a small knowledge archive from relevant papers, and starts with a minimal implementation and tree-based grouping designed to support later operations.
- Make the project testable by the agent: provide ways to create and manipulate objects and capture screenshots; test fillets on random solids from multiple viewpoints and check screenshot consistency. For responsiveness, he uses cached meshes for movement and low-resolution rendering while moving, followed by a GPU render after 500 ms.
- For software intended for AI use, he advocates a well-documented, simple direct API—such as socket commands rather than MCP mediation—and useful outputs for the model, including depth-shaded or wireframe views and measurements.
Sanfilippo used AI as an editor for a story published in Urania—not to rewrite or polish the prose, but to assess its narrative weaknesses—and said the approach worked very well.
- Claude Haiku 5.5 costs $0.10/$0.50 per million input/output tokens through 100,000 tokens, then $0.50/$2.50; its tokenizer uses about 1.25× as many tokens as Haiku 4.5 for the same long prompt. GPT-6 Luna matches Haiku’s price through 100,000 tokens and raises its price only at 272,000, making it a better deal for longer-context workloads.
-
llm-anthropicnow supports new models without a plugin release for each one: update it, refresh the model list, then invoke Haiku throughllm, for example with-o thinking_effort low. Haiku defaults to medium reasoning and does not allow reasoning to be disabled. - Anthropic is adding monthly Claude Platform API credits for Max and Team subscribers: $100 for Max 5x, $200 for Max 20x, and up to $500 pooled for Team; credits do not roll over. Auto-reload can be disabled so requests stop when the balance runs out.
- In his own Terminal Bench 4 tests, Theo says 6.1 Soul’s max run scored best with its native harness and its x-high run tied Opus 5.5; he ran the models through Codex and Claude Code and argues that 6.1 Soul’s cost per task was roughly 13× lower in this comparison. His takeaway is to compare completed-task cost under the relevant harness, not just token or subscription value.
- Theo’s cross-model review workflow uses Claude to launch Soul subagents to review its work, and Claude calls Codex through T3 Code to review Claude’s work. He says OpenAI models have been more thorough in his analysis and deeper issue-finding, while he considers the Claude plan the better choice for regularly shipping serious engineering work.
-
Use
devenv.nixas the shared source of truth for toolchains and dependencies across local development, CI/CD, and ephemeral agent sandboxes; Huntley’s example enables Rust, Postgres,prek, and arustfmthook withdevenv up, avoiding separate environment configurations that drift apart. -
Huntley runs coding agents on NixOS and explicitly allows them to use
sudo, relying on the system’s ability to roll back changes; his loop-engineering approach tests the whole OS, usingrunNixOSTestto assert multi-machine networking, firewall rules, and application interoperability before deployment. - In agent sandboxes, he customizes Git with a Nix overlay that removes force-push functionality, then tests the restriction in a NixOS VM; he presents overlays as a way to customize software throughout the dependency stack.
- Trade-off: Nix’s derivation store can complicate incremental caching, an area where Huntley says Bazel and Buck2 excel, though he still considers Nix the best bang for the buck.
- Huntley proposed moving beyond conventional IDEs toward tools designed for collaboration with many LLMs; he says agentic coding products were getting strong results without language-server AST or other symbolic-context injection, so that context may not be essential to agent workflows.
- He cautions that a stronger model alone cannot verify correct production behavior under suboptimal conditions such as flaky networks; agent-built software still needs verification that accounts for runtime conditions.
Reacting to a post naming GPT-6 in Chat, ThePrimeagen said it and “Astra ultrafast” seemed to be creating a “speed addiction epidemic.” The quoted announcement reported a new high of 40M active users across Codex and ChatGPT Work and a banked reset in paid accounts.
- Stacklok’s open-source Mecatl harness, begun in June, is designed to run agents in the cloud: its agent loop is independent of the client, model provider, state store, and execution environment, while tool calls and bash, session management, and memory are separated into components that can be centrally managed rather than kept in local JSONL files.
- Stacklok’s open-source ToolHive manages MCP servers; its Kubernetes-based offering includes a gateway, registry, and operator helper, with input/output authentication support. Stacklok’s AI Gateway handles access control, budgets, reporting, and provider routing, but does not currently select models by task; the company favors semantic model selection within the harness, where more context is available. The gateway was not yet open source and was described as planned.
Theo disputes the $132M/year calculation for a Cursor profile reporting 1.5T tokens: he says it overstates costs by applying an inaccurate rate to usage that includes cached tokens, and estimates API costs of $100K–$500K/month instead. He estimates 1.5T tokens would cost about $800K with expensive models or as little as $100K with cheaper models and more cache hits. He says his own 33.5B-token usage “came out to $0.56/m”; in a separate reply, he reports a blended $0.46 for real-world Opus 5.5 use, nearly 20× below the $8/M assumption. Theo cites a 30× drop in cost at a given intelligence level over five months and projects that a $100K/month workload could fall to $3K/month in a few months and $100/month next year if the trend continues.
For model-generated components or migration plans, separate review into two filters: use experience-based taste to reject weak drafts quickly (potentially most of them), then use judgement to choose what to deliver by weighing quality against time, risk, and user or team costs.
Theo says agents worked on tsc-rs, a Rust rewrite of the TypeScript compiler, type checker, and LSP, for five months. He reports spending about $400k in Codex tokens without getting there, then about $20k on Opus to complete it in two weeks; he says he had not read any of the code. He describes tsc-rs as an open-source drop-in replacement for tsc, available now.
Addy Osmani says his team uses Claude Haiku 5.5 as a subagent alongside Opus 5.5; Haiku 5.5 is about 75% cheaper to run than Haiku 4.5 and has an adjustable effort setting. Claude says it is available in Claude Code and recommends it for high-volume, cost-sensitive subagent tasks such as summaries, compactions, and database queries, alongside Opus 5.5 or Sonnet 5.5.
ThePrimeagen tested Claude Haiku 5.5 against OpenAI GPT-6 Luna, both with reasoning set to none, on an automation task; swapping in Haiku required only a one-line config change, but it performed significantly worse and ran significantly slower on both passes and failures. The reason for the difference was still unclear.
Haiku 5.5 is described as a significant step up from Haiku 4.5 for coding, computer use, and knowledge work . Alex Albert says Haiku 5.5 is much faster and 75% cheaper, noting that Haiku 4.5 came out on October 15, 2025, less than a year before .
ThePrimeagen tested OpenAI’s new decision model on Omarchy QA results, using the same prompt and image, and found it generally faster and more accurate at identifying screenshot targets to click; the post gives no quantitative comparison. In one example, it correctly inferred the terminal background from an image and clicked it.
Geoffrey Huntley says /z80 was productionized as an MCP server and a paper was published; he characterizes it as using an LLM to clone software and intellectual property (“running the LLM like a Bitcoin mixer”).
Does the $200 Codex plan suck now?
- In his own Terminal Bench 4 tests, Theo says 6.1 Soul’s max run scored best with its native harness and its x-high run tied Opus 5.5; he ran the models through Codex and Claude Code and argues that 6.1 Soul’s cost per task was roughly 13× lower in this comparison. His takeaway is to compare completed-task cost under the relevant harness, not just token or subscription value.
- Theo’s cross-model review workflow uses Claude to launch Soul subagents to review its work, and Claude calls Codex through T3 Code to review Claude’s work. He says OpenAI models have been more thorough in his analysis and deeper issue-finding, while he considers the Claude plan the better choice for regularly shipping serious engineering work.