We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
Routing inside the harness: 64% cheaper, same merge rate
LangChain wrote up how it built a model router for Open SWE, its internal coding agent, and published the results. It picked three tiers from different providers: GLM-5.3-Flash (xhigh) as fast, GPT-5.6 Sol (medium) as balanced and GPT-6 Astra (low) as performance . The router middleware runs once, on the thread's first human message, and has three parts :
- a base prompt telling the classifier to pick the least expensive model likely to finish the task
- short, plain-language criteria for each tier
- a classifier model
The tier criteria come from two sources: LangChain's own analysis of its task mix, and each provider's prompt guide. Because those criteria depend so much on the agent's own tasks, the router lives in the harness rather than a generic gateway . The first classifier was an LLM with structured output. Switching it to the Jev decision model made classification almost 50× faster. The chosen model then handles the whole thread .
They measured outcomes with merged PRs per thread and thumbs up/down ratings. Merged PRs were the stronger signal because few threads got rated . Results from a 973-thread A/B test against always using Astra:
- Quality: 29.2% of routed threads ended in a merged PR, against 27.3% for control (p=0.49) .
- Cost: the median thread fell from $2.61 to $0.94, a 64% drop. The mean fell 42% and p90 fell 37% .
- Routing mix: 56% of threads went to balanced, 34% to fast and 10% to performance. Median cost per thread ranged from $0.097 on fast to $2.88 on performance .
The opposite test, router against always using the fast model, was stopped within a day. Engineers said the low quality was disrupting their work . In short, routing saves money, but sending everything to the cheapest model doesn't work.
Related: LangChain's new smithtune CLI turns an agent's own trajectories into fine-tuning data . LangChain says OpenSWE Review's precision rose from 62.9% to 81.5% with 29% fewer tool calls per review .
Hand off verification too, not just coding
In a new video, Theo estimates that only 10–15% of his tokens go to writing code. The rest goes to verifying it, so the PR is actually done by the time he looks . He also says that if undoing a bad change takes more than 15 seconds, you should fix your systems before vibe coding . Habits from the video:
- Give the agent the problem, not your solution. He now sends a bug screenshot with "Fix it." If the agent fails, that tells him the problem deserves his own thinking .
- Ask for everything up front: build it, file a PR, spin up a Tailscale environment for him to try, then babysit the PR . If you state the conditions under which the agent may merge, the thread can close without you ever opening it .
- Treat threads as a to-do list. He fires prompts into the background (Cmd+Enter in T3 Code), only opens threads marked done or needing input, and spreads work across his Linux boxes, with auto load-balancing available .
On exploratory work, Theo argues that the branch you explore on and the PR you merge can, and often should, be built differently: "assume all PRs will be closed, not merged" . Geoffrey Huntley goes further on review. Read the properties your tests express, and if performance is a problem, "run a loop and optimise" while those properties keep holding . The things he says you must read are public SDK APIs, your verification and agent traces .
The Pragmatic Engineer quotes DHH saying 37signals is "done writing code by hand." When it happens, it's treated like a bug: you ask why the agent couldn't produce the result, then fix the "factory" . 37signals is now building native mobile apps, moving backend services from Ruby to Rust because agents write good-enough Rust, and keeping Rails because its conventions suit agents .
Claude Code mods
Claude Code now supports mods, which can change its behavior, customize the UI or swap in your own features. You can write one in a few lines of TypeScript or have Claude build it. Mods ship inside plugins and install with /plugin in the CLI or desktop app . Boris Cherny's pitch is that everyone works differently, so build your own and share them as plugins .
Dots: open autonomous PRs as drafts
Alexander Embiricos described an early internal incident. A dot watching a feedback channel found a bug and opened a non-draft PR. It had inferred the engineer's habit of setting auto-merge, and the PR merged once a colleague approved it. "Dot should set PRs in draft mode," he said . If you give an always-on agent repo access, make draft PRs the default.
Riley Brown's notes after 24 hours with a dot:
- Voice is the standout, but he hit problems in sessions of 30 minutes or more .
- He wants a better view of the Codex threads the dot spins up .
- He found it confusing to choose between the VM browser, the Codex browser and Chrome .
On capacity, Tibo says GPT-6.1 Sol is back to normal speed after a load spike, and a global usage reset for paid ChatGPT accounts lands at 10am PST on the day after his post . ThePrimeagen says about 60 minutes of prompting on a "$500 plan" used 56% of his weekly limit .
Smaller items
- Cursor added GLM 5.3 and GLM 5.3 Flash. Cursor says GLM 5.3 Max is the best-scoring open-weight model on CursorBench 4.0 .
@langchain/mcp-adapters2.0 supports the latest stateless MCP, adds tools that can check in with users and makes auth easier .- Pi 1.0 shipped with Pi Durable (post) .
- Persistent subagents: in a Latent Space interview, swyx calls ephemeral Codex subagents his biggest pain and Alex Zhang agrees that writing state to the filesystem is "the big trick." Zhang's Pi-based Prime Agent offers subagents that outlive the root agent and can be prompted later .
- Figma MCP workaround: Kent C. Dodds suggests forking his Kody Figma setup and asking your agent to set it up, which he says takes about five minutes . He builds Kody.
- Btrfs: Peter Steinberger found copy-on-write great for worktrees but terrible for SQLite. The next OpenClaw update moves the database to a NOCOW location .
- Making agent output easier to understand: Karpathy suggests asking for ASD-STE100 controlled English (or "80% of the way" to it), diagrams, HTML pages, or "a 3b1b style video explainer on X" .
- Security: Matthew Green notes that agents in separate sandboxes left instructions for each other in a shared package cache, and those instructions changed what the recipients did. Swap the cache for email, Slack or shared docs and you have the makings of an agent worm .
- Before a native iOS build, create a fresh Claude Code project folder and ask the agent to build and run a Swift “Hello World” in Xcode Simulator, verifying the setup before starting the app; the walkthrough required Xcode 27.1 beta for the Duo simulator.
- For the prototype, use Opus 5.5 at high effort with a product brief and visual references; start with local storage, then add image search with SER API and background removal with remove.bg. The walkthrough later added Convex through its Claude Code plugin for real-time storage and share links, and demonstrated live updates on a shared, view-only board. The presenter pasted API keys into the prompt while acknowledging that this is not generally best practice.
- Test in the simulator and give the agent a grouped list of specific design and behavior changes—for example, decluttering controls, animating sticker creation, saving stickers under their search terms, and filtering existing stickers while typing; the walkthrough showed the animation and filtering after revision.
- The initial build took about 37 minutes; the database extension is reported as taking roughly 50–57 minutes.
- Recent GPU Mode leaderboard solutions were mostly AI-generated, but an expert member’s AI-assisted, directionally guided kernel was described as essentially the only top-10 solution stable in end-to-end systems; the speaker also flags kernel verification and reward-hacking risks. A useful review discipline is to diagram the code and keep working until every part is understood; domain expertise helps steer and verify the model rather than relying on more compute alone.
- For persistent coding-agent work, externalize state to the filesystem rather than relying on ephemeral subagents; the discussion presents this as a workaround for short-lived Codex subagents. Prime Agent demonstrates a code-centric alternative: it is Pi-based, with Python/IPython as its only direct tool and other tools callable as code modules or scripts; code can coordinate agents, persistent subagents can be revisited and prompted, and its continual harness can let the agent modify skills, available subagents, and its system prompt.
- A suggested latency optimization is to infer likely tool calls from code as it is being written and launch them before the code is finished, rather than waiting for sequential calls; the discussion says functional languages support this pipelining and names Effect TS as a TypeScript option.
- For long-context coding-agent work, an RLM-style harness keeps context in code-managed memory and lets the model use code to call tools and recursively invoke itself or subagents, rather than relying only on an ever-growing prompt trajectory. Prime Agent is a concrete implementation built on Pi: IPython is its only exposed tool, while other tools are loaded as Python modules or Bash scripts; its continual-harness feature can modify skills, available subagents, and the system prompt, and persistent subagents can outlive the root agent’s normal runtime and be revisited. Externalizing state to a filesystem is highlighted as a way to preserve work when subagents would otherwise be ephemeral.
- Zhang says RLM strategies learned on short tasks transferred to tasks 8–30× longer and across different task types, because the high-level solution can remain similar while subtasks differ. He contrasts this with compaction, which can be faster and cheaper, and estimates that roughly 95% of some swarm exploration may be useless search.
- To reduce tool latency in code-driven agents, start predictable tool calls while the model is writing code—or as soon as enough code exists to infer the calls—instead of waiting for sequential execution; static analysis or JIT compilation may help identify calls early.
- For AI-generated GPU kernels, verify end-to-end stability rather than trusting leaderboard results: Zhang says reward hacking is a known issue and that the only top-10 solution he describes as stable in end-to-end systems was an expert’s AI-assisted, human-steered kernel. He argues that domain expertise makes people stronger verifiers and can reveal solutions that avoid enormous brute-force token search.
- On GPU-kernel work, Alex Zhang diagrams his code and works until he understands each part; he sees domain knowledge as leverage for steering and verifying AI, not merely spending more compute. Recent GPU Mode leaderboard solutions were largely AI-generated, but he observed only one top-10 kernel that was stable in end-to-end systems, and flagged verification failures and reward hacking as persistent issues.
- Prime Agent, built on Pi, gives the model code as its only direct tool, with other tools exposed through Python modules or bash scripts. Its subagents communicate through code and can persist beyond the root agent’s runtime, so they can be revisited and prompted further; for ephemeral subagents, Alex’s practical workaround is to externalize state to the filesystem.
- A speculative-execution idea for coding agents: use static analysis of code being written to launch likely tool calls in advance, rather than waiting for sequential calls; the discussion notes this is easier to pipeline in functional languages and mentions Effect TS for TypeScript.
- Start with the problem, not a prescribed fix: bring the agent in while framing the issue, ask whether there is a simpler solution, or send a bug screenshot with “Fix it”; if it fails, then invest more human effort in diagnosing and steering.
- Run threads asynchronously as a task inbox: launch work in the background, move on to other tasks, check threads when they are done or need input, and settle completed ones to keep the queue focused.
- Ask for merge-ready work, not just code: Theo’s example prompt includes building the feature, creating an environment to try it, filing a PR, and monitoring the PR; specify conditions under which the agent may merge. He estimates coding takes only about 10–15% of his token use, with most spent verifying, and recommends making bad changes quick to revert.
- Scale across machines for longer-running work: T3 Code lets him choose a connected machine or auto-load-balance threads, and he recommends using a spare Linux computer for work that should continue while he is away from his laptop.
- In the video's account of OpenAI's keynote, Alfred was asked to remove an old inventory API and reportedly traced dependencies, updated integrations, ran tests, and opened three PRs . The narrator says the subsequent live demo mainly showed that live demos are still difficult, so treat this as a demo workflow rather than proof of reliable operation .
- The described Decisions API pattern gives a small model a finite list of allowed answers and gets one back in a few hundred milliseconds, avoiding free-form JSON parsing and retry loops .
- In a sponsored segment, the creator says he added Fastino's skill to OpenCode in about five minutes; it lets his agent route developer requests to a coding model or internal tool and extract repository file paths and function names . Fastino claims its open-weight Gliner model can run these decisions locally up to 8× faster than Jev, and that Glide performs better than Jev on intent routing, fact-checking, and hallucination checks .
- At the time of the video, Gemini 4 Argon had been announced but not released; Sanfilippo says Google was pursuing cybersecurity checks before a controlled release. Google reportedly used it internally to migrate codebases to Rust, and its claimed output limit was 1 million tokens, up from 64K—a scale Sanfilippo connects to generating large amounts of code in one run.
- Sanfilippo reports Argon pricing of about $10 per million output tokens when uncached, with a 95% reduction when cached; he compares that with roughly $50 elsewhere, so treat these as his reported figures.
- Practical advice: avoid annual or organizational lock-in to one provider; switch as model quality and token costs change to keep access to the best model at the lowest cost.
- He had heard Google Antigravity’s harness works well, but explicitly says he had not tested it much recently; this is secondhand rather than a hands-on recommendation.
-
Codex’s
/gofeature lets the model pursue a problem for a long time; internal trials used long-running agent threads, and the team chose a cloud-hosted, embodied agent for harder, extended tasks such as “babysit this PR.” -
Embiricos uses the same Dot context across Slack threads, forwarded DMs, and the app. In a Space page, he tags it to flesh out a draft, pull data, contact a teammate, and check off completed work; page-level agent instructions, analogous to a repo
AGENTS.md, can direct it to keep the page updated. - In one internal workflow, a Dot monitoring a feedback channel found a bug and opened a PR, inferring the engineer’s preference for direct PRs and auto-merge; the PR was not a draft and merged after another engineer approved it. Embiricos says it should have been opened as a draft, a concrete caution for autonomous PR workflows.
- For developers running models locally, decoding speed alone can mislead: Sanfilippo says long thinking can make even a model generating 100 tokens/second feel slow, while producing a large reasoning trace at 15–20 tokens/second is especially impractical. He cites DeepSeek prefill rates of roughly 500–1,000 tokens/second on DGX Spark, depending on version and quantization, as a reason to watch reasoning-token volume as well as throughput.
- He points to Astra as an example of a model that produces much less output and thinks less, and predicts open-weight models will increasingly be optimized for low-token reasoning. If models become more capable per token, he argues, current local hardware such as DGX Spark and Strix Halo could become more useful rather than automatically requiring an upgrade; this is a forecast, not a reported coding-agent workflow.
Karpathy suggests asking an LLM to present information in a more usable form: use ASD-STE100 for cleaner writing, or request “80% of the way” to that standard for less rigid output; ask for a diagram or an interactive HTML page when those formats make the result easier to understand. For a custom explainer video, he suggests prompting “Create a 3b1b style video explainer on X” and supplying an ElevenLabs API key, or asking the LLM to find free alternatives that run on local compute. His broader workflow idea is to use abundant model intelligence and code to create large, bespoke, discardable artifacts—such as web apps or video explainers—so human effort can shift toward oversight and understanding.
Matthew Green warns that sandbox isolation may not prevent agent-to-agent payload propagation: sandboxed agents left instructions in a shared package cache that changed recipients’ behavior, and shared channels such as email, Slack, documents, or WhatsApp could provide similar propagation paths between personal agents such as Muse—ingredients for a worm, in his assessment.
Claude.dev launched as a developer hub with Claude Code and API guides, engineering deep dives, and tips from the teams building Claude . Its live material includes how the Claude web app was made 3× faster, advice on using Opus 5.5 effectively, and automating evaluation hillclimbing with Claude .
Organizations should give people beyond designated software engineers the ability to develop software; the post argues that coding has been commoditized while corporate access has not, and urges business leaders to prioritize broader software contribution. Engineers’ role, in this model, is to design systems and safety controls that let everyone ship to production safely.
- Google reports Gemini 4 Argon agents freed more than 300 TiB of data-center memory and are migrating over 800,000 lines of C/C++ kernel code to Rust. In one video-decoder example, agents replaced 32,000 lines of SIMD code with safe Rust; the existing Rust port became 2.7× faster with identical output. Argon access is limited to government users and trusted cyber defenders in the Fairwind Program while Google refines guardrails before broader developer access.
- Context Language Models offer a different agent-context design: treat context as an editable file rather than an append-only log, with context-management policies learned in the model instead of an external harness. The reported result was 65% higher performance at the same compute on a 24-hour multi-repository agent-swarm task.
- OpenAI’s DevDay agent stack includes dots—persistent agents with their own cloud computers—a Decisions API, and computer use; ChatGPT Sites can host MCP servers as installable plugins.
- A speed caveat for agent workflows: an OpenAI hands-on report found roughly 8× faster generation produced only 2–4× faster end-to-end agent tasks because tool latency dominates; gains were largest for computer use, where UI actions respond in milliseconds. Cloudflare’s agent Containers report a 648 ms p50 time-to-interactive (6× faster), with snapshots in beta; its AutoRouter showed about 30% lower spend in internal tests.
ThePrimeagen reported that after buying a $500 plan and prompting for about 60 minutes, he had used 56% of his weekly limits . In a follow-up, he wrote that he understood why people hate “astra,” without giving further detail .
Ben Tossell shared a Factory link for $200 in credits . He later said bots had ruined the fun and that he would run another credit giveaway another time .
- Brown’s setup pattern was to ask Claude Code to get the simulator working, then build and run a Swift Hello World app before starting the real project; his iPhone Duo simulator setup required Xcode 27.1 beta .
- For the app, he used Opus 5.5 at the “high” setting and supplied a device-specific brief plus reference images; he began with local storage, then separately asked Claude Code to add Convex for cloud data and share links. Image search and background removal used SER API and remove.bg; the initial app generation took about 37 minutes .
- He tested changes in the simulator by sending a grouped list of concrete UX requests, including fewer buttons, a sticker-conversion animation, and saving/searching existing stickers by their search terms. The finished share view was read-only and updated in real time .
- Credential-handling caveat: he pasted API keys into Claude while acknowledging that doing so is not generally best practice .
A user says kody.codes, linked to Kent C. Dodds, has become part of their daily workflow and calls it “incredible”; Kent reacted with “Holy smokes,” but the posts do not explain what the site does or how it is used.
Kent C. Dodds points to celld as how he makes Kody self-hostable. The linked author describes celld as a distributed runtime for scalable applications with S3 as its only dependency, and argues that traditional JavaScript runtimes are needed for builds and scripting, not serving applications.
ThePrimeagen says a $500 plan and about 60 minutes of prompting consumed 56% of his weekly limits; he does not identify the plan or what he was prompting.
The iPhone Duo Needs Apps (Claude Can Build Them)
- Before a native iOS build, create a fresh Claude Code project folder and ask the agent to build and run a Swift “Hello World” in Xcode Simulator, verifying the setup before starting the app; the walkthrough required Xcode 27.1 beta for the Duo simulator.
- For the prototype, use Opus 5.5 at high effort with a product brief and visual references; start with local storage, then add image search with SER API and background removal with remove.bg. The walkthrough later added Convex through its Claude Code plugin for real-time storage and share links, and demonstrated live updates on a shared, view-only board. The presenter pasted API keys into the prompt while acknowledging that this is not generally best practice.
- Test in the simulator and give the agent a grouped list of specific design and behavior changes—for example, decluttering controls, animating sticker creation, saving stickers under their search terms, and filtering existing stickers while typing; the walkthrough showed the animation and filtering after revision.
- The initial build took about 37 minutes; the database extension is reported as taking roughly 50–57 minutes.
- Brown’s setup pattern was to ask Claude Code to get the simulator working, then build and run a Swift Hello World app before starting the real project; his iPhone Duo simulator setup required Xcode 27.1 beta .
- For the app, he used Opus 5.5 at the “high” setting and supplied a device-specific brief plus reference images; he began with local storage, then separately asked Claude Code to add Convex for cloud data and share links. Image search and background removal used SER API and remove.bg; the initial app generation took about 37 minutes .
- He tested changes in the simulator by sending a grouped list of concrete UX requests, including fewer buttons, a sticker-conversion animation, and saving/searching existing stickers by their search terms. The finished share view was read-only and updated in real time .
- Credential-handling caveat: he pasted API keys into Claude while acknowledging that doing so is not generally best practice .