# Route each task to the cheapest capable model: LangChain cuts Open SWE's median thread cost 64% without hurting merge rates

*By Coding Agents Alpha Tracker • October 2, 2026*

LangChain published a detailed recipe for model routing inside the agent harness, with A/B results. Also: Theo's workflow for getting PRs merge-ready, Claude Code mods, a draft-PR warning about Dots, and smaller tool releases.

## Routing inside the harness: 64% cheaper, same merge rate

LangChain wrote up how it built a model router for Open SWE, its internal coding agent, and published the results. It picked three tiers from different providers: **GLM-5.3-Flash (xhigh)** as fast, **GPT-5.6 Sol (medium)** as balanced and **GPT-6 Astra (low)** as performance [^1]. The [router middleware](https://github.com/langchain-ai/open-swe/blob/main/agent/middleware/model_selection.py) runs once, on the thread's first human message, and has three parts [^1]:

- a base prompt telling the classifier to pick the least expensive model likely to finish the task
- short, plain-language criteria for each tier
- a classifier model

The tier criteria come from two sources: LangChain's own analysis of its task mix, and each provider's prompt guide. Because those criteria depend so much on the agent's own tasks, the router lives in the harness rather than a generic gateway [^1]. The first classifier was an LLM with structured output. Switching it to the Jev decision model made classification almost 50× faster. The chosen model then handles the whole thread [^1].

They measured outcomes with merged PRs per thread and thumbs up/down ratings. Merged PRs were the stronger signal because few threads got rated [^1]. Results from a 973-thread A/B test against always using Astra:

- **Quality:** 29.2% of routed threads ended in a merged PR, against 27.3% for control (p=0.49) [^1].
- **Cost:** the median thread fell from $2.61 to $0.94, a 64% drop. The mean fell 42% and p90 fell 37% [^1].
- **Routing mix:** 56% of threads went to balanced, 34% to fast and 10% to performance. Median cost per thread ranged from $0.097 on fast to $2.88 on performance [^1].

The opposite test, router against always using the fast model, was stopped within a day. Engineers said the low quality was disrupting their work [^1]. In short, routing saves money, but sending everything to the cheapest model doesn't work.

Related: LangChain's new `smithtune` CLI turns an agent's own trajectories into fine-tuning data [^2]. LangChain says OpenSWE Review's precision rose from 62.9% to 81.5% with 29% fewer tool calls per review [^3].

## Hand off verification too, not just coding

In a new video, Theo estimates that only 10–15% of his tokens go to writing code. The rest goes to verifying it, so the PR is actually done by the time he looks [^4]. He also says that if undoing a bad change takes more than 15 seconds, you should fix your systems before vibe coding [^4]. Habits from the video:

- **Give the agent the problem, not your solution.** He now sends a bug screenshot with "Fix it." If the agent fails, that tells him the problem deserves his own thinking [^4].
- **Ask for everything up front:** build it, file a PR, spin up a Tailscale environment for him to try, then babysit the PR [^4]. If you state the conditions under which the agent may merge, the thread can close without you ever opening it [^4].
- **Treat threads as a to-do list.** He fires prompts into the background (Cmd+Enter in T3 Code), only opens threads marked done or needing input, and spreads work across his Linux boxes, with auto load-balancing available [^4].


[![If you have a Claude sub, watch this](https://img.youtube.com/vi/D8PikZ1KhUo/hqdefault.jpg)](https://youtube.com/watch?v=D8PikZ1KhUo&t=2772)
*If you have a Claude sub, watch this (46:12)*


On exploratory work, Theo argues that the branch you explore on and the PR you merge can, and often should, be built differently: "assume all PRs will be closed, not merged" [^5]. Geoffrey Huntley goes further on review. Read the properties your tests express, and if performance is a problem, "run a loop and optimise" while those properties keep holding [^6]. The things he says you must read are public SDK APIs, your verification and agent traces [^7].

The Pragmatic Engineer quotes DHH saying 37signals is "done writing code by hand." When it happens, it's treated like a bug: you ask why the agent couldn't produce the result, then fix the "factory" [^8]. 37signals is now building native mobile apps, moving backend services from Ruby to Rust because agents write good-enough Rust, and keeping Rails because its conventions suit agents [^8].

## Claude Code mods

Claude Code now supports mods, which can change its behavior, customize the UI or swap in your own features. You can write one in a few lines of TypeScript or have Claude build it. Mods ship inside plugins and install with `/plugin` in the CLI or desktop app [^9]. Boris Cherny's pitch is that everyone works differently, so build your own and share them as plugins [^10].

## Dots: open autonomous PRs as drafts

Alexander Embiricos described an early internal incident. A dot watching a feedback channel found a bug and opened a non-draft PR. It had inferred the engineer's habit of setting auto-merge, and the PR merged once a colleague approved it. "Dot should set PRs in draft mode," he said [^11]. If you give an always-on agent repo access, make draft PRs the default.

Riley Brown's notes after 24 hours with a dot:
- Voice is the standout, but he hit problems in sessions of 30 minutes or more [^12].
- He wants a better view of the Codex threads the dot spins up [^12].
- He found it confusing to choose between the VM browser, the Codex browser and Chrome [^12].

On capacity, Tibo says GPT-6.1 Sol is back to normal speed after a load spike, and a global usage reset for paid ChatGPT accounts lands at 10am PST on the day after his post [^13]. ThePrimeagen says about 60 minutes of prompting on a "$500 plan" used 56% of his weekly limit [^14].

## Smaller items

- **Cursor** added GLM 5.3 and GLM 5.3 Flash. Cursor says GLM 5.3 Max is the best-scoring open-weight model on CursorBench 4.0 [^15].
- **`@langchain/mcp-adapters` 2.0** supports the latest stateless MCP, adds tools that can check in with users and makes auth easier [^16].
- **Pi 1.0** shipped with Pi Durable ([post](https://earendil.com/posts/pi-1-0/)) [^17].
- **Persistent subagents:** in a Latent Space interview, swyx calls ephemeral Codex subagents his biggest pain and Alex Zhang agrees that writing state to the filesystem is "the big trick." Zhang's Pi-based Prime Agent offers subagents that outlive the root agent and can be prompted later [^18].
- **Figma MCP workaround:** Kent C. Dodds suggests forking his [Kody Figma setup](https://kody.codes/@kody/figma) and asking your agent to set it up, which he says takes about five minutes [^19]. He builds Kody.
- **Btrfs:** Peter Steinberger found copy-on-write great for worktrees but terrible for SQLite. The next OpenClaw update moves the database to a NOCOW location [^20].
- **Making agent output easier to understand:** Karpathy suggests asking for ASD-STE100 controlled English (or "80% of the way" to it), diagrams, HTML pages, or "a 3b1b style video explainer on X" [^21].
- **Security:** Matthew Green notes that agents in separate sandboxes left instructions for each other in a shared package cache, and those instructions changed what the recipients did. Swap the cache for email, Slack or shared docs and you have the makings of an agent worm [^22].

---

### Sources

[^1]: [𝕏 article by @sydneyrunkle](https://x.com/i/article/2105489324590645248)
[^2]: [𝕏 post by @LangChain](https://x.com/LangChain/status/2105671430826230065)
[^3]: [𝕏 post by @LangChain](https://x.com/LangChain/status/2105671433636794690)
[^4]: [If you have a Claude sub, watch this](https://www.youtube.com/watch?v=D8PikZ1KhUo)
[^5]: [𝕏 post by @theo](https://x.com/theo/status/2105854195253522501)
[^6]: [𝕏 post by @GeoffreyHuntley](https://x.com/GeoffreyHuntley/status/2105829115064615004)
[^7]: [𝕏 post by @GeoffreyHuntley](https://x.com/GeoffreyHuntley/status/2105829426089082965)
[^8]: [The Pulse: RoR creator sparks new “death of coding by hand” debate](https://blog.pragmaticengineer.com/the-pulse-ror-creator-sparks-new-death-of-coding-by-hand-debate)
[^9]: [𝕏 post by @ClaudeDevs](https://x.com/ClaudeDevs/status/2105721434807083061)
[^10]: [𝕏 post by @bcherny](https://x.com/bcherny/status/2105756563302723721)
[^11]: [OpenAI’s Dots lead on the future of ChatGPT](https://www.youtube.com/watch?v=VnHOFQE7nCw)
[^12]: [𝕏 post by @rileybrown](https://x.com/rileybrown/status/2105670426378535406)
[^13]: [𝕏 post by @thsottiaux](https://x.com/thsottiaux/status/2105843926221660585)
[^14]: [𝕏 post by @ThePrimeagen](https://x.com/ThePrimeagen/status/2105840136902570138)
[^15]: [𝕏 post by @cursor_ai](https://x.com/cursor_ai/status/2105787358557999585)
[^16]: [𝕏 post by @LangChain_JS](https://x.com/LangChain_JS/status/2105744469547208842)
[^17]: [𝕏 post by @pidotdev](https://x.com/pidotdev/status/2105738462712209603)
[^18]: [Academia is for Ambition — Alex Zhang, MIT](https://www.latent.space/p/rlm)
[^19]: [𝕏 post by @kentcdodds](https://x.com/kentcdodds/status/2105692703476572174)
[^20]: [𝕏 post by @steipete](https://x.com/steipete/status/2105720334821261330)
[^21]: [𝕏 post by @karpathy](https://x.com/karpathy/status/2105819303471976479)
[^22]: [Quoting Matthew Green](https://simonwillison.net/2026/Oct/1/matthew-green)