We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
Astra and Sol are faster by default, including in third-party harnesses
OpenAI's Tibo says default speed is now about 50% higher for GPT-6 Astra and GPT-6.1 Sol on subscriptions. This covers OpenAI's own products and partners that use Sign in with ChatGPT, "including OpenCode, Pi, Amp, Devin." You don't need to change anything, and he said it would arrive within two hours . He gives throughput as 50 TPS instead of 30 TPS . This was "Day 1" of the 28-day ship-or-reset pledge from the last brief. It goes at the main complaint about Astra raised there, that it is slow without fast mode. If you dropped Astra for speed, it may be worth re-timing it in your harness.
Cursor SDK: steer running agents and get subagent results back
Cursor shipped a set of SDK changes aimed at people who build their own harnesses:
-
Steering:
run.steer()adds your message to the agent's next turn. If a subagent is mid-task, it moves to the background and keeps working . -
Background subagents report back. Their results return to the parent as a follow-up turn on the same run, so
stream()andwait()keep going until every subagent finishes . -
MCP annotations on custom tools. Tools can carry hints such as
readOnlyHintanddestructiveHint, so the model can tell a lookup from a delete before it calls one . - Your own system prompt. You can replace Cursor's system prompt, and rules, skills and tool schemas still load. It is being enabled account by account .
Steering and custom system prompts work with local TypeScript agents. Background subagent results come back locally in both TypeScript and Python. Details are in the changelog .
Give the agent a way to check its work, and keep that check fast
Addy Osmani's list of tests worth investing in :
- End-to-end tests that simulate real user flows and serve as ground truth.
- Property-based tests that state what must never happen and generate thousands of cases to try to break it.
- When replacing a system, old-vs-new comparisons on randomly selected inputs.
- Fast, deterministic runs. A slow or flaky loop teaches the agent to retry rather than fix the problem.
ThePrimeagen's Omarchy QA harness shows how hard the speed part is. An agent crawling the desktop in QEMU (unlock the screen, type the password, confirm the desktop is back) first took about 15 minutes. Changes got the median to about 5, and a custom harness got it to about 2 . He is now trying Clef, which he describes as Cloudflare's jev and chose for its image capabilities. He is aiming for 30 seconds over login, system menu, lock screen, password, back to desktop and a done check; that is a goal, not a result . He says the harder open question is "What is correct? What do we even test?" .
Huntley's Jiti: grow a running Lisp app by talking to it
Geoffrey Huntley released Jiti (write-up, repo). It is a small kernel where you ask for a capability, the model writes Lisp, and the running application keeps that capability until you ask to remove it . An OpenAI model receives the request, operating instructions, tool definitions and the app's current state. It can inspect functions, read definitions, propose source or run expressions . Accepted definitions are ordinary Lisp, so calling them needs no further inference .
Two design choices carry over to other harnesses:
-
Separate tools for changing and using the app.
develop_formchanges the application andexecute_formuses what already exists. Both share one evaluator, and the split helps the model tell "change the app" from "use it" . - Safety checks decide what is kept; goals decide when the work is done. Accepted changes become durable revisions .
You can also ask it to save a composition, such as shout-backwards built from uppercase-string and reverse-string, as a new function that later requests can find and build on . Huntley calls this "the endgame pattern," but says he "left out how to do verification with it" .
Keep your setup vanilla
@thdxr argues that models improve faster than tinkerers can keep up. In his view, custom workflows mostly solve problems that no longer exist, and someone "naively using vanilla codex" is more likely to be getting the state of the art . Kent C. Dodds agrees: aim for as vanilla a setup as possible . One way to get there, he says, is to give agents tools to use instead of instructions they must follow to the letter .
Memory and observability
- Devin "Dreaming." Across sessions, Devin builds a memory graph of how you like to work. Overnight it removes stale records and looks for latent information . Cognition is open-sourcing the memory format: graph relationships, history kept as records change over time, stored in git and Markdown, and usable with any agent .
-
LangSmith. At Interrupt NYC, LangChain said its LLM Gateway integrates with Codex, Claude Code and Cursor. It offers per-team, per-user and per-key spend limits, rate limiting and provider fallbacks. Stateful fallbacks, which skip a failing provider for a set period, are coming soon . A new
smithtuneCLI filters LangSmith traces, converted into a standard "trajectory" format, and sends them to Fireworks or Baseten for fine-tuning .
Choosing a language when agents write the code
Guillermo Rauch says Vercel's Turborepo port from Go to Rust was internally controversial because of human migration costs. He now argues that "What's 'best for humans' is no longer necessarily 'best for business'" . He doubts Rust is the endpoint either . DHH's case: Rust used to be painful for web apps, but if agents make it productive and you aren't reading the verbosity, "the equation has changed" . He finished his Campfire conversions with a JavaScript/Express port . He invites people to point an agent at once-campfire and request other languages, and he will take PRs as-is .
Smaller items
-
Reviving an old project: Simon Willison gave Claude Opus 5.5 one prompt on his dormant pure-Python WASM engine: "Evaluate current state of pwasm - then consider what it would take to get the MicroPython and micro JavaScript experiments … working under it - and what it would take to speed it up." After 42 commits with minimal follow-up, it handles almost all of the WASM spec. He says he "wouldn't trust this thing at all," hence the alpha tag . His llm-anthropic 0.30 adds
llm anthropic refresh, which pulls the model list from Anthropic's API, andllm anthropic count, which counts tokens before you send a prompt . - Cowork moves to the cloud: Anthropic now runs both inference and the VM in the cloud, with a sandbox per session. The desktop app handles access to local files, so work continues when your laptop is closed and you can use Cowork from a phone .
- T3 Code added in-thread visualizations that agents build, using your theme's CSS variables .
- Beam: Reflection AI announced an open agentic model with 501B total and 23B active parameters, aimed at coding and agentic tasks. Full weights are due this month .
- Small fixes: Theo had Opus 5.5 write a working Chrome extension in 1.5 minutes to stop GitHub screenshots opening in a new tab .
- Sanfilippo says he built an SDF-based CAD and its 3D engine from scratch in a few hours across two days; he credits SDF with enabling substantial functionality using little code.
- He started with Fable 5.1, then did most of the work with Astra after the first tool could no longer continue; he found Astra capable of continuing in the CAD domain. The prototype still had bugs in adaptive triangle simplification, including at higher precision, and mesh generation was slow.
-
To revive a dormant codebase, Simon asked Claude Opus 5.5 to assess his pure-Python WebAssembly engine, get MicroPython and JavaScript experiments working, and improve speed. After 42 commits and minimal follow-up prompting,
pwasmhandled almost all of the WASM specification and its PyPI wheel included working WASM builds of MicroPython, QuickJS, and Micro QuickJS; he cautions that it remains alpha and untrustworthy. -
llm-anthropic 0.30addedllm anthropic refreshto fetch Anthropic’s model list directly from its API, avoiding a package release just to add a new model, andllm anthropic countto use the free token-counting API before sending a prompt. -
Simon used GPT-6 Astra to build an experimental face-blurring tool with MediaPipe’s C++ library compiled to WebAssembly via
@mediapipe/tasks-visionand the BlazeFace model. - For a local model test, Simon used a Codex Remote session with GPT-6 Astra to rerun an addition-in-words experiment on a DGX Spark with Qwen3.8-27B-Q4_K_M: the no-reasoning run scored 23.57% across 5,070 attempts, while a medium-reasoning pilot scored 167/169 with only one sample per digit-length pair. He notes that the one-shot result could change on a rerun, so it should not be treated as a robust comparison.
- Deep Agents is an off-the-shelf harness shaped by coding agents: it provides a filesystem and sandbox for execution plus context management such as compaction, and is intended for building domain-specific agents.
- Harrison Chase presented decision models as complements to frontier models used in Claude Code, Codex, and Deep Agents: they can score outcomes for in-loop evaluation, flag inputs as low-latency guardrails, and route among models, agents, skills, or tools.
-
LangSmith Trajectories standardize agent interaction loops, including data from coding agents and different SDKs; the UI condenses tool calls to make traces easier to scan and annotate, and the format can feed fine-tuning. The announced workflow uses the
smithtuneCLI to select, filter, and preprocess traces, then sends them to Fireworks or Baseten for training and evaluates the resulting model. - LangSmith LLM Gateway integrates with Codex, Claude Code, and Cursor, with cost visibility and spend limits, rate limiting, and provider fallbacks; stateful fallbacks that temporarily bypass a failing provider were described as coming soon.
Simon Willison tested local Qwen 3.8 27B in reasoning and non-reasoning modes on long-number addition. In the non-reasoning test, it scored 1,195/5,070 (23.57%) across 30 fixed pairs per ordered digit-length cell. A medium-reasoning pilot had two misses among 169 pairs, but used one pair per cell and selected the easiest cases first, so it is not a like-for-like comparison with the larger test.
- Riley used Grockbot as a single intake for a coding task: he asked it to improve Native Note’s My Mind list view, and the primary bot immediately delegated to a project-specific developer bot, which worked in Cursor with Opus 5.5. The bot’s coding tab let him open Cursor and inspect the work; in this demo, he also instructed the agent to push directly to production, then confirmed the updated view was live.
- For integration portability, Riley used Composio as a shared home for plugins and skills: he had 28 apps there and connected GPT, Grok, Claude, and Cursor, making it easier to try different agent platforms without recreating the integrations.
Anthropic’s newer Cowork runs model inference and the VM in the cloud, with an isolated sandbox per session; the desktop app handles access to files on the user’s device. The prior setup ran the VM locally, which users found costly in disk, battery, and performance and which stopped working when the laptop was closed. Anthropic says the cloud setup is intended to support phone access and work that continues when the laptop is closed, without the local VM’s battery cost.
Geoffrey Huntley’s Jiti is a small kernel for extending a running Lisp application through conversation with an LLM. The model receives the request, operating instructions, tool definitions, and observations of the application’s current state, then can inspect functions, read definitions, propose source, or execute expressions; a persistent worker returns actual results to inform the next step. Accepted functions and managed data remain available for later requests, and defined functions run as ordinary Lisp without requiring another inference call.
The interface separates develop_form (add, redefine, or remove functionality) from execute_form (call existing functionality, including combinations); both use the same evaluator and transaction machinery. A reusable workflow is to compose existing functions, then ask the kernel to save the composition as a function—for example, shout-backwards combines uppercase-string and reverse-string and joins the catalog for later inspection and reuse. Caller-supplied safety checks govern whether managed changes are accepted, while goals determine completion; accepted changes become durable revisions.
- LangChain’s open-source Deep Agents harness borrows coding-agent patterns for domain-specific agents: it provides a filesystem and sandbox execution environment, plus context management using compaction.
-
LangSmith Trajectories standardizes message-loop histories from coding agents and other agent SDKs; its UI condenses tool calls to make complex runs easier to inspect and annotate, and the format is designed to support fine-tuning. The
smithtuneCLI filters and preprocesses LangSmith traces before sending them to Fireworks or Baseten for training, followed by model evaluation. - LangSmith LLM Gateway integrates with Codex, Claude Code, and Cursor, supports normalized OpenAI and Anthropic formats as well as API passthrough, and offers spend controls, rate limiting, and provider fallbacks. Stateful failover was described as coming soon.
- LangSmith Engine v2 automates an agent-improvement loop: it can reproduce trace issues on a deployment preview, test a fix on a preview branch, and red-team the deployment with simulated failure hypotheses—also helping create an initial eval dataset before launch.
- Cognition introduced Devin “Dreaming”: Devin builds a memory graph of how a user likes to work across sessions, then at night removes stale records and discovers latent information; Cognition says it is creating an open-source Agent Memory Repo standard.
- The open-source memory format supports graph relationships and updates over time with historical records, is backed by Git and Markdown, and can be used with any agent—not only Devin.
Omarchy’s QA agent takes about 15 minutes to crawl the desktop in QEMU; after changes it reached a roughly 5-minute median, and a custom harness reduced the run to about 2 minutes. The remaining challenge is both agent speed and deciding what counts as correct and what to test. ThePrimeagen is experimenting with Clef’s image capabilities to target a 30-second run covering login, system menu, lock screen, password entry, return to desktop, and completion confirmation; this is a goal, not a reported result.
Give coding agents fast, deterministic ways to check their work: use end-to-end tests that simulate real user flows as ground truth, and property-based tests that generate thousands of cases to probe what must never happen. When replacing a system, compare old and new versions on randomly selected inputs. Slow or unreliable test loops can teach an agent to retry rather than fix the underlying problem.
- For a self-hosted agent model, PewDiePie says he taught tool use with supervised fine-tuning on successful interaction examples: he wanted 20,000 clean examples, collected about 300 useful ones himself, then generated and filtered synthetic data to roughly 2,000 examples. He says collecting high-quality fine-tuning data was difficult.
- He then used GRPO: have the model attempt the same task multiple times, score the attempts, and train it to favor results above the group average; he says this needs no separate critic or teacher model. The overall fine-tuning effort took months.
- The video’s sponsor, Namespace, describes its devbox as giving agents a full VM with the real codebase, test suite, databases, and network access, with controls over what enters and leaves the environment.
Badlogicgames called for an official, well-maintained, capable Google Drive/Gmail MCP; the request accompanied Google’s announcement that Markdown files can be previewed in Drive and opened, edited, commented on, and collaboratively worked on in Docs without conversion .
Kent C. Dodds recommends keeping coding-agent setups as vanilla as possible, echoing @thdxr’s view that models improve faster than custom workflows and that vanilla Codex may be closer to the state of the art than setups addressing problems newer models have already overcome. In a reply, Dodds suggests giving agents usable tools rather than relying on instructions they must follow, and points to kody.codes as an example.
Joyce Erhl’s takeaway on AI, highlighted by Kent C. Dodds: less emphasis on typing code and more on the “why” and the “who” behind the work.
- In Riley’s Native Note workflow, a primary Grockbot routed a UI-change request to a project-specific developer bot, which used Cursor; Riley could open its coding tab to inspect progress. He specified Opus 5.5 and explicitly authorized shipping to production; the bot completed the change and the updated view was live.
- Riley uses Composio to centralize plugins and skills, then connect them across GPT, Grok, Claude, and Cursor, reducing the friction of switching agent platforms; he had 28 apps set up there.
Reflection AI introduced Beam, a 501B-parameter agentic open model with 23B active parameters, trained from scratch and described as advancing the Western open frontier on coding and agentic tasks; the company said full weights would be released that month. Geoffrey Huntley reacted, “it’s out. congrats jake.”
Geoffrey Huntley describes an in-app coding-agent workflow: an LLM loop connected to a Lisp REPL adds functions as needed, and those functions persist across application restarts. This replaces the read-file/edit-file/bash-compile cycle with prompting inside the running application. Further reading: Lisp write-up and code to try .
Geoffrey Huntley proposes growing applications in Lisp through iterative conversation with LLMs in a live, interactive environment without compilation steps . He calls this an “endgame pattern” and predicts a YC batch may adopt it as a fundraising pitch or the whole company within nine months, while noting he left out how to verify the approach .
Geoffrey Huntley proposes prompting an LLM to program Lisp for outcomes rather than writing code and enduring compilation-loop pain . He links a Lisp explainer and the ghuntley/jiti GitHub project for trying the idea .
Interrupt NYC: Opening Keynote
- Deep Agents is an off-the-shelf harness shaped by coding agents: it provides a filesystem and sandbox for execution plus context management such as compaction, and is intended for building domain-specific agents.
- Harrison Chase presented decision models as complements to frontier models used in Claude Code, Codex, and Deep Agents: they can score outcomes for in-loop evaluation, flag inputs as low-latency guardrails, and route among models, agents, skills, or tools.
-
LangSmith Trajectories standardize agent interaction loops, including data from coding agents and different SDKs; the UI condenses tool calls to make traces easier to scan and annotate, and the format can feed fine-tuning. The announced workflow uses the
smithtuneCLI to select, filter, and preprocess traces, then sends them to Fireworks or Baseten for training and evaluates the resulting model. - LangSmith LLM Gateway integrates with Codex, Claude Code, and Cursor, with cost visibility and spend limits, rate limiting, and provider fallbacks; stateful fallbacks that temporarily bypass a failing provider were described as coming soon.
- LangChain’s open-source Deep Agents harness borrows coding-agent patterns for domain-specific agents: it provides a filesystem and sandbox execution environment, plus context management using compaction.
-
LangSmith Trajectories standardizes message-loop histories from coding agents and other agent SDKs; its UI condenses tool calls to make complex runs easier to inspect and annotate, and the format is designed to support fine-tuning. The
smithtuneCLI filters and preprocesses LangSmith traces before sending them to Fireworks or Baseten for training, followed by model evaluation. - LangSmith LLM Gateway integrates with Codex, Claude Code, and Cursor, supports normalized OpenAI and Anthropic formats as well as API passthrough, and offers spend controls, rate limiting, and provider fallbacks. Stateful failover was described as coming soon.
- LangSmith Engine v2 automates an agent-improvement loop: it can reproduce trace issues on a deployment preview, test a fix on a preview branch, and red-team the deployment with simulated failure hypotheses—also helping create an initial eval dataset before launch.