We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
GPT-6.1 Sol: early users give it review, investigation and computer use, and keep Opus 5.5 for writing code
OpenAI launched GPT-6.1 Sol at DevDay. Tibo (OpenAI) describes it as "near Astra intelligence at one fifth of the price of Astra" with a 95% cache-read discount . Theo gives the prices as $2/M input, $10/M output and $0.10/M cache reads. Cache reads used to be 10% of the input price, so this halves them. That matters because agent work is mostly cache reads . When Opus 5.5 recalculated the cost of Theo's past sessions at these prices, its estimate fell from $5,700 to $1,550 .
Theo's position after testing: "I still think Opus 5.5 is the current GOAT for coding." He doesn't default to Sol for writing code, but does for code reviews, architecture analysis, computer use and email . His workflow:
- Opus writes the code, Sol audits it. On a T3 Code PR, Sol found two bugs that Fable and Opus had missed: follow-up messages could stay blocked after the agent started, and recovered setup progress disappeared too early. He pasted Sol's findings into Claude to fix . He calls Opus "a more pleasant collaborator" and says Sol is better at digging into details .
- Let Opus call Sol. He plans to let Opus 5.5 call Sol for root-causing bugs, triage and reviewing Opus's own work, and already runs Sol inside Claude Code for this .
- Don't use Sol for long unattended rewrites or UI. It ran for days on his TypeScript-to-Rust compiler port without progress . He says it "sucks at front end" and fills UIs with unnecessary text .
- Use it to clean up after other agents. In an audit of that port, Sol found that about 1.3M of 1.8M lines were unused. Opus had rewritten the port from scratch in a new crate and never deleted the old code .
The harness changes the benchmark result. Theo first reported Sol beating Opus 5.5 on Terminal-Bench 4 for about 1/30th of the price . He then noticed that he had run it in Codex, while Artificial Analysis used mini-swe-agent . Harbor Hub runs each model in its official harness; there every model scored higher than in Artificial Analysis's runs, and Sol gained the most from moving to Codex . When you compare models, run them in the harness you actually use.
Codex Cloud 2.0, the Agents API and dots
- Codex Cloud 2.0 has configurable cloud environments. Tibo says that once you've configured one, "it's impossible to go back to building on your laptop" . It runs on the same infrastructure as dots . At the keynote, Romain Huet sent "rewrite the entire back end in rust" to the cloud to check on later . He also used an "app shot" to have Codex audit an app across screen sizes in the simulator, running in the background . Simon Willison tried Codex Cloud for a phone-built live-blog tool, ran into problems, and switched to Claude Code for web .
- Agents API (preview): the same technology that powers OpenAI's cloud agents, with computer use .
- Dots are always-on agents running on Astra, with their own computer and browser. They are included in Pro without drawing on usage . A Codex task that a dot creates is billed as normal usage . Alexander Embiricos's rule: use Codex when you want to do the work yourself, and a dot when you want to delegate . He says Codex is like a principal engineer that won't check Linear or Slack unless you explicitly tell it to .
- Sign in with ChatGPT lets you spend your included subscription usage inside partner products such as Devin and OpenCode . Tibo gave "over 60" partners in one post and "over 16" in a later one .
- Open models in Codex: enterprise teams can run GLM-5.3 Flash and Kimi K3 natively, and the spend counts against their OpenAI commit .
- Ultrafast runs 8× faster at 6× the cost . Matthew Berman says he burned $1,000 in about 90 minutes with early access. He found that with inference this fast, local tool calls and terminal commands became the slow part .
- Plans: Plus is 1×, Pro 100 is 5× and Pro 200 is 10×. Existing Pro 200 subscribers keep 20× for a while and get extra credits . Tibo says the reopened Pro $200 plan works out to half the API-dollar value of the old one, with no five-hour limit .
Pruning tests written by agents
Peter Steinberger says OpenClaw deleted about 400k lines of its own tests "without much change in code coverage." Models write tests for every tiny change, and his test-audit skill helped remove them . Kent C. Dodds dropped about 70k lines of tests from Kody (PR #2707) .
Steinberger also moved his team's agent sessions into one shared place. People were embarrassed for a day, and then learned from how differently everyone worked with their agents .
Sonnet 5.5: better as a subagent than as your main model
Theo's advice on Sonnet 5.5 is to let Opus or Fable call it rather than prompting it yourself . On his benchmark for planning how to land a huge PR, Sonnet scored slightly above Opus at about half the price, in about five minutes. That is codebase comprehension, not implementation . On a game build it took 43 minutes against Opus's 36 . He recommends avoiding both max and low effort. In one benchmark, going from xhigh to max raised token use by 1,500% .
A reusable pattern for agents that change state
The Box/LangChain contract-review agent shares read tools with subagents. Tools that change state or wait for a human approval stay with the lead agent . The rule: "The LLM extracts, the playbook decides." Business rules live in code you can inspect and test .
Smaller items
- Cursor
/visualizedraws charts and diagrams inline in the Agents Window . - Pi: MCP and codemode are shipped extensions, and you can turn them off with
pi config. - LangChain Managed Deep Agents has a terminal UI for test-driving agents locally:
mda chat. - Kody can run Codex agents . Dodds says the Decisions API isn't documented anywhere he could find yet, and suggests jev in the meantime . His broader point is to keep integrations, memories and tokens outside any one agent so switching agents stays cheap . He builds Kody, so he isn't neutral on this.