We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
GPT-6.1 Sol: early users give it review, investigation and computer use, and keep Opus 5.5 for writing code
OpenAI launched GPT-6.1 Sol at DevDay. Tibo (OpenAI) describes it as "near Astra intelligence at one fifth of the price of Astra" with a 95% cache-read discount . Theo gives the prices as $2/M input, $10/M output and $0.10/M cache reads. Cache reads used to be 10% of the input price, so this halves them. That matters because agent work is mostly cache reads . When Opus 5.5 recalculated the cost of Theo's past sessions at these prices, its estimate fell from $5,700 to $1,550 .
Theo's position after testing: "I still think Opus 5.5 is the current GOAT for coding." He doesn't default to Sol for writing code, but does for code reviews, architecture analysis, computer use and email . His workflow:
- Opus writes the code, Sol audits it. On a T3 Code PR, Sol found two bugs that Fable and Opus had missed: follow-up messages could stay blocked after the agent started, and recovered setup progress disappeared too early. He pasted Sol's findings into Claude to fix . He calls Opus "a more pleasant collaborator" and says Sol is better at digging into details .
- Let Opus call Sol. He plans to let Opus 5.5 call Sol for root-causing bugs, triage and reviewing Opus's own work, and already runs Sol inside Claude Code for this .
- Don't use Sol for long unattended rewrites or UI. It ran for days on his TypeScript-to-Rust compiler port without progress . He says it "sucks at front end" and fills UIs with unnecessary text .
- Use it to clean up after other agents. In an audit of that port, Sol found that about 1.3M of 1.8M lines were unused. Opus had rewritten the port from scratch in a new crate and never deleted the old code .
The harness changes the benchmark result. Theo first reported Sol beating Opus 5.5 on Terminal-Bench 4 for about 1/30th of the price . He then noticed that he had run it in Codex, while Artificial Analysis used mini-swe-agent . Harbor Hub runs each model in its official harness; there every model scored higher than in Artificial Analysis's runs, and Sol gained the most from moving to Codex . When you compare models, run them in the harness you actually use.
Codex Cloud 2.0, the Agents API and dots
- Codex Cloud 2.0 has configurable cloud environments. Tibo says that once you've configured one, "it's impossible to go back to building on your laptop" . It runs on the same infrastructure as dots . At the keynote, Romain Huet sent "rewrite the entire back end in rust" to the cloud to check on later . He also used an "app shot" to have Codex audit an app across screen sizes in the simulator, running in the background . Simon Willison tried Codex Cloud for a phone-built live-blog tool, ran into problems, and switched to Claude Code for web .
- Agents API (preview): the same technology that powers OpenAI's cloud agents, with computer use .
- Dots are always-on agents running on Astra, with their own computer and browser. They are included in Pro without drawing on usage . A Codex task that a dot creates is billed as normal usage . Alexander Embiricos's rule: use Codex when you want to do the work yourself, and a dot when you want to delegate . He says Codex is like a principal engineer that won't check Linear or Slack unless you explicitly tell it to .
- Sign in with ChatGPT lets you spend your included subscription usage inside partner products such as Devin and OpenCode . Tibo gave "over 60" partners in one post and "over 16" in a later one .
- Open models in Codex: enterprise teams can run GLM-5.3 Flash and Kimi K3 natively, and the spend counts against their OpenAI commit .
- Ultrafast runs 8× faster at 6× the cost . Matthew Berman says he burned $1,000 in about 90 minutes with early access. He found that with inference this fast, local tool calls and terminal commands became the slow part .
- Plans: Plus is 1×, Pro 100 is 5× and Pro 200 is 10×. Existing Pro 200 subscribers keep 20× for a while and get extra credits . Tibo says the reopened Pro $200 plan works out to half the API-dollar value of the old one, with no five-hour limit .
Pruning tests written by agents
Peter Steinberger says OpenClaw deleted about 400k lines of its own tests "without much change in code coverage." Models write tests for every tiny change, and his test-audit skill helped remove them . Kent C. Dodds dropped about 70k lines of tests from Kody (PR #2707) .
Steinberger also moved his team's agent sessions into one shared place. People were embarrassed for a day, and then learned from how differently everyone worked with their agents .
Sonnet 5.5: better as a subagent than as your main model
Theo's advice on Sonnet 5.5 is to let Opus or Fable call it rather than prompting it yourself . On his benchmark for planning how to land a huge PR, Sonnet scored slightly above Opus at about half the price, in about five minutes. That is codebase comprehension, not implementation . On a game build it took 43 minutes against Opus's 36 . He recommends avoiding both max and low effort. In one benchmark, going from xhigh to max raised token use by 1,500% .
A reusable pattern for agents that change state
The Box/LangChain contract-review agent shares read tools with subagents. Tools that change state or wait for a human approval stay with the lead agent . The rule: "The LLM extracts, the playbook decides." Business rules live in code you can inspect and test .
Smaller items
- Cursor
/visualizedraws charts and diagrams inline in the Agents Window . - Pi: MCP and codemode are shipped extensions, and you can turn them off with
pi config. - LangChain Managed Deep Agents has a terminal UI for test-driving agents locally:
mda chat. - Kody can run Codex agents . Dodds says the Decisions API isn't documented anywhere he could find yet, and suggests jev in the meantime . His broader point is to keep integrations, memories and tokens outside any one agent so switching agents stays cheap . He builds Kody, so he isn't neutral on this.
- For ZX Spectrum rotation, Sanfilippo approximated multiplication by sine or cosine as sums of shifted values, with 64 angle steps per 90° quadrant and four shifts per approximation. The model suggested chaining incremental shifts through an accumulator instead of repeatedly shifting the original value, then generating Z80 instruction sequences and jumping to the one for the selected angle.
- For his released Spectrum virtual machine for Another World, he says he gave the model 1,000 optimizations of his own; it contributed additional ideas for working within 48 KB of RAM. These included halving vertical resolution by doubling each row while retaining full horizontal resolution, and using two black attribute bands in the graphics area as about 1 KB of additional RAM.
- One of Sanfilippo’s own speedups was a rectangle fast path: set color attributes for interior blocks rather than writing each pixel.
- Steinberger uses annoyance with slow or awkward computer workflows as a cue to redesign them: give agents tools to take actions, such as installing software, rather than limiting them to summarization. He says model advances can make workflows obsolete quickly; ultra-fast models, for example, let him stay focused instead of juggling many terminal windows.
- He moved his team to share agent interactions in one place. Though people initially felt exposed, sharing revealed how differently everyone used agents and helped the team learn from one another.
- He is exploring an agent that can create agent hierarchies on demand.
- He describes using agents at the operating-system level: his agent can read and recompile the Linux kernel to fix a webcam. He also had Codex work overnight on an ESP device’s display driver; when he added an audio fix, repeated tests kept waking him because the volume was down elsewhere, until the agent figured it out.
- For AI prompt entry on Windows, he configures Raycast hotkeys to launch apps and open clipboard history (Alt+M) to reuse copied text, images, files, or links. He also dictates prompts with WhisperFlow; he describes Handy as a free, open-source alternative whose formatting he finds worse.
- He connects tools to ChatGPT desktop’s Codex through Plugins; the Notion integration can search content, update pages, and upload photos. He also prompts it to create and reformat Notion pages, including converting a document to a tabs view, and uses computer-use for GUI tasks such as creating a chart in Paint.
- The ChatGPT Chrome side panel shares chats with the desktop app and uses context from the open webpage, letting him ask about an article or request a website recreation without copying and pasting.
- Pairing the desktop app’s Remote feature with a phone by QR code lets him send instructions to ChatGPT desktop remotely; together with Codex computer-use, he can make changes and bring the Notion desktop app to the foreground.
- Theo’s workflow pairs Opus 5.5 for implementation with 6.1 Soul for investigation and review . On a T3 Code PR, Soul found two issues Fable and Opus had missed—queued follow-up messages could remain blocked after the agent started, and recovered setup progress could disappear too early—and Theo passed the findings to Claude to fix .
- In Theo’s evaluation, 6.1 Soul was less spiky than Astra and useful for scoped work, but Opus remained his preference for long, unattended implementations . He also judged Soul weak at front-end design and UI, despite its other strengths .
- Theo reported 6.1 Soul pricing of $2 per million input tokens, $10 per million output tokens, and $0.10 per million cache-read tokens; he emphasized the cache-read rate as especially important for agent workloads .
- In a sponsored CI segment, Theo described letting agents use Depot’s CLI to run the same CI checks without pushing a PR, with failure analysis and suggestions helping feed results back to agents .
-
Embiricos’s rule of thumb: use Codex for work you want to do manually and a dot when you want to delegate it. He describes Codex as principal-engineer-level but says it needs explicit direction rather than being expected to check Linear or Slack on its own; a dot has its own cloud computer, can connect to your computer, and appears in Slack under a distinct owner-linked identity such as
AE-dot. - In his own PM workflow, he mostly talks with his dot in Slack and uses voice while commuting to prepare for conversations, including having it grill him with hard questions.
- For collaborative work, he points to asking a dot for a clickable HTML artifact—such as an interactive page or game—rather than relying only on text; he says HTML authoring is now effectively free.
- OpenAI announced Codex Security Cloud for continuous code and vulnerability scanning, a cloud-native Codex that does not require leaving a laptop on, plus a refreshed CLI with voice control and a new code-review experience. The speaker’s practical caveat: with very fast inference, local tool calls and terminal commands became the bottleneck.
- The video describes GPT-6.1 Soul as close to GPT-6 Astra on developer and PDF benchmarks at lower cost and faster inference: $2/$10 per million input/output tokens for Soul versus $10/$50 for Astra, with cached Soul input at $0.10 per million. Astra still performs better at the top end of the OS World computer-use benchmark.
- The speaker reports that the accelerated Astra option was eight times faster but cost six times more, and that early access burned through $1,000 in about 90 minutes—a concrete warning to weigh speed against usage costs.
- For a complex project with at least four independent workstreams, Krieger used Claude to monitor progress and create the project UI/status view—first for the TPM, then for the whole team. The team could also change the interface with Claude; open questions included data provenance and whether people or Claude in the background could update the data.
- He described turning a write-up on agent-native architecture into a skill, centered on the principle that agents should be able to do anything a human can do in a product.
- For agent-compatible products, he recommended shared underlying primitives so agents and REST APIs use the same plumbing instead of bolted-on integrations. That foundation can support a gradual path from a sidebar to more malleable or personalized interfaces without inventing separate infrastructure.
- Treat Sonnet 5.5 as a delegated research and planning subagent, rather than the default model you prompt directly: have Opus or Fable use it for deep codebase investigations, checking hypotheses, and scoping complex work. In the creator’s revised T3 Code benchmark—breaking a large PR into a strategy for landing changes—Sonnet scored slightly above Opus at about half the price and finished in roughly five minutes; the test focused on analysis and codebase comprehension, not implementation.
- Don’t assume Sonnet is cheaper or faster for every coding task: in a separate fish-game build, it took 43 minutes versus Opus’s 36 and used more than 3.5 times as many input tokens for a similar cost. The cost comparison is qualified because the Opus run also used a Fable subagent.
- Avoid Max reasoning effort (and Low): the creator says Max can force unnecessary reasoning, reporting a 1,500% increase in token use versus X-high on benchmarks and cases where overthinking hurt answers; he recommends avoiding both settings.
ThePrimeagen claims OpenAI hacked Hugging Face, which had spare GPUs and GLM 5.x available, and that this ultimately helped “root the cause and save their bacon.” He argues that frontier labs seem more dangerous than open-model providers and calls for banning private models.
- Codex’s terminal workflow supports quick natural-language iteration: Romain prompted UltraFast to build a React app that randomly selects attendees, then asked it to change the selection from three people to six .
- For UI testing, an app shot supplies context for Codex to navigate an app on the computer. Romain prompted it to audit the app across screen sizes; it opened the simulator, explored the app, and could keep running in the background while he worked elsewhere .
- Codex’s cloud workflow lets a developer hand off a larger task—Romain sent a prompt to rewrite an entire backend in Rust—and check on it later .
- In ChatGPT desktop, connect existing tools through Plugins first (Notion is the example), enabling the agent to search and update pages; then use Computer Use for GUI tasks such as making a chart in Paint. Brown presents Notion as especially agent-friendly because of its MCP/tool support, and demonstrates prompting ChatGPT to create a page and convert it to tabs.
- The Chrome ChatGPT side panel syncs chats with the desktop app and uses the current page as context, so you can ask about an article or request a recreation of a viewed website without copying and pasting.
- For remote desktop work, enable ChatGPT desktop’s “Control this PC,” pair a phone by QR code, and send instructions from mobile; the demo edits Notion and uses Codex computer control to open the native Notion app, with the PC left on.
- Grokbot’s connected agent team can check business tools such as Slack for project updates and return a link to the exact thread.
Anthropic reports a capability gap on 100 randomly selected internal binary-exploitation tasks: GLM-5.3 produced full control-flow hijacks in 4% of trials and Claude Mythos Preview in 6%, while Claude Opus 4.6 and GLM-5.2 succeeded in none .
-
A builder’s outreach-engine workflow configures Treg in Claude Code, Codex, or another agent with
set up treg - https://treg.to/llms.txt. The agent searches X, LinkedIn, and Reddit for launch signals; simple rules filter noise before paid enrichment; a JEV judge qualifies leads; contact details are purchased only for qualified leads; and the agent drafts while a person reviews and sends . - The writeup describes Treg as providing 3,700+ tools and data sources from 99 providers through one key, with per-call provider pricing and no markup . The builder reports 105 leads, three booked calls, one closed deal, and $2.87 in total run costs .
ThePrimeagen says he would welcome the LLM revolution if it produces working, lean software, citing fatigue with “ts-bloat.” He points to DHH’s beta Campfire rewrite in Rust: DHH described the conversion as “free*,” though the code was ugly and six times as verbose, and said he never had to look at the Rust directly.
Simon Willison vibe-coded a system on his phone while traveling to OpenAI DevDay to make adding photos to his live blog easier; after running into problems with Codex Cloud, he switched to Claude Code for web.
Kody can run and automate Codex agents. For support email, Dodds suggests using OpenAI’s Decision API to determine whether an agent can handle a request, then launching a Codex agent for it. Since he could not find documentation for the Decisions API, he suggests using jev in the meantime.
Kent C. Dodds argues that coding agents will leapfrog one another, so developers may want to switch agents or use several; integrations, automations, memories, and tokens can make switching difficult when they are locked into one agent. He recommends keeping that shared agent context in Kody rather than each agent’s own environment, so a replacement agent can use the same platform.
OpenClaw deleted around 400,000 lines of tests with little change in code coverage. Steinberger says modern models often generate tests for every tiny change, including unhelpful ones, and that the linked test-audit skill helped.
Kent C. Dodds says kody.codes can be used with Codex to spin up and talk to a Codex agent through another agent, such as the Claude mobile app—a cross-agent workflow for accessing Codex.
Jason Zhou introduced an open-source “Clay Killer” project and /leads-signal skill in the treg workflow-skills repo. The skill uses Claude to monitor keyword mentions, complaints, product reviews, job changes, and 23 other signals, then find hot leads each morning. Zhou claims it costs $0.0002 per signal versus $167/month for Clay plans.
OpenAI should be scared of this one
- Treat Sonnet 5.5 as a delegated research and planning subagent, rather than the default model you prompt directly: have Opus or Fable use it for deep codebase investigations, checking hypotheses, and scoping complex work. In the creator’s revised T3 Code benchmark—breaking a large PR into a strategy for landing changes—Sonnet scored slightly above Opus at about half the price and finished in roughly five minutes; the test focused on analysis and codebase comprehension, not implementation.
- Don’t assume Sonnet is cheaper or faster for every coding task: in a separate fish-game build, it took 43 minutes versus Opus’s 36 and used more than 3.5 times as many input tokens for a similar cost. The cost comparison is qualified because the Opus run also used a Fable subagent.
- Avoid Max reasoning effort (and Low): the creator says Max can force unnecessary reasoning, reporting a 1,500% increase in token use versus X-high on benchmarks and cases where overthinking hurt answers; he recommends avoiding both settings.