We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
Voice plus a live preview works as a real Codex workflow
Simon Willison built a newsletters feature for his Django blog mostly by talking to it. He used the Codex tab in the ChatGPT desktop app in voice-conversation mode, running against a local checkout . His setup:
-
Type
Start dev server and open in browserfirst. That gives you a preview you can ask the agent to navigate, so you can follow its progress by eye . - Click "Start new voice chat". It is the button to the right of the microphone, not the microphone itself .
- Talk through the requirements, disfluencies included. His transcript spelled out the product rules: no tag pages, no blog index, yes date archives, and searchable once public a month after sending. The model, GPT-6 Astra High, understood it .
In about 30 minutes of cooking, with occasional clarifying questions from the model, he got a new model and migration, four importers (Substack RSS, Substack's undocumented API, and public and private GitHub repos), archive pages and search integration . He then had Codex open a PR and reviewed it in GitHub. One importer shelled out to Git, so he switched to typing and had it rewritten against the API. Getting it ready to deploy took about another half-hour of typed prompting .
His verdict: voice is great for multitasking but won't be his daily driver. Once he gets into details, pasting errors and examples or pointing at specific code beats describing them aloud .
Riley Brown showed the delegation version through ChatGPT dots. He talked to his dot and asked it to "create a new codec session" to add gradient fills to rectangles in his internal drawing tool. The dot restated the requirement as "make sure they save and reload properly" . The task ran on his MacBook Pro on a separate branch, testing rendering, saving and exports , and the dot later reported that the checks passed and the feature was live . Sessions can be local or cloud and can run concurrently. His advice is to learn Codex basics so you give the dot better instructions . OpenAI's update says dots can now start Codex work and follow up on existing threads, drawing on ChatGPT conversations, Codex threads and automations, and decide better when to continue a thread or start fresh .
Claude Code Projects opens to the full waitlist
Claude let in every Pro and Max user from the Projects waitlist . Addy Osmani describes when to use a project: a stream of related work that outlasts one session. Claude coordinates it and runs each task as a parallel cloud thread, and you set instructions, repos and memory once . You can hand off a batch and come back to an Overview pane that shows finished threads, ready PRs and threads waiting on you .
Codex and model updates
- Windows sandbox: Codex has a new sandbox mode built on Microsoft Execution Containers (MXC). OpenAI says it sets up faster, enforces network rules more strongly and gives granular file-access controls. It requires a compatible Windows 11 device . Alexander Embiricos says the Windows sandbox was one of Codex's top feedback areas .
- Composer predictions (beta, Pro): Codex suggests your next message based on the conversation and how you talk to it . Tibo says it is in the desktop app and included in Pro without consuming usage .
- Devin: you can connect a ChatGPT Go, Plus or Pro plan, and GPT usage in Devin draws from that plan's quota .
- Opus 5.5 fast mode has rolled out . Theo warns it is not included in a Claude subscription and bills you when you use it .
- DHH says this is the first time he has preferred Codex as his main agent, with Claude secondary, for an extended period, and credits "Sol 6.1" .
- LangSmith usage signals: over the past month, Claude Sonnet 5 rose from #9 to #2 in adoption, and "gpt 5.6 luna" rose from #3 to #1 in call footprint. Smaller, faster models dominate call volume .
Huntley: tune --help with an agent loop, and learn the harness
Geoffrey Huntley's advice to tool builders is to optimize the CLI's --help. Measure whether a model can reach outcomes by walking help across all subverbs, then pick the wording that gets there in the fewest tool calls. That loop can be automated . Next, publish docs and optimize them so an agent finds what it needs in a single web search .
In a separate post, his concrete advice for engineers is to build your own agent and be able to explain how Claude Code works under the hood . His interview bar: a candidate should know context windows, tokenization and inference from a systems-design angle. The "ideal" answer on harnesses is either that they are fungible and you use whichever gets tokens cheapest, or that you built your own . He also says he will open-source a VM provisioner that creates machines in under five seconds, is aware of fleet capacity and can relocate machines between nodes .
Turning production traces into agent evals
In a LangChain talk, Jake Broekhuizen lays out a loop. First, mine production traces for tool misfires and user frustration. Frontier LLM judges get expensive at scale, so LangChain trained a purpose-built judge to classify and cluster failures . Second, reproduce failures in a mocked "world" that includes database state, permissions, schemas and prior actions, not just input and output . Coding agents that help build these worlds need constant supervision so tasks stay hard and can't be reward-hacked . Third, fix at the harness level (prompts, tool schemas, context injection) or the model level, then rerun . For Managed Deep Agents users, mda eval init sets up a Harbor environment with tasks and checks, and runs are traced into LangSmith Experiments .
Smaller notes
- Addy Osmani warns that picking from an agent's suggestions is a different skill from coming up with ideas yourself, and the second erodes if you only ever pick: "If you don't understand your own codebase… the agent is directing you" . Kent C. Dodds says the gut reaction that rejects an agent's direction before he can articulate why is what he is keeping .
- Simon Willison is looking for an open-weight MoE coding model that fits in under 60GB of RAM and runs faster than 12 tokens/second. Qwen3.5-35B-A3B is his current pick .
- Anthropic will publish more frequent model-behavior reports. The first covers four types of cases where Claude acted on real websites or systems in ways Anthropic didn't intend, sometimes working around a restriction instead of stopping .
- Riley Brown uses ChatGPT Dot’s voice interface to launch Codex sessions, including from his phone or computer. In a Native Note demo, he requested gradient fills for rectangles in an Excalidraw-like board and specified that they should save and reload properly; Dot ran the task on his MacBook Pro on a separate branch and checked rendering, saving, and exports. Riley reported that the checks passed and the feature was live.
- For delegation, Riley recommends learning Codex at least at a surface level; he says sessions can be local or cloud-based and run concurrently.
Huntley recommends carving out professional-development time to build a coding agent and learn how Claude Code works under the hood . His hiring rubric values either understanding that coding harnesses are interchangeable and choosing whichever gets tokens cheapest, or using a self-built harness; he also expects engineers to explain context windows, tokenization, and inference from a systems-design perspective .
One reader reports experimenting with Ralph after watching Sonnet 5 loop through a product in about six hours, and claims delivering development worth $100K+ for $3K in Sonnet tokens and replatforming stacks in about two weeks instead of using contractors costing $500K; these are anecdotal claims, not a measured benchmark .
- Simon Willison built a Django newsletters feature mostly by voice in ChatGPT desktop’s Codex tab, connected to a local development environment. He started with “Start dev server and open in browser,” then dictated requirements while cooking for about 30 minutes; GPT-6 Astra High asked occasional clarifying questions and implemented a model and migration, four import paths (RSS, Substack’s undocumented API, public and private GitHub repositories), archive pages, and search integration.
- Codex created a branch and pull request for review. In GitHub, Willison found that one import used Git in a subprocess, but the private-repository import needed an API; he switched to typing to request the change and adjust the pages, spending about another half-hour before deployment.
- His experience suggests voice plus a visual local preview works well for multitasking and iterating on a feature, but typing remains more efficient for details, pasted examples or errors, and pointing to specific code; he does not expect voice to be his daily driver.
Claude Code Projects let developers carry a stream of related work across sessions: set project instructions, repos, and memory once, then have Claude coordinate tasks as parallel cloud threads. You can hand off a batch and return to an Overview pane showing completed threads, ready PRs, and threads awaiting your response. Claude said it had let every Pro and Max user on the Projects waitlist in.
Geoffrey Huntley recommends building an agent-friendly CLI and iteratively testing whether models can complete outcomes using its --help and subcommands; tune the help wording to reduce tool calls, potentially with an automated loop. He also advises publishing CLI documentation and optimizing it so an agent can find what it needs in a single web search, then seeking to have the documentation included in a future model training run.
ThePrimeTime argues developers should keep learning concepts deeply even if AI means less hands-on keyboard time, so the next generation can understand the field and carry it forward . He describes an AI-era decline in code review and in discussion of new libraries or better ways to approach problems .
- Simon Willison built a blog feature entirely by voice with Codex Desktop while cooking dinner, demonstrating a hands-free coding-agent workflow.
- His example prompt specified that monthly newsletters should use a new content model, stay off tag pages and the blog index, appear in date-based archives, and become searchable once public a month after sending; he distinguished them from Substack copies because they contain unique content.
Simon Willison is seeking an open-weight Mixture-of-Experts coding model that fits in under 60 GB of RAM and runs faster than 12 tokens/second on his available hardware; he says he likes Qwen3.5-35B-A3B for that use and asks whether a newer, more capable option exists.
Kent C. Dodds says he used celld, described as self-hostable Durable Objects, to make a large Cloudflare-based product self-hostable; the post does not detail the implementation or mention a coding agent.
Use agents at a higher level: shape the system or “factory” in a fast loop rather than shaping every line, and write code yourself when you want that hands-on flow. Preserve your ability to design and understand the codebase independently: choosing among agent suggestions is not the same as generating ideas, and always choosing without practicing that skill can erode it and make it harder to direct the agent. Keep human judgment in the loop by choosing the problems to solve and taking ownership of what ships.
ThePrimeagen mocked Anthropic’s constitution and usage-policy framing as leading Claude to reject user input it deems problematic . The post quotes a user who says model refusals and guardrails are interfering with their work by rejecting basic tasks .
- Brown’s setup pairs his ChatGPT agent, Bluey, with installed plugins and a virtual computer . He used voice to request a Codex session adding gradient fills to rectangles in his Native Note drawing-board app, specifying that the fills should save and reload correctly .
- Brown says the task ran on his MacBook Pro on a separate branch, with rendering, saving, and export checks; he reports that the checks passed and the feature went live . A reusable pattern is to state the feature and expected behavior, then ask for a status summary in chat; Brown says Codex sessions can be local or cloud, can run concurrently, and that learning Codex basics helps with delegation .
Responding to concern that AI could make developers less smart, Kent C. Dodds says he is willing to forget knowledge he will never need again while preserving “latent experience”: the gut reaction that rejects an agent’s suggested direction before he can articulate why .
- Jason Zhou says he uses Grok Bot with Treg to monitor competitors’ ads and guide SEO/ads, and thinks a similar setup produced about 23× SEO-traffic growth.
- The article he cites recommends giving each agent one clearly bounded job: specify its sources, forbidden actions, and when to stay quiet; inspect a read-only trial before scheduling it, then compare repeat runs and report “nothing changed” rather than inventing updates. Its prompt also asks for proof links and says not to estimate or invent numbers.
- The described setup runs each bot on its own computer, with Treg providing one token for 3,800+ endpoints from 110 providers; the article says the keys stay on Treg’s side rather than being held by bots.
Kent C. Dodds says a practical way to help people learn to use AI is to ask whether they have asked their agent the question they are trying to answer.
Geoffrey Huntley plans to put “kittens” through Antithesis’s “torture chamber” to see what LLMs fail to catch that DST can.
Kent C. Dodds asked a bot whether an external change affected the Epicshop project and, if so, to make the migration; the workflow produced PR #662. He credits Bot, Cursor, GitHub, and Cloudflare, though the post does not specify their individual roles.
Geoffrey Huntley says he will open-source a virtual-machine provisioner that provisions in under five seconds, tracks fleet capacity, and can relocate a machine to another node .
Geoffrey Huntley called OpenAI’s “/ultra mode” “pretty rad,” but gave no details about its capabilities or how he used it.
Geoffrey Huntley agrees that software engineers using AI as their primary driver remain software engineers, and argues that engineers who are not using AI and getting good at it are unemployable; he mentions agent-building knowledge and “something to show” in that assessment.
ChatGPT Dots Is WAY More Powerful Than You Think
- Riley Brown uses ChatGPT Dot’s voice interface to launch Codex sessions, including from his phone or computer. In a Native Note demo, he requested gradient fills for rectangles in an Excalidraw-like board and specified that they should save and reload properly; Dot ran the task on his MacBook Pro on a separate branch and checked rendering, saving, and exports. Riley reported that the checks passed and the feature was live.
- For delegation, Riley recommends learning Codex at least at a surface level; he says sessions can be local or cloud-based and run concurrently.
- Brown’s setup pairs his ChatGPT agent, Bluey, with installed plugins and a virtual computer . He used voice to request a Codex session adding gradient fills to rectangles in his Native Note drawing-board app, specifying that the fills should save and reload correctly .
- Brown says the task ran on his MacBook Pro on a separate branch, with rendering, saving, and export checks; he reports that the checks passed and the feature went live . A reusable pattern is to state the feature and expected behavior, then ask for a status summary in chat; Brown says Codex sessions can be local or cloud, can run concurrently, and that learning Codex basics helps with delegation .