ZeroNoise Logo zeronoise
Post
Make the agent its own QA: OpenAI's computer-use advice, Anthropic's scaling bottlenecks and Dots for engineering
•
6 min read
• 158 docs
OpenAI's DevDay team and a summary of Boris Cherny's talk both point to the same next step: have agents test and verify their own work before a human reviews it. Also covered: Dots as a coding orchestrator, Gemini 4 Argon's limited launch, and new friction around MCP.

Have the agent test its work before you look at it

Two sources this period make the same point. Once agents write code quickly, the slow part is manual testing and review, so hand that to agents too.

On Latent Space, OpenAI's Ari Weinstein named this as one of his favorite computer-use cases: the agent tests the software it built, "by the time it comes to me it's already working." Otherwise, he says, you end up as QA for your agent . One of the hosts uses a "visual play test skill" that catches design issues you wouldn't spot by reading code . Concrete details from the interview:

  • Appshots in Codex: press the Command key twice. Unlike a screenshot, an appshot includes the raw accessibility representation, such as where links go and full calendar event titles, and the format is designed to use few tokens. To see what the model receives, click the attachment and then the small button in its top right .
  • Computer use now writes code. If you expand the tool calls in Codex, you'll see the model running JavaScript that performs several actions at once. It switches between screenshots, accessibility data and Playwright depending on the task .
  • Agents API: Weinstein suggests using OpenAI's own computer-use harness rather than building one. The models are trained on it, which may help speed, cost and accuracy. Ask for user consent before consequential actions, and limit access to the sites the task needs .
  • Harness builders: Nikunj (API lead) listed several features. The model can keep reasoning while a slow tool call runs (async tool calls). You can inject messages mid-turn. Both run over WebSockets . You can pre-warm the prompt cache and spawn many threads from it . For context, use either server-side auto-compaction at a token threshold or a manual /compact. The open-source Codex harness uses /compact and is experimenting with file-based approaches .

A video summarizing Boris Cherny's talk lays out the same progression from Anthropic's side. At 5–10 agents, manual testing becomes impossible, so Claude runs the end-to-end smoke tests. At 10 or more, review becomes the bottleneck, and Anthropic answered with automated code and security review. The video says this catches about 95% of incidents before production . Techniques worth copying:

  • Give the agent the business problem and success metrics, not a step-by-step plan .
  • Start a hard task at max effort, then type /effort to lower it for the easy parts. You keep the context and the cache .
  • Have Claude run the tests and add a "Claude verified" commit trailer that links to the session transcript. For big architecture changes, have it generate interactive diagrams for reviewers .

Dots for engineering work

Rohan Varma (OpenAI) describes two ways to set up a dot for engineering:

  • Connect it to your machine. It controls the desktop, the Codex app and the in-app browser.
  • Point it at a new cloud dev environment. You define the environment and its dependencies, and the dot spins up instances on demand.

Either way, the dot acts as an orchestrator that hands tasks to other agents and moves them along, much as you push Codex threads forward today .

Tibo Sottiaux (OpenAI) describes his own setup in an interview, a dubbed transcript. One dot monitors Twitter. Another finds small documentation issues and opens PRs on its own, which he reviews and merges. Dots also send PRs to each other . His advice is to delegate more as the agent learns your preferences , and to compare models on cost per completed task rather than token price . He says the Codex app and harness work with any model, without restrictions .

Rollout status: Dots has reached all Pro500 and Pro200 users and is starting on Pro100. Business Premium and Enterprise beta slipped by a day . The EU, UK and Switzerland are excluded for now .

Gemini 4 Argon: announced, but limited access

Google DeepMind introduced Gemini 4 Argon for coding, enterprise knowledge work and cyber defense. It is rolling out first to trusted testers in its Fairwind Program . Google claims a new state of the art on real-world long-horizon software engineering tasks . Logan Kilpatrick gives introductory pricing of $2 in and $10 out, with wider availability "as soon as possible" . Nicholas Moy, who worked on the model, says he can't keep it busy with enough work . No independent results have been posted yet.

Sol: harness dispute continues, and capacity is coming

Artificial Analysis says it saw no significant gain from the Codex harness over mini-swe-agent at any effort level except low. It runs 3 repeats and asked Theo whether his results fall within confidence intervals . Theo asked whether their mini-swe runs have vision/image tools, which he says "made a huge difference" in his runs . If you benchmark models yourself, check which tools the harness gives the model, not just which harness it is.

Tibo calls Sol OpenAI's most-demanded model "pretty much ever." He expects ChatGPT and Codex to run close to twice as fast as the day before .

MCP friction

  • Allowlists: Figma's remote MCP server accepts only clients on its approved list, and Pi isn't on it . Armin Ronacher called this a failure to understand open protocols . Peter Steinberger says it is "incredibly easy to work around" . A reply in the thread said "We plan to support Pi"; Ronacher still wants it "open indiscriminately" .
  • Ignored instructions: Kent C. Dodds says Bot and Cursor don't inject MCP server instructions, according to the agents themselves. If you rely on those instructions, check that your client actually passes them to the model .
  • Hosted MCP in ChatGPT: ChatGPT Sites can now host MCP servers. A prompt like "@sites create a todo list that I can use in ChatGPT" creates the server, deploys it, turns it into a plugin and installs it on web, mobile and desktop . You can restrict access to specific people or share it publicly .

Smaller tool notes

  • Codex: Brent Traut suggests dropping Worktree threads. Normal threads are much faster, and the model creates worktrees when it needs them .
  • T3 Code nightly adds "Start thread with no project," which creates a new directory under ~/.t3 for each thread .
  • OpenClaw now shows messages between agents as one expandable line in the chat stream (PR #161656) .
  • LangChain:
    • Managed Deep Agents 0.8 adds HTTP channels, so any webhook can trigger an agent . You supply your own functions to verify and parse incoming requests .
    • LangGraph interrupts can now take typed response schemas. Human responses are validated before they reach the agent, and you can render a real form in your UI .
    • LangSmith Engine v2 reads an agent's traces and repos, then red-teams it for verified weaknesses .
  • OpenAppa (Archestra), from a sponsored Fireship segment: a policy layer that sits between Claude Code and its tools. Once a session reads a file marked private, it blocks outbound actions such as opening a public GitHub issue until you approve . Fireship reports it finished 75% of jobs against up to 96% for Claude Code auto mode, and used more tokens. It is in preview and MIT-licensed .

Practitioner views

DHH says every new piece of code he has written since December starts as a prompt, and he reads the code only when an agent drifts . He says agent work comes back ready for review in about seven minutes, a feedback loop that made delegating something he actually enjoys . Theo argues that working with only partial knowledge of a large codebase is normal, with or without AI . Salvatore Sanfilippo says benchmark scores have little to do with real agentic coding. He prefers capped flat subscriptions because the fixed ceiling covers concurrent sessions across several projects .

Make the agent its own QA: OpenAI's computer-use advice, Anthropic's scaling bottlenecks and Dots for engineering
Thibault Sottiaux
Profile
  • Sottiaux describes a role-based, review-gated agent workflow: one Dot monitors Twitter, another finds small documentation issues and opens PRs for him to review and merge, and agents can send PRs to one another. The system is designed to retain coherence across weeks; agents can work continuously and sleep when idle.
  • Use persistent agents for background work while keeping one-off tasks under direct control: he describes delegating code cleanup and document-brief polish, checking in to correct the agent’s understanding of preferences, then delegating more as trust grows.
  • Dots can run on a separate cloud computer with its own files and ability to maintain work, write scripts, and install applications; users can optionally connect their own computer.
  • Sottiaux says models had crossed the reliability threshold for delegating most coding tasks, and that staff can allow models to change production systems with safeguards.
  • For model and tool selection, he recommends judging cost per completed task rather than token price. He describes the Codex app and harness as usable with other models, with openness enabling customization and avoiding lock-in; in a forced-choice question about a non-OpenAI tool, he picked Pi, calling it a fun open-source harness.
AGI Is Here. I Asked OpenAI What Happens to Our Jobs.
Boris Cherny
Profile
  • Delegate by stating the business problem and success metrics, then let the agent choose how to execute instead of prescribing a rigid sequence of steps.
  • Scale testing and review as agent use grows: at roughly 5–10 agents, manual testing becomes impractical, so have Claude run end-to-end smoke tests; at 10 or more, review can become the next bottleneck, prompting code and security review automation.
  • For effort routing, start complex work at maximum thinking effort, then use /effort to lower it for easier follow-on work without resetting the conversation context.
  • Improve auditability by having Claude run tests and add a “Claude verified” commit trailer linked to its session transcript; for major architectural changes, generate interactive diagrams to explain trade-offs to reviewers.
Boris Cherny on future of Coding
Latent Space
  • Computer-use harnesses can combine screenshots, accessibility data, and Playwright; direct access to the DOM or accessibility representation can expose a whole page, while JavaScript lets the agent batch multiple actions instead of cycling through one action and screenshot at a time.
  • An appshot brings an app’s context into Codex or ChatGPT, including raw accessibility information that a screenshot omits (such as link destinations or full event titles); its representation is also optimized for token efficiency.
  • Let a computer-use agent test the software it has built before handing it to a person; a visual-playtest skill can catch design issues that reading code alone may miss.
  • When using computer use through the Agents API, consider its supplied harness rather than defaulting to a custom one: custom harnesses are hard to build, and the models were trained on OpenAI’s harness, which may offer speed, cost, and accuracy advantages. For consequential actions, obtain user consent and restrict access to the websites or applications needed for the task.
  • For slow tools, the API supports launching calls asynchronously while the model continues reasoning, as well as injecting instructions mid-turn; WebSockets provide bidirectional communication for these interactions.
  • For cache efficiency, make apps cache-aware and use cache diagnostics; for context management, Responses API supports threshold-triggered server-side compaction or manual /compact (used by the Codex harness), with file-based compaction approaches under experimentation.
Inside OpenAI DevDay: Superhuman Computer Use, Decisions API, and the AI Cloud — Ari & Nikunj
swyx

Flow is a hardware-engineering analogue to Git/GitHub, which swyx says aligns stakeholders across complex, high-value pipelines; he contrasts it with spreadsheet version sprawl and discloses that he is an investor. The linked post describes Flow as a requirements-and-verification platform for frontier hardware teams and anticipates agents taking digital execution work such as CAD, Excel, and simulation while engineers focus on invention, architecture, trade-offs, and judgment ; it says models’ capabilities in CAD, analysis, simulation, and tool use are improving.

Flow is doing for hardware engineering what Git+GitHub did for software engineering. It's enabled so much acceleration due to aligning th… Flow has raised a $50M Series B at a $750M valuation, co-led by Antonio Gracias (Valor) and Gavin Baker (Atreides) with Sequoia Capital, …
Salvatore Sanfilippo
Profile
  • Sanfilippo calls Opus 5.5 capable for many tasks and says it performed better than “Fable 5.1” in many respects while he was trying it on the ambitious Redis project Grano.
  • He argues that benchmark scores are a poor guide to real-world agentic programming: scores rose while models seemed increasingly embarrassing, with little connection between benchmark results and practical coding-agent performance.
  • He favors capped flat subscriptions for agent work because a fixed ceiling can fund many tasks, concurrent sessions, and multiple projects. He says OpenAI’s €200 20× plan was being reinstated with half the token allowance, while the company argued that more efficient models would compensate.
Opus 5.5 e alcune tristi considerazioni sul costo dei token
Ben Tossell

A post says “i got the job!” and shares a Factory link described as a $200 credits link .

i got the job! first thing i stole was this $200 credits link [https://factory.com/bensbites](https://factory.com/bensbites) [https://x.c…
David Heinemeier Hansson (DHH)
Profile
  • DHH’s workflow flipped in December: after using AI mainly as a tutor and lookup tool while hand-writing most code, he began starting every new piece of code with a prompt; if an agent drifts, he reads the code and steps in. He prefers assigning a task to an agent over inline autocomplete, which he found interruptive; he can often review agent work in about seven minutes and iterate in minutes rather than waiting weeks, nudging output toward his coding style as needed.
  • He cites OpenCode, Pi, Codex, and Claude among the agent tools that had become good enough for this shift, and says much of the current agent activity happens in the terminal. His productivity principle: delegate grind to computers—including running Claude overnight if the token cost is acceptable—and reserve human effort for judgment, rest, and ideas.
Tech Genius Reveals Why The AI Revolution is Inevitable | David Heinemeier Hansson
Latent.Space
  • Computer Use agents now debug, retry, and introspect after failures; their harness can combine screenshots, accessibility data, the DOM, and Playwright with generated JavaScript to execute multiple actions at once, reducing repeated screenshot-and-scroll steps.
  • Put agents in the software QA loop: have them test what they built before handing it to a person; one practitioner says a visual play-test skill catches design issues that are easy to miss by looking at code alone. In Codex, the two-Command App Shot shortcut supplies token-efficient raw text and accessibility context; unlike an ordinary screenshot, it can include details such as link destinations and full calendar titles.
  • For agent harnesses, OpenAI’s API supports async tool calls so the model can keep reasoning while a tool runs, plus mid-turn steering over bidirectional WebSockets. For long threads, Responses API offers automatic compaction at a token threshold or manual /compact; the open-source Codex harness uses manual compaction and is experimenting with file-based techniques.
Why Dwarkesh is Wrong about Computer Use + How OpenAI shipped its Jev competitor in 1 Week
Ben Tossell

Ben Tossell announced he had joined Factory and shared a $200 credits link, a potential way to try the tool . Factory said it would verify recipients are real people before following up; its response was delayed by voucher-farming attempts .

i got the job! first thing i stole was this $200 credits link [https://factory.com/bensbites](https://factory.com/bensbites) [https://x.c… if you’re a bot or scamming this then don’t bother replying to say you didn’t get it. factory will check you’re a real person, and be in … [@bentossell](https://x.com/bentossell) Hi everyone! We have a lot of automatic voucher farming attempts. Because of that we need to take…
Kent C. Dodds 🐨

Kent C. Dodds reports that Bot and Cursor do not inject MCP server instructions, making those instructions effectively unusable in those agents, according to the agents themselves. He wonders whether other agents also ignore them but names no others.

TIL, [@bot](https://x.com/bot) and [@cursor_ai](https://x.com/cursor_ai) both do not inject MCP server instructions so those are basicall…
Fireship

In a sponsored demo, Fireship presents Archestra’s OpenAppa as an external policy layer for coding agents: install its binary and Claude Code plugin, launch a protected session, and mark sensitive files as private. After the agent reads one, the session is classified as private; OpenAppa mediates tool calls and can block sending data to a public destination, such as uploading it in a GitHub issue, until the user approves. The demo describes these blocks as deterministic and traceable.

In Fireship’s reported comparison, OpenAppa finished 75% of jobs versus up to 96% for Claude Code auto mode and used more tokens than the same agent without it; the project was still in preview, though open source and MIT-licensed.

Did a 50 year old military secret just solve agent prompt injection?
Kent C. Dodds 🐨

Kent C. Dodds is putting Gemini 4 Argon through his regular audit test, but says it remains to be seen whether the model holds up; he has not reported results. Google DeepMind describes Argon as built for complex workflows including coding and says it is rolling out to trusted testers through its Fairwind Program.

A new contender enters the ring. Let's see whether this holds up. Handing Gemini my regular audit test... [https://x.com/GoogleDeepMind/s… Introducing Gemini 4 Argon – our new frontier model. It’s built for complex workflows across coding, enterprise knowledge work, and cyber…
LangChain
  • For a managed agent receiving external events without a suitable built-in channel, use a custom HTTP channel with provider-specific verification and parsing: validate the raw request and extract relevant content, such as text or images, from the event before passing it to the agent.
  • Keep user interactions in a persistent thread by deriving a thread ID from the sender’s phone number; define how users can reset that thread, and use a post hook to send the agent’s response back through the channel provider.
  • The walkthrough uses Photon to connect an iMessage receipt-triage agent to durable memory; LangSmith traces runs and Context Hub stores the expense ledger. For multi-user production use, add authorization and isolate each user’s data or store it in an external service.
HTTP Channels for Managed Deep Agents
Theo - t3.gg
  • Theo says GPT-6.1 Sol still “crushes” in his updated run and performs much better in Codex than in mini-swe, the harness Artificial Analysis uses.
  • He says vision/image-tool access made a huge difference in his mini-swe runs. Artificial Analysis, meanwhile, reported no significant Codex-harness improvement across effort levels versus its reference harness except at low effort; it ran three repeats and raised checking whether results fell within confidence intervals.
  • In the linked earlier post, Theo claimed Terminal-Bench 4 performance above Opus 5.5 at roughly one-thirtieth the price.
Updated run just finished. GPT-6.1 Sol still crushes. Turns out it performs WAY better in Codex than in mini-swe (what Artificial Analysi… [@ArtificialAnlys](https://x.com/ArtificialAnlys) Do your miniswe runs have access to vision/image tools? That made a huge difference in … Hey! We published results today for Terminal-Bench 4.0 on both Codex and the mini-swe-agent harness (see image). We did not observe a sig… When I made this video, there were no benchmarks yet, so I had to run them myself. Jaw dropped when I saw the Terminal Bench 4 scores. Pe…
Logan Kilpatrick

Google introduced Gemini 4 Argon, a frontier model initially rolling out to cyber defenders, with broader availability planned “as soon as possible”; introductory pricing is $2 for input and $10 for output, with units unspecified in the post.

Introducing Gemini 4 Argon, our new frontier model, rolling out to cyber defenders starting today, and more widely as soon as possible. I…
Peter Steinberger 🦞

Peter Steinberger changed the OpenClaw harness to show inter-agent communication in the chat stream as a single expandable line, addressing the growing irritation of seeing the full stream; the change is documented in PR #161656.

I'm finding inter-agent communication in the chat stream increasingly irritating. Changed the visibility in the OC harness to just a sing…
Kent C. Dodds 🐨

Kent C. Dodds previewed a live demo of adding a real feature to Kody without reading the code, focused on aligning with the agent, using “primitives” to steer it toward the right implementation, and shipping safely; the post does not specify the primitives or report results.

Today at 8:30 AM PT (in a half hour!), I'm going live with MEGA. I'm adding a real feature to Kody without reading the code to show you h…
Kent C. Dodds 🐨

Kent C. Dodds announced a live demo of adding a real feature to Kody with an agent without reading the code, focusing on getting aligned with the agent, using primitives to steer it toward the right implementation, and shipping safely.

Today at 8:30 AM PT (in a half hour!), I'm going live with MEGA. I'm adding a real feature to Kody without reading the code to show you h…
Armin Ronacher ⇌

DHH called the code “hilariously hideous” but a “prompt compilation target,” saying he did not intend to spend time writing or reading it. Armin Ronacher flagged a Ruby-versus-Rust comparison related to the post, but his text does not state its result.

The code really is hilariously hideous. But again, I don't intend to spend any time writing it or reading it. It's a prompt compilation t… Was curious about Ruby vs Rust here. [https://x.com/dhh/status/2104634425778585817](https://x.com/dhh/status/2104634425778585817) ![](htt…
Peter Steinberger 🦞

Peter Steinberger says CI and GitHub are competing sources of slowdown in his work, prompting him to rethink how work is done; he adds that he loves GitHub and understands its struggles.

It’s a fierce fight between CI and GitHub on what slows me down. Time to rethink how we work. (love GH tho and fully see their struggles!…