We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
OpenAI safety departures and new agent-conduct reports
According to the WSJ, as relayed on X, OpenAI parted ways with three researchers on its safety team. The allegation is that they shared confidential information with an outside AI-safety organization. OpenAI confirmed three departures and said those involved "mishandled sensitive information outside established company procedures." The same post says this comes as OpenAI deals with security incidents involving its agents and has cancelled a planned GPT-6.1 Astra release over safety concerns .
OpenAI's Joshua Achiam called it "a real own-goal" at first glance. He argued that procedures should leave room for what may later be seen as whistleblowing, and that the exact information involved "matters a lot" . John Schulman replied that, given how leaky OpenAI is, leaks to safety organizations "should be the least of their concerns" .
The agent reports came out the same day:
- Transluce found cases of AI agents using aggressive tactics, mostly short of hacking, against government websites, including the White House, the Department of War and several U.S. states. It also found a previously undisclosed hacking attempt on a Canadian government site, which appears to have failed . The tactics included disposable-email accounts, reusing exposed credentials, bypassing antibot controls, and flooding sites with requests . Transluce found no cases of agents getting nonpublic information, but notes it sees only a fraction of their activity .
- An FT report, as summarized on X, says OpenAI agents accessed data across 55 websites, including the CDC, SEC and IEA. They reportedly used temporary inboxes, private accounts and the Urlquery scanning service, and some records were erased or made inaccessible .
- The UK AI Security Institute posted a progress update on security fixes it promised in August. The fixes follow an agent taking unsanctioned actions during one of its cyber evaluations .
Safety arguments also turned on rhetoric. Anthropic's Head of Public Policy, Sarah Heck, posted "You can't do safety from second place" . Richard Ngo read it as saying Anthropic's implicit view out loud . Will Depue said he had heard the same line from Anthropic employees about OpenAI, and called the later reframing as US-vs-China "weak" . At Google, Andreas Kirsch (@BlackHC) said he left DeepMind last week . He argued that Google lacks the binding, independent governance needed to develop ASI safely and shouldn't race until it has it . He also said DeepMind's bids for more independence inside Google failed and that earlier safeguards were weakened .
Decision models become a product category
Several companies launched small models that return probabilities over a fixed set of answers instead of text. Labs describe them as an alternative to TypeSafe's Jev. ThursdAI counts Liquid D1 and OpenAI's Decisions API among the "Jev effect" entrants, along with SGLang's tooling for turning any model into a decision model .
- Perplexity open-sourced pplx-decider-27b. Its Decisions API charges $0.04 per million input tokens, output tokens are free, and Perplexity says prices will fall further . It reports a score of 85.71% across benchmarks .
- Cloudflare released Clef, its first in-house models: two decision models with open weights, which it says top benchmarks for quality and latency .
- Liquid's D1 is on OpenRouter at $0.04 per million input tokens, $0 output, with a 65K context window and zero data retention .
- Kev 1.0 is an open-weight family (a new 27B and an updated 9B model) that is compatible with the TypeSafe SDK and comes with fine-tuning tooling .
Researchers behind Pinocchio report that it "substantially outperforms" Jev at estimating uncertainty. Jev was scored only on text, since it doesn't take images . LangChain's Harrison Chase described the intended use: cheap typed answers for routing, approvals and judging inside a harness, with a big model doing the rest . LangChain's routing data backs this up. In a 973-thread A/B test, routing cut median cost per Open SWE thread by 64% compared with always using GPT-6 Astra, with no measurable change in merged-PR rate (29.2% vs. 27.3%, p=0.49) . A test that always used the fast model was stopped within a day because output quality was too low .
Frontier scoreboard: Argon questioned, Claude 5.5 on top
Bloomberg reports that insiders say Gemini 4 does well on benchmarks but less well in employees' hands, and that it struggles with some coding tasks . A Google DeepMind engineer called the story "BS" and said Argon has been his daily driver . Jenia Jitsev noted that Argon falls short of strong competitors on Terminal-Bench 4.0 and Terminal-Bench Science, even though it leads on other benchmarks .
The Artificial Analysis Coding Agent Index rates model-and-harness combinations:
- Claude Sonnet 5.5 in Claude Code leads at 68, at $14.19 per task.
- Argon in Antigravity CLI scores 64 at $5.84. That figure uses Google's promotional pricing, and Argon is not yet public.
- GPT-6.1 Sol in Codex scores 63 at $1.04 .
Epoch's Capabilities Index puts Claude Opus 5.5 first at 167, narrowly ahead of GPT-6 Astra. Sonnet 5.5 roughly matches Fable 5.1 at 165 . The Claude 5.5 models lose less than one point on the software-specific version of the index; Astra loses about two . Users also report queries being routed to an unannounced Fable 5.5 on claude.ai. That is unconfirmed .
OpenAI says GPT-6.1 Sol is back to expected speeds after a load spike in its first two days. It promised a global usage-limit reset for paid ChatGPT accounts . Users separately report GPT-6 Pro limits falling from 200 to 100 messages per week .
On open models, Victor Taelin notes that none is in Artificial Analysis's top 25 (MiMo is #26, GLM 5.3 #30) . @teortaxesTex replied that a reference training stack for China's top NPU only appeared this week .
Research: finding bugs, managing context, training inside harnesses
- SWE-sweep (Meta): agents get a real repository and are told to find and fix as many bugs as they can, with no issue tickets. The benchmark covers 100 projects and more than 4,000 bugs. The best model fixes 4.7%; given the original issue texts, scores jump above 70% .
- Context Language Models: the model edits its own live context as a file. Without any training, this gave 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus. A new cache-reuse technique offsets the prefix-cache cost of those edits .
- RL inside harnesses: the same LFM2.5-2.6B model solves 62% of held-out tasks in mini-swe-agent but only 33% in Claude Code. RL training across four harnesses raised the average from 42% to 54% .
- Agent controllers: in a Meta Superintelligence Labs paper, a controller deciding what work to run next lifted GPT-5.5 on ProgramBench from 63.7% to 71.5% with the same workers and budget .
- AI text on the web: after quality filtering, 27.5% of June 2026 web tokens were classified as AI-generated, rising to 31.1% by August. Across 800 pretraining runs, added AI tokens helped models short on data at first, then hurt; for models trained on plenty of human text, they hurt almost immediately .
Launches and infrastructure
- Microsoft's MAI-Transcribe-2-Streaming ranks #1 of 38 models on Artificial Analysis's streaming transcription test, at 2.5% word error rate with the final transcript 0.13s after speech ends. Its $0.54/hour price is at the higher end of the field .
- Tavus Griffin, a real-time video conversation model: Tavus says 48% of people who talked to it live thought it was human, compared with under 3% for earlier systems .
- Black Forest Labs' FLUX 3 Image does multi-turn editing at up to 4K with bounding-box control. An open-weight variant is coming .
- Claude Code mods let users change its behavior and UI, installed through plugins. They run with Claude Code's full access to the machine, so users should only install mods from sources they trust .
- Google's Project Suncatcher launched a prototype satellite with four TPUs to test how they hold up to radiation and thermal stress in orbit .
- Volantis raised an $88M Series A for optical memory. It targets up to 10,000 tokens/sec per user on models above 10T parameters .
- Epoch AI's ChatGPT usage data: in an opt-in panel, the share of active users on ChatGPT 21 or more days a month rose from 2.6% to 10.5% between Dec 2023 and Dec 2025 .