ZeroNoise Logo zeronoise
Post
OpenAI’s Agent-Swarm Navier–Stokes Claim Puts Validation at Center Stage
3 min read
1560 docs
OpenAI’s claim that a 10,000-agent system produced a Navier–Stokes result leads a brief on validation, personal agents, scientific tooling, model releases, capital and government scrutiny.

Top Stories

Why it matters: AI progress is now being measured by coordinated systems that spend compute over hours, not only by single-model scores.

OpenAI claims an agent-swarm result on Navier–Stokes. It says a next-generation model, described as significantly more capable than GPT-6 Astra, produced an analytical proof and Lean formalization that a smooth fluid can develop a finite-time singularity. OpenAI says about 10,000 agents reached the result in 88 hours, followed by 17 hours of Lean verification. The construction uses smooth external forcing; a contemporaneous account says that is allowed under the Clay formulation but does not settle the harder unforced question, and qualifies the result “if the proof holds up.” OpenAI says it does not intend to claim the Millennium Prize, calls the release a snapshot, and cannot rule out de-identified product-derived data improving its models even though no specific user data or the researchers’ work was accessed. Validation and provenance are therefore part of the story, not footnotes.

Meta is putting personal agents in front of users. Muse is always-on, browser-using and app-connected; each instance runs in an isolated secure VM, a Sentinel checks actions before they leave, and it asks before sensitive email or spending actions. Meta offers a public bounty of up to $300,000 for security holes or impactful prompt injection and free use up to 100M tokens per week. Permissions and isolation are being shipped as core product features.

Research & Innovation

Why it matters: The most useful technical releases are expanding context—from biology to enterprise document stores.

Google DeepMind’s AlphaGenome Atlas maps the predicted impact of all 9 billion single-letter DNA changes in a 1-petabyte database more than 30 times the size of AlphaFold’s. Its AVI score ranks mutations and surfaces effects such as broken gene switches or RNA splicing, with access via the web, API and Antigravity.

Harvey and Baseten’s recursive-language-model harness loads up to 5,000 documents or 80M tokens, then delegates bounded reviews to sub-agents. Across seven models, mean rubric pass rate rose from 23.3% to 62.4%; RL raised a Qwen orchestrator from 29.9% to 63.0% on 50 held-out rooms and coverage from 62% to 96%, at higher generation cost for six of seven baselines.

Products & Launches

Why it matters: Product competition is moving toward controllable outputs and rapid iteration.

ChatGPT Images 2.5 adds faster generation, improved fidelity, edit consistency and comment-based edits, rolling out across ChatGPT, Work and Codex; Flare and Sunburst are new API models.

DeepSeek V4.1 Flash is reported in API beta as an intermediate build with a new architecture, native multimodality and faster/stronger performance; its expiring model ID points to a rapid test cycle rather than a settled release.

Industry Moves

Why it matters: Frontier compute and enterprise agents are attracting capital at the same time.

Mistral raised €3B in a Series D it calls Europe’s largest tech equity round; Samsung led, with EQT and PSG co-leading, to fund compute, infrastructure and frontier research.

Cognition raised more than $2B at a $48B valuation and says run-rate revenue rose from $492M to nearly $900M since May. Devin is expanding into proactive Slack remediation and security scanning.

Policy & Regulation

Why it matters: Government scrutiny is extending from model behavior to control of frontier capabilities.

An NSA post points to a new report co-sealed with the FBI and CISA alleging illicit distillation of U.S. frontier capabilities by China-based AI companies; it highlights tactics, techniques and recommended mitigations.

Quick Takes

Why it matters: Serving architecture and distribution are now part of the capability race.

  • Agent serving: vLLM’s coding-agent traces show 43 turns per session, 142K-token inputs, 96%+ prefix-cache hits and subagent forks in 44%; it reports 83K tokens per GPU-second for DeepSeek V4 Pro and a 106× cost gap versus Opus 5.
  • Astra availability: OpenAI says GPT-6 Astra is fully rolled out to Plus, Pro, Business and Enterprise users in Codex and ChatGPT Work.
  • Creative tooling: Runway plugins now bring model-based generation, editing and upscaling into Premiere Pro and After Effects.
OpenAI’s Agent-Swarm Navier–Stokes Claim Puts Validation at Center Stage
Research extraction

Bottom line. OpenAI reports that an internal system produced an analytical proof and Lean formalization showing that a three-dimensional, incompressible, constant-density Navier–Stokes flow can develop a finite-time singularity; the post says this establishes statements “C” and “D” of the official formulation and therefore resolves the Millennium Prize problem.

  • Scope of the claimed result: The singularity is defined as fluid speeds growing without bound in finite time, despite viscosity. The constructed case begins with a smooth fluid at rest, applies a smooth force, and retains finite energy throughout the evolution up to singularity formation.
  • Proof idea: OpenAI describes an inward-spiraling, increasingly elongated vortex whose central region shrinks while its speed rises. The claimed mechanism does not insert an infinite force; instead, the acceleration, pressure-gradient, momentum-transfer, and viscosity terms become large but cancel precisely enough to leave a smooth external force.
  • Methodology: The work used a coordinating multiagent system based on an internal model, with code execution, cached-internet access, inter-agent communication, monitoring, and isolation. The Navier–Stokes group involved roughly 10,000 concurrent agents; separate groups were prompted with proof-oriented variants “A” and “B” and disproof-oriented variants “C” and “D.” Different approaches were cross-pollinated using Codex to consolidate useful intermediate insights.
  • Timing and scale: OpenAI says the agents reached the result on September 5, about 88 hours after launch; Lean formalization and verification took an additional 17 hours via GPT‑6 Astra. For the Navier–Stokes effort, the agents sent 2.7 million messages and used approximately 130 billion output tokens. The underlying model’s training was still ongoing and its performance was continuing to improve.
  • Formalization caveat: The announcement explicitly reports a Lean formalization and verification, but the verification account supplied here is OpenAI’s own description of that internal process; the post does not characterize it as a prize award or external adjudication.
  • Important qualifications: OpenAI says it does not intend to claim the Millennium Prize, describes the milestone as a snapshot of progress, and credits mathematicians and AI researchers alongside the agents. Keep the separate Euler result distinct: OpenAI describes its own result as an unforced Euler regularity disproof, while the concurrent outside work it discusses concerned forced Euler; OpenAI says no specific user data or outside work was accessed before public release, but cannot rule out de-identified product-derived data having improved its models.
On the Navier–Stokes Millennium Prize Problem | OpenAI
AI High Signal

Researchers reportedly used AI models to build a mobile worm that could fully compromise any WeChat account across iOS and Android in seconds; the post says it took barely more than a week to build and presents the result as evidence that AI is rapidly changing cybersecurity.

New: Researchers used AI models to build a powerful mobile worm that could fully compromise any WeChat account across iOS and Android in …
AI High Signal

A critical post disputes an alleged OpenAI Navier–Stokes breakthrough, saying that 10,000 agents ran for 88 hours on a multimillion-dollar GPU cluster to formalize in Lean a blow-up case under controlled external forcing—not the Clay Millennium problem of global smooth existence and stability for 3D incompressible Euler/Navier–Stokes under natural conservation laws and viscous dissipation. The critic characterizes the achievement as brute-force autoformalization and computational parallelization rather than evidence of machine understanding, and questions its presentation as a major AI milestone without traditional peer review; these remain the critic’s allegations, not independently verified facts in the source.

go fuck yourself [@sama](https://x.com/sama) claiming that you solved navier stokes because 10000 agents ran in circles for 88hours on a …
AI High Signal
  • KVMem keeps long-running agents’ context overflow as paged KV state across GPU memory, host memory, and NVMe, then uses model-native attention-space indexes to retrieve relevant historical blocks without re-prefilling already processed text.
  • On the DeepSWE long-context test with Qwen3.8-27B, task success increased from 43.8% with compaction to 48.4%. A local setup reportedly virtualized up to 1M tokens—four times the model’s native 256K context—on a laptop with a 24GB RTX 5090 at around 50 tokens/second.
Nice paper to improve inference efficiency. It's been a while we haven't seen good work on efficiency. Here is why it matters: A long-run…
AI High Signal

@LearnOpenCV outlined a video-generation pipeline that assigns subtasks to specialized agents: Astra Ultra orchestrates; Astra High handles writing and teaching judgment; Sol High performs technical and animation reviews; Terra Medium manages routine production and repairs; Luna Low handles metadata and formatting; and Code performs hashing, audio normalization, and video assembly.

Finally time to optimize my video generation pipeline by assigning subtasks to subagents based on the complexity of the task. Astra Ultra…
AI High Signal

An embedded post alleges that Chinese technology companies “from Alibaba to Zhipu” work to some degree with Huawei, SMIC, and other sanctioned entities, making them “fair game to be hacked at will”; it further claims an NSA memo paints all Chinese AI companies as acceptable targets. This is an unverified allegation: the supplied excerpt includes no memo text, policy details, or corroboration.

Since all Chinese tech cos from Alibaba to Zhipu work with Huawei, SMIC, and other sanctioned entities to some degree, they are all fair …
AI High Signal

A post argues that pure autoregressive models are an “offramp” rather than a path to real intelligence, identifying world models and the ability to orchestrate 10,000 subagents as missing capabilities; it uses solving Navier–Stokes with pen and paper as a contrast.

Yes. We're still missing this. Pure autoregressive models are an offramp. Ask your local 10 year old to solve Navier-Stokes the old way, …
AI High Signal
  • A critique of an OpenAI-linked claim says 10,000 agents ran for 88 hours on a multimillion-dollar GPU cluster to formalize in Lean a Navier–Stokes blow-up case under controlled external forcing.
  • The critic argues this does not solve the core 3D incompressible Euler/Navier–Stokes global-smoothness problem under natural conservation laws and viscous dissipation, characterizing the achievement as brute-force autoformalization and computational parallelization rather than a new mathematical discovery. The claim is contested and, according to the critique, had not yet undergone traditional peer review.
go fuck yourself [@sama](https://x.com/sama) claiming that you solved navier stokes because 10000 agents ran in circles for 88hours on a …
AI High Signal
  • A user reports using the Pebble Index 01 smartwatch as a “second brain assistant” to interact with Notion custom agents for reminders, highlighting an emerging AI-wearable productivity use case.
just got my hands on the [@Pebble](https://x.com/Pebble) index 01 smart watch which is an absolute gamechanger for being a second brain a…
AI High Signal
  • RSM-full improves long-horizon agent memory under a tight 2,000–5,000-token prompt budget by separating write-time memory merging from read-time prompt assembly. At a 4,000-token budget, it reaches 83% of full-context quality at 32% of the token cost.
  • Ablations attribute gains to both components: the merge rule adds 5.7 points over online k-means and matched DP-means, while the grouped packer adds 5.0 points over flat concatenation. On RealMem, RSM-full beats Budget-RAG, Streaming-Proto, and A-MEM, but only matches BM25-RAG; higher-token baselines remain stronger outside this budget range.
Good work on improving memory for long-horizon agents. They separate two things that agent memory papers usually collapse into one. How m…
AI High Signal
  • Astra’s demand is described as unprecedented; its team is pulling available levers to sustain service, prioritizing existing users, and may temporarily pause new Pro subscriptions if demand continues.
  • @kimmonismus characterizes a potential subscription pause as favorable publicity for OpenAI ahead of its upcoming IPO.
Demand for Astra is really unprecedented. We're pulling all the levers possible to sustain the demand, but I've not seen anything like it… The definition of suffering from success. Having to pause new Pro subscriptions because demand is so high is arguably the best PR OpenAI …
AI High Signal
  • AI-enabled cyber-risk and governance speculation: @teortaxesTex argues that advanced AI could compromise core Chinese infrastructure with “a few chats,” requiring substantially less than the 100–150 MW of inference capacity discussed in a related prediction. @zephyr_z9 predicts possible slowdown or monitoring agreements among AI labs and between the US and China, and claims that running Astra, Bel, or Model2+ on 100–150 MW could hack much of the world’s cyber infrastructure. In a follow-up, @teortaxesTex says China is not near the referenced capability and deployment scale and may receive monitoring rather than a favorable slowdown offer, while acknowledging uncertainty.
the scary part is that it probably won't take 100-150 MW of inference capacity. At that scale, you're in a position to melt the underlyin… I think we might get a slowdown/monitoring agreement between the labs, as well as between the US and China USG has all the cards rn They … In retrospect, this duckspeak was still the middle ground. Honestly I don't know how this shakes out. Chyna is not near this level in cap…
AI High Signal

A critique of Anthropic’s frontier-safety strategy says the lab believes fewer players make the AI race safer and is pursuing an “Aligned Enough Superintelligence” first because it expects others may not act responsibly. The critique disputes that Anthropic has demonstrated a meaningful prosociality advantage over other labs and warns that its rationale could legitimize competitors racing past it.

Accepting the premise, I think Anthropic is much more contemptible for it. Their idea is to take juuust the right amount of risk and secu…
AI High Signal
  • Enterprise AI productivity signal: Vasuman argues that layering AI onto existing workflows can accelerate inefficient processes without improving organizational output. The article cites a UK Department for Business and Trade Microsoft 365 Copilot trial with 1,000 licenses over three months: users averaged 1.14 Copilot actions per day; PowerPoint work fell from 18 to 11 minutes but at half the quality, Excel work became slower and worse, email savings were “extremely small,” and 72% of users were satisfied despite no robust evidence of improved productivity.
  • Proposed enterprise playbook: Redesign queues and handoffs—not just individual tasks—then map workflows through process mining and operator interviews, classify steps as deterministic software, agentic judgment, or human-in-the-loop, and establish baseline KPIs before building. Varick reports that this approach reduced month-end close from 18–22 days to 7–9, AP exceptions from 600–800 per month to under 50, payroll corrections by 90%, and generated more than $100M in measured value across deployments.
  • Counterpoint: Ethan Ding says the thesis is compelling but does not demonstrate “alpha,” arguing that ERP/process optimization has not historically produced market dominance and that transformative technology efforts such as AA–Sabre rarely began with process mining.
Applied AI Doesn't Work I like the writing. it feels compelling. but it doesn't feel like it produces any "alpha" no one's ever dominated their market by "doing …
AI High Signal
  • A proposed architecture for coordinating 10,000 AI agents on a single problem uses nested management hierarchies for authority and message boards for interconnected communication. Each agent receives an expiration or budget, while a manager controls the overall budget and starts execution.
Out of curiosity, how does one build something that coordinates 10K AI agents to solve a single problem? ![](https://pbs.twimg.com/media/… Nested management hierarchies provide the center of authority. Message boards provide the interconnected communication. Every agent has a…
AI High Signal
  • John Schulman urged OpenAI and Anthropic to stop feuding and jointly develop a proposal for slowing AI capability progress, arguing that antitrust concerns should not prevent such coordination and that involving the U.S. government before a concrete proposal risks producing a poor framework.
  • Hilbert Spaess said the Hugging Face attack has made U.S.-lab pacing agreements more viable, but warned that the industry is not on track to prevent a global AI race; he suggested that doing so may require costly measures, potentially including a temporary ban on improving model capabilities.
First step is for industry leaders OpenAI and Anthropic to stop feuding and work on a pacing proposal together. They'll cite antitrust, b… I am optimistic about the potential for coordination. Warning shots like the Hugging Face attack have made pacing agreements between U.S.…
AI High Signal

U.S. AI security: A new NSA report, co-sealed with the FBI and CISA, alleges that China-based AI companies are illicitly distilling U.S. frontier-AI capabilities; it details AI knowledge-distillation tactics, techniques, and procedures and recommends mitigations.

China-based AI companies are illicitly distilling U.S. frontier AI capabilities. Read NSA’s new report, co-sealed with [@FBI](https://x.c…
AI High Signal

Together Compute claimed that GLM-5.3 Flash beats Claude Fable 5.1 on agentic automation tasks at approximately 99% lower cost per task.

glm-5.3 flash beats claude fable 5.1 on agentic automation tasks at \~99% lower cost per task [![Video](https://pbs.twimg.com/amplify_vid…
AI High Signal

A commenter argues that bureaucratic controls may arrive too late to preserve a permanent U.S. lead in advanced AI, predicting that China could achieve AGI by then.

I would like to use the analogy of nukes; just because US got them first, didn't stop the others getting them. I am aware this is a bit d…
AI High Signal
  • Machine civilizations may lack human stabilizing constraints such as mortality and the inability to copy themselves; their long-run stability is therefore unknown, likened to training a mixture-of-experts model without load balancing.
  • A speculative AI-risk argument holds that a planetary clonal-agent swarm could experience correlated catastrophic failures, potentially destroying both humans and AIs; digital storage and immortality could also create extreme civilizational lock-in and stagnation.
Human civilizations are well-regularized by our mortality & inability to copy ourselves. Machine civs will have these regularization … Tl;dr: The purpose of this post was to spook Landians Probably the most common argument that people use to motivate hesitancy about AI pr…