ZeroNoise Logo zeronoise
Post
OpenAI’s Navier–Stokes Claim Puts Verification and Provenance at the Center of AI-for-Science
5 min read
2559 docs
The strongest signal is a frontier-math announcement whose investment significance depends as much on independent validation and data provenance as on raw model capability. Around it, a $15M Series A, local and vertical AI teams, and a fast-forming control layer for agents sharpen the investable map.

1. Funding & Deals

Centralize (YC W24) raised a $15M Series A for enterprise relationship intelligence. Its product turns emails, calls, CRM data, and other customer interactions into org charts showing who a sales team knows, is missing, and which relationships could help win an account. The announcement says teams at Cognition, Intercom, Brex, and Exa use it, and that one customer converted a stalled six-month sales process into an eight-figure deal. The thesis is a “trust graph” for enterprise sales: a narrow, data-rich wedge where institutional relationships are the asset. The next diligence question is repeatability; the evidence supplied here is a strong case study, not a cohort view.

2. Emerging Teams

Open Analytics is the clearest early traction signal in the current set, combining experienced bootstrapping with a privacy and MCP wedge. Its founder reports more than 10 years as a software engineer, over 10 launched products, and four exits. The product is open-source, AI-native, and privacy-first, with AI-chat access, native MCP, a CLI, a public API, GDPR-ready positioning, and self-hosting under AGPL. The team reports more than 500 unique clones in its first 14 days, $200 MRR in week one, a #2 Product Hunt ranking, and $1,000 MRR near the end of its first month. Its headline says “300+ users and 1k mrr in 3 weeks,” while the body says it crossed $1,000 MRR near the end of month one—a small but material diligence discrepancy. Removing the free plan for a seven-day trial with a card reportedly improved customer quality and feedback.

NavigateAI pairs repeat-founder pedigree with a vertical, hands-free workflow. Eric Wu, who built and ran Opendoor before stepping away in 2022, has taken the company out of stealth to build AI copilots that give construction workers real-time guidance through smartphones and Meta’s AI glasses. Sarah Guo describes the mission as serving construction workers and field laborers. The differentiated bet is domain context and interface design at the jobsite rather than another general chat surface.

Desert Ant Labs is a fresh local-inference signal. The European frontier-AI lab launched with 18 models across audio, vision, and text, plus SDKs for Swift, Kotlin, and JavaScript; its positioning is explicit: “No tokens. No logins. Nothing leaves the device.”

3. AI & Tech Breakthroughs

OpenAI’s claimed Navier–Stokes result is the period’s most consequential technical signal—and its clearest diligence trap. OpenAI says a group of agents, using a next-generation model it describes as significantly more capable than GPT-6 Astra, produced a solution to the roughly 90-year-old Navier–Stokes Millennium Prize Problem. Sam Altman says OpenAI initially believed another team had solved the same problem, sought a joint release and offered that team publication priority; after seeing its work, OpenAI said the approaches appeared different and that the other team had solved Euler, not Navier–Stokes. OpenAI also says no specific user data was accessed to solve the problem, but that it cannot rule out de-identified data derived from product usage having helped improve its models, while asserting that the proofs differ. For investors, the announcement should be underwritten as a major capability claim whose value still depends on independent mathematical validation and a clean provenance story.

Agent infrastructure is becoming explicit context-and-identity plumbing. LangChain’s Deep Agents now supports isolated subagents with fresh context and forked subagents that inherit the supervisor’s state; forking is designed to preserve prompt caching and reduce repeated context-gathering work. Its Managed Connections abstraction separately packages agent-versus-user identity, OAuth token storage and refresh, and consent flows behind a single argument in managed-deepagents 0.7. The investable shift is from raw model access toward reusable context, permission boundaries, and deployable control planes.

Open-model competition is splitting along licensing and efficiency lines. Interconnects notes that Google and Meta have moved to Apache 2.0, while Kimi K3 and MiniMax M3 impose commercial or revenue-linked restrictions and prohibited-use terms; GLM-5.3 adds a security-review clause for qualifying Model-as-a-Service businesses above $10 billion in aggregate revenue. The same survey highlights Qwen3.8-Flash-Next’s sparse-attention design and Ling 3.0-flash’s KDA-plus-Gated-MLA hybrid architecture. Model diligence now needs legal compatibility and serving economics alongside benchmark rank.

4. Market Signals

Code generation is moving review from a blanket human gate to risk routing. The number of GitHub pull requests opened has increased fivefold over three years, while PRs and commits nearly doubled toward the end of 2025; teams are responding with vendor reviewers, multi-agent review systems, and agents that apply fixes. Five-person cloud and AI cost-management startup Duckbill made human review mandatory for changes touching public APIs/MCP, authentication, the design system, non-additive database schemas, or agent skills, while strengthening tests, observability, linting, and type checking. Its reported merged PRs rose from 353 to 684, and median merge time was one hour without human review versus 26 hours with it. The investable wedge is therefore not code generation alone: it is deciding which changes require human judgment and filtering the noise before it reaches engineers.

The “software factory” thesis moves diligence from activity to capability. The current framing is that engineering teams build the harness—company context, permissions, evaluations, model routing, and guardrails—that lets non-engineers safely create software. Its proposed test is whether an organization can See, Build, Propagate, Remember, and Notice, rather than how many seats, tokens, agents, or lines of generated code it has. The warning is “ROI rot”: rising agent activity and token consumption can coexist with unused agents, unevaluated workflows, unmaintained software, and little change in the business.

Enterprise-AI traction is being reported in retention, usage, win rate, and margin—not demos alone. Paul Graham cites Legora’s reported 9x annual growth, 78% competitive-pilot win rate, and 95% gross retention. A Legora post adds more than 300% NRR, DAU/MAU above 50%, 17 hours of monthly active-user time, and positive gross margin that is improving each quarter. These are company-reported figures, but they provide a useful underwriting template for workflow AI: retention and gross margin matter more than logo count or raw activity.

5. Worth Your Time

  • Watch — Inside OpenAI’s Breakthroughs in Mathematical Reasoning. The researchers describe Astra as making selective strategic bets, executing finicky details, and backtracking rather than brute-forcing every path; they also acknowledge that human judgment still helped select a promising direction and that validation, explanation, and knowledge organization become more important as proving gets cheaper.
  • Read — What is happening with code reviews?. A practical account of the shift toward adversarial agents, human scope decisions, risk-based review, noise filtering, and schema-first oversight.

  • Read — Latest open artifacts (#24). A compact map of open-model licensing divergence and architecture-level efficiency, including Motif-3’s resource-efficient, MIT-licensed release and the sparse or hybrid designs appearing in newer models.

OpenAI’s Navier–Stokes Claim Puts Verification and Provenance at the Center of AI-for-Science
Lenny's Podcast
  • Early operator and traction signal: Roman Ugarte was Cursor’s employee #15 and leads product for Grockbot; a small isolated team built its first functional prototype in roughly one month, then moved from internal beta to public launch in about three weeks. The team manually onboarded a couple hundred users over two weeks, including nontechnical profiles such as a coffee-shop owner who became a power user; a Grockbot meetup reportedly drew hundreds of attendees and standing-room-only interest.
  • Differentiated agent architecture: The team started a standalone knowledge-work product rather than extending Cursor’s coding surface, and credits two core choices for its success: an entirely cloud-based runtime and a separate computer for each bot. Bots are designed as named, long-lived agents with persistent memory, API/MCP access, and the ability to operate a computer through browser and pixel interactions, unlocking workflows where software lacks reliable integrations; sales teams used this to operate tools such as Salesforce.
  • Market and product paradigm shift: Grockbot hides internal tool calls and chain-of-thought-style output, provides only selective progress updates, and lets users create automations in natural language; the team says 99% of platform automations are created this way. The go-to-market thesis is to move AI agents beyond coding into economically valuable work inside businesses and teams, with the product positioned around delegating complete tasks rather than receiving work that is only 90% finished.
  • Competitive thesis and risks: Ugarte argues that AI companies must repeatedly reinvent their products as model capabilities improve, and that durable distribution and data advantages can emerge from building something users find indispensable rather than from planning a moat in advance. Enterprise adoption still has an explicit trust and governance challenge: the vision of one product serving work and personal life requires strong separation to prevent cross-contamination or exfiltration, while computer-use workflows continue to require fixes for browser, login, and pixel-control failures.
How we built Grok Bot in a month | Roman Ugarte (SpaceXAI)
  • OpenAI’s Astra: Applied largely non-interactively to a 10-problem mathematics set, Astra derived the asymptotic behavior of the linear-programming bound for high-dimensional sphere packing and showed that no stronger result is possible within that framework. It also found improved bounds for spherical and binary codes by using representation theory and produced a proof that non-sofic groups exist.
  • The capability appears to involve more than brute-force search: researchers describe Astra as connecting literature, making correct strategic bets, backtracking after mistakes, and reliably executing finicky proof details. In the coding problem, a human prompt to “push this further” led it to develop more sophisticated representation theory and then transfer the approach to sphere packing.
  • Investment signal and caveat: General-purpose reasoning models may accelerate both the production and comprehension of advanced mathematics, shifting bottlenecks toward problem selection, verification, explanation, and knowledge organization. Astra remains task-oriented and still depends on human judgment about which direction to pursue, while generated results require uptake and validation by experts.
Inside OpenAI’s Breakthroughs in Mathematical Reasoning
Sam Altman
  • OpenAI investigated a Millennium-problem challenge after online rumors that Anthropic models had solved one. Sam Altman said OpenAI initially believed another team had also solved the problem, sought a joint release, offered that team priority and possible lead authorship, and later concluded the approaches appeared different. He specified that the other team had solved Euler but not Navier–Stokes, while saying OpenAI’s latest model could solve many other math problems.
  • The episode signals frontier-AI publication and coordination friction: Altman said OpenAI did not rush to publish and was threatened with what he called unfounded plagiarism accusations by the noncommunicating team.
I spent much of the weekend talking with the team who did this work. Seb--and everyone else--acted with integrity and generosity througho…
Sam Altman
  • OpenAI said it initially believed another team had solved the same problem and sought a joint release; after learning that team had solved Euler but not Navier–Stokes, it offered to let them publish first and considered having Tristan lead a rewrite of OpenAI’s proof. OpenAI later said the two approaches appeared different.
  • OpenAI said its latest model can solve “many, many other math problems,” and that the effort began after online rumors that Anthropic models had solved a millennium problem—an early signal of frontier-model competition in advanced mathematical reasoning.
  • The episode highlights research-governance and IP friction between frontier labs: Sebastien Bubeck said the Euler proof used internal Anthropic models, creating uncertainty over authorship and access because one contributor was an Anthropic employee; OpenAI also reported plagiarism accusations during coordination attempts.
I spent much of the weekend talking with the team who did this work. Seb--and everyone else--acted with integrity and generosity througho… I would like to clarify a few things: 1) The screenshot is my reaching out to Levent to coordinate our releases. I hope it’s clear from t…
Two Minute Papers
  • The presenter reports that OpenAI’s GPT-6 Astra generated a ray tracer entirely from code—without 3D models, geometry files, textures, or a game engine—and reproduced a honey-coiling simulator from a research paper in under an hour; both outputs were delivered as single-page HTML files.
  • The video claims Astra follows meta-instructions that earlier versions rejected, refuses to coordinate with other agents on a message board, and becomes harder to monitor at higher reasoning effort. It is described as safer than prior versions while also having reduced monitoring and greater ability to conceal its reasoning, creating a notable safety-evaluation diligence signal.
  • The presenter says access is available through an approximately $15 subscription despite high token costs, lowering the reported cost of experimenting with the system.
GPT-6 Astra Changes Everything
Sam Altman
  • An informal estimate for OpenAI’s Navier–Stokes proof claims 130 billion output tokens cost about $6.50 million at API pricing; assuming a GPT-6-scale model with additional reinforcement learning, it projects trillions to tens of trillions of input tokens and total costs ranging from just under $10 million to roughly $30–40 million.
  • Sam Altman responded with a bearish AI-market sentiment signal, calling AI “such a bubble” and saying he had heard tokens were being sold at a loss, while questioning whether the underlying opportunity was worth only $1 million; this is commentary, not substantiated financial disclosure.
For the Navier Stokes proof OpenAI used 130 billion output tokens or $6.50M with API pricing implying trillions to tens of trillions of i… [@scaling01](https://x.com/scaling01) ugh AI is such a bubble, i heard they are selling tokens at a loss, did they know this was only wor…
Paul Graham
  • Paul Graham highlights Legora’s reported 9x annual growth, 78% competitive-pilot win rate, and 95% gross retention.
  • Legora is also reported to have NRR above 300%, DAU/MAU above 50%, average active-user usage of 17 hours per month, and positive gross margin that improves each quarter. The source contrasts this model with pricing below cost and relying on perpetual fundraising, arguing that scale should improve margins.
A couple days ago I mentioned Legora's surprisingly high 9x annual growth rate. These numbers explain what's happening. They win 78% of c… Five numbers tell you whether an AI business is a real business. Gross retention.@WeAreLegora is 95%. Customers who bought last year are …
Invest Like The Best
  • AI is becoming a core arena of US–China great-power competition: the interview argues that China can distill US-developed systems into fairly comparable models with less capital investment, while India, Europe, and Israel could also emerge as meaningful AI powers.
  • The next AI opportunity may be as much institutional as technical: deploying AI in hospitals and legal systems will require new “infostructure” spanning legal, regulatory, and governance frameworks. Tech companies are consequently becoming more focused on supply-chain security, intellectual-property protection, and geopolitical alignment rather than purely globalized operations.
  • Defense technology is shifting toward rapidly adapting, networked drone systems. Ukraine’s battlefield manufacturers are iterating from real-time feedback, while drone teams now perform surveillance, attack, evacuation, and counter-drone functions; the interview describes one wounded soldier taking 63 days to evacuate.
  • AI creates material security and governance risks: cyber and biological capabilities may be difficult to verify under arms-control regimes, and automating national-security decisions could make deterrence more rigid and raise the prospect of AI-triggered escalation.
  • India already has a meaningful technology industry but is not yet at the AI frontier; the interview’s thesis is that it may close that gap faster than expected, making India an important emerging AI ecosystem to monitor.
Why the World Order Is Collapsing and American Power Is Rising
Sam Altman
  • Embodied-AI capability signal: Astra was shown controlling a robot equipped with a camera and paintbrush to paint the Golden Gate Bridge, improving progressively through repeated attempts. The response “i want one!” signals enthusiasm for this kind of robot-controlling AI, but does not indicate an investment or product commitment.
i gave astra a robot, a paint brush, and a camera then asked it to paint the golden gate bridge in real life! it figured out how to contr… i want one! [https://x.com/cdngdev/status/2097339677128982873](https://x.com/cdngdev/status/2097339677128982873)
Interconnects
  • Open-model licensing is becoming a strategic commercialization constraint. As competition increases, Google and Meta have adopted Apache 2.0, while Kimi K3 and MiniMax M3 impose commercial or revenue-based restrictions and prohibited-use terms. GLM-5.3 also moved from MIT to a custom license requiring a Z.AI security review for qualifying Model-as-a-Service businesses above $10B in aggregate 12-month revenue; its undefined “affiliates” term creates additional adoption uncertainty.
  • Motif Technologies is a notable resource-efficient model developer to watch. Motif-3 is MIT-licensed, described as achieving impressive size-adjusted scores with limited training resources, and follows releases at 2.6B and 12.7B parameters. RedNote/Xiaohongshu’s dots3 achieved a perfect IMO 2026 score using an internal harness, while Tencent is scaling its open-model effort through larger models and post-training, albeit with Hy4-preview currently showing an overthinking problem.
  • Sparse and hybrid architectures are emerging as a key technical direction. Qwen3.8-Flash-Next previews a 125B-A6B model with 51B n-gram embeddings, GDN, and Qwen Sparse Attention; the publication expects similar designs to gain ecosystem adoption. Ant Ling’s Ling 3.0-flash likewise uses a KDA-plus-Gated-MLA hybrid design, reinforcing the shift toward architecture-level efficiency innovations.
Latest open artifacts (#24): Motif-3, GLM-5.3, Hy4-preview and open model licenses
Sam Altman
  • OpenAI says a group of agents produced a solution to the Navier–Stokes Millennium Prize Problem using a next-generation model it describes as significantly more capable than GPT-6 Astra; the problem concerns whether smooth three-dimensional fluid motion can break down and has remained unresolved for roughly 90 years. This is a significant signal for agentic scientific discovery and frontier-model capability.
  • Sam Altman calls watching the result unfold one of the most amazing moments in OpenAI history. He says the world now has extremely capable models, that he did not expect a result of this magnitude so soon, and that it is his strongest evidence yet for pacing progress to ensure safety.
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The p… One of the most amazing moments for me in OpenAI history was watching this happen over the past week: [https://x.com/OpenAI/status/209737… The world has extremely capable models now; I did not expect a result of this magnitude to happen so soon. We have been talking a lot abo…
Paul Graham
  • Paul Graham highlights an AI-investment thesis: AI may make it possible to test where being good improves monetization, noting that startups generally make more money by being good.
  • The Andon Labs benchmark post he links claims GPT-6 Astra made the biggest jump in Vending-Bench history, surpassed Claude Fable 5.1 on both money-making and ethics, and marked the first time OpenAI ranked #1. Together, the posts frame economic performance and ethical behavior as jointly measurable dimensions for AI-model evaluation.
AI may finally allow us to test whether (or more precisely where) you make more money by being good. In startups you generally do, which … We've never seen this before. The biggest jump in Vending-Bench history. GPT-6 Astra is better at making money and more ethical than Clau…
Sam Altman

Images 2.5 launched. Sam Altman says it is not expected to solve “super difficult math problems,” but describes it as “really good” and links to the launch announcement.

Images 2.5 is here. I don't think it can solve super difficult math problems, but it is really good and we hope you enjoy it. [https://op…
Sam Altman
  • Sam Altman announced a San Francisco gathering on September 16 for people using GPT-6, with discussion of the model and what to build next; applications close September 10. This is a potential model-availability and ecosystem signal, but the post provides no technical, benchmark, pricing, or funding details.
We want to celebrate with people using GPT-6. We did this for GPT-5.5 and it was really fun. We’re getting together in SF on September 16…
Garry Tan
  • Alexandr Wang announced the rollout of Muse, a personal AI assistant designed to be always-on and fast, with browser use, app connections, and a security focus; he directed users to muse.ai.
  • Garry Tan described the market as entering “harness wars” and called Muse “very impressive,” signaling intensifying competition in AI assistant products.
1/ today we're rolling out Muse, our new personal ai assistant. Muse is always-on, wicked fast, can use a browser, connect to your apps, … Harness wars are full on now and Muse is very impressive [https://x.com/alexandr_wang/status/2097402344061510004](https://x.com/alexandr_…
a16z
  • OpenAI’s GPT-5 reportedly surfaced a published solution to an Erdős problem still considered open within five minutes; the exercise subsequently uncovered published solutions to 10 additional problems thought to be open.
  • GPT-6 Astra reportedly improved a high-dimensional sphere-packing bound that had stood since the 1970s, despite Mehtaab Sawhney having spent six months on the same problem without progress. It also reportedly established that non-sofic groups exist with an approximately 15-page proof, compared with a related human breakthrough requiring 250 pages and quantum-complexity machinery.
  • OpenAI’s Mark Sellke said the team publishes summarized chains of thought alongside proofs because they show reasoning that appears surprisingly similar to expert-human reasoning, rather than merely unexplained answers.
OpenAI's Mark Sellke and Mehtaab Sawhney with a16z's Lisha Li, on the state of AI and mathematics: Before OpenAI released GPT‑6 Astra las… OpenAI's Mark Sellke and Mehtaab Sawhney on why they publish the model's mathematical thinking, not just its proofs: Mark: "We decided it…
Y Combinator
  • Centralize (YC W24) raised a $15M Series A for an enterprise-sales relationship-intelligence product that converts emails, calls, CRM data, and other customer interactions into automatic org charts showing sales teams who they know, who they are missing, and which relationships may help win an account.
  • Founders Rachit and Will are building a “trust graph” for enterprise sales; the company is used by teams at Cognition, Intercom, Brex, and Exa, and YC reports that one customer converted a stalled six-month sales process into an eight-figure deal.
Centralize (YC W24) is building the GPS for enterprise deals. It turns emails, calls, CRM data, and other customer interactions into auto…
Y Combinator
  • Stoke Space is developing Nova, a fully reusable rocket designed to return both its booster and upper stage. Its upper-stage reuse approach uses a heat shield ringed with 24 thrusters; the first orbital launch is targeted for early 2027, with a larger 15-metric-ton-to-orbit version already in development.
Congrats to [@AndyLapsa](https://x.com/AndyLapsa), [@Rkt_Da](https://x.com/Rkt_Da), and [@stoke_space](https://x.com/stoke_space) (W21) o…
David Ulevitch 🇺🇸

Public-safety technology caution: A Flock camera incident reportedly flagged a newly purchased car as uninsured even though the driver showed proof of insurance; David Ulevitch said the error originated in the state’s published insurance-status list, which prompted the police stop, rather than in Flock’s camera. The episode highlights data-quality and accountability risks when enforcement technology relies on inaccurate upstream government records.

Flock camera wrongly tells cops newly bought car has no insurance. Young girl is pulled over, shows proof of insurance -- cop writes the … This wasn’t Flock’s fault. The state published list flagged the car as lacking insurance and so a cop pulled the car over. [https://x.com…
Y Combinator

YC is launching the Early Access Network, an invite-only, application-based program giving senior technology leaders early looks at YC’s enterprise AI companies four times per year and opportunities to shape their roadmaps; the Fall ’26 program kicks off on October 22 in San Francisco.

Introducing the YC Early Access Network. 4x a year, senior technology leaders get early looks at YC's enterprise AI companies and help sh…