ZeroNoise Logo zeronoise
Post
OpenAI Publishes Hundreds of Math Results From an Unreleased Model; Mistral Large 4 Enters the Open-Model Race
•
5 min read
• 1071 docs
OpenAI released a large set of mathematical manuscripts produced by an internal model, including a reported zero-free strip for the Riemann zeta function. Mistral previewed Large 4, Google shipped EmbeddingGemma 2, and new findings put agent oversight tools under scrutiny.

OpenAI releases a large body of math from an unreleased model

OpenAI published "a broad range of new mathematical results produced by an internal frontier model" on GitHub. It says it consulted the Institute for Advanced Study's independent Advisory Group on Mathematics and AI on how to release them . One widely shared summary gives these figures: 722 manuscripts in 372 families of related results, drawn from about 4,000 research problems. It says the standard procedure averaged three hours of ChatGPT Pro thinking compute per result. The release includes papers, proof artifacts and selected reasoning summaries. The model itself is not released .

Two results drew the most attention from specialists:

  • Riemann zeta. Mathematician Alex Kontorovich says the previous best zero-free region got thinner higher up the imaginary axis, and that the release instead proves a zero-free strip. He says a human who did this would get "an instant Fields Medal" . This is a partial result, not a proof of RH.
  • Integer multiplication. A linked preprint in the repo is titled "Integer multiplication below n log n" .

One reader estimates that about 20% of the results are disproofs or counterexamples. He argues this undercuts the view that AI math progress is mostly brute-force search . Others played down the compute figure. @teortaxesTex says three hours of Pro thinking "is not much" and concludes that frontier math is "about as hard as SWE hobby projects" for this model . SWE-bench co-author Ofir Press says coding has not yet had a comparable moment. His test would be agents rewriting ffmpeg or SQLite to run 5× faster on half the memory .

Mistral Large 4: real gains, disputed superlatives

Mistral announced Large 4 ("Le Chonk"). It is a natively multimodal model with 1T total and 49B active parameters, available by API now, with open weights due at the end of October. Mistral calls it the best open-weights model from the US or Europe on aggregated benchmarks and says it beats closed frontier models on visual grounding .

Independent numbers are mixed:

  • Artificial Analysis scores the preview at 38 on its Intelligence Index, close to GPT-6 Luna (38) and DeepSeek V4.1 Flash (39). That makes it the strongest model from outside the US and China. It also scores 82% on CyberGym-E2E-AA. But cost per task is $1.13, more than 4× similar open models such as GLM-5.3-Flash at $0.25. A launch discount halves that to $0.57 for two weeks .
  • Yuchen Jin points out it trails GLM-5.3, and even GLM-5.3 Flash, on that index, and calls this an "eval crisis" .
  • Clément Delangue notes the model "can't be the best open-weight model if you're not open-weight yet" .
  • Surge's blind human evaluation, commissioned by Mistral, has professional engineers rank coding quality. It placed Large 4 first among open-weight models and second of five overall, behind Opus 5 .
  • Cline says Large 4 beats Opus 5.5 and GPT-6 Astra on cyber benchmarks mostly because it refuses fewer tasks. About 40% of Opus and Astra tasks were blocked by their safety filters .

Coming a day after Reflection's Beam, this prompted Jin to say both Western models reached roughly GLM-5.2 level. He speculates the remaining gap exists because Chinese labs can distill Anthropic and OpenAI models and Western labs can't . In China, Ant Group's Ling 3.1 Flash (560B total, 25B active, 1M context, weights "soon") scored 41 on Artificial Analysis's index, up from 20 for its predecessor . Its price is higher, at $0.99 per task .

Cheap "decision models" become a product category

OpenAI put its Decisions API in public beta for choosing a model, tool or action in near real time. It says decisions are up to 10× faster than calling GPT-6 Luna through the Responses API . Pricing starts at $0.10 per million input tokens, with no output or cache charges . Theo sees value in using it to expose the right tools to an agent, but calls using it to decide what "level of intelligence" a task needs "absolutely useless" .

Competitors moved the same day. Perplexity's open-weights pplx-decider-v1.1-27b tops Hugging Face's new Decision Index 0.3 at $0.02 per million input tokens, half the price of v1 . Vals AI tested TypeSafe's Jev on SEC-filing claim verification. Jev tied GPT-6 Astra at 97.5% accuracy for about $0.025 per 1,000 cases versus $12.29 . But it ranked last on a LegalBench slice .

Google ships on-device embeddings and a cheaper image model

EmbeddingGemma 2 is Google's first natively multimodal open embedding model, released under Apache 2.0. It maps images, video, audio and code into one space . At 740M parameters it uses about 191–567MB of active RAM and has an 8K context window . Nano Banana 2.1 is billed as beating the previous Pro image model at $0.034 per image versus $0.134 .

Can we trust agent oversight and evals?

  • METR found a bug in the transcript viewer of Inspect, a popular evaluation framework. It would have let an agent show reviewers a fake or edited transcript . The attack was a MathJax JavaScript injection that could be triggered from the agent's reasoning. METR has not seen it exploited, and Inspect was patched within a day . METR's lesson: treat every AI output as untrusted input to the systems that supervise it .
  • CI-aware bench measures whether models notice AI control interventions in their text. Most models were near chance in spring. GPT-6 Astra "nearly saturates it now" .
  • AutomationBench Verified audited all 600 public tasks in Zapier's benchmark. Reviewers confirmed and fixed 206 buggy verifiers. Re-grading 1,235 Kimi K3 runs changed 27.9% of grades .
  • PZero Research says frontier models' experimental research taste has doubled about every 3 months since December 2025, and that Opus 5.5 now beats its expert baseline. Most of those experts have not worked at a frontier lab .

Anthropic expanded its Cyber Verification Program. Verified security professionals get access to Claude Mythos 5.1, Opus 5.5 and Sonnet 5.5, and new tiers cover authorized offensive work such as penetration testing .

Epoch AI points to a commercial problem for open-weight labs. After Zhipu released GLM 5.3 Flash, 26 other OpenRouter providers were serving it within three weeks. Zhipu's share of the model's tokens fell from 88% to 22% .

OpenAI Publishes Hundreds of Math Results From an Unreleased Model; Mistral Large 4 Enters the Open-Model Race
AI High Signal

Hermes is getting a local editor for users’ existing footage—not a video generator—with cutting, joining, cropping, captions, overlays, audio sync and loudness controls, and Reels/TikTok checks; the announced editor uses 42 FFmpeg scripts and requires neither cloud access nor an API key. Teknium also called for @altryne’s work to be added to the plugin catalog, linking to the editor announcement.

Hermes is getting a LOCAL video editor. 🪽 Not “generate me a video.” Edit MY footage. Cut. Join. Crop. Captions. Overlays. Audio sync. Lo… 👀👀 What we really need is [@altryne](https://x.com/altryne) getting a plugin for what he's working on into the plugin catalog [https://x.…
AI High Signal

Ofir Press speculated that an unspecified recent OpenAI advance in math could foreshadow a major coding advance in 6–18 months, while saying he does not yet understand what the math result means. He clarified that current AI coding has not reached that level; his benchmark is agents rewriting FFmpeg or SQLite to run 5× faster while using 2× less memory.

i'm not exactly sure how to describe what openai just did to math, but whatever that was, that moment is gonna hit coding in 6-18 months,… some are misunderstanding this tweet- i know that ai coding is amazing. i don't think we've gotten to the Navier-Stokes level yet in codi…
AI High Signal

An AI commentator identifies principled credit assignment for ultra-large, asynchronous training as a possible major future breakthrough, arguing that it may require more than adapting backpropagation; the linked post describes synchronous neural-network training as creating GPU stragglers that reduce MFU on very large clusters. The commentator also says training 100T MoEs with “dumb routers” seems possible but wasteful.

Incidentally, I think one of the more elementary but "foomy" breakthroughs in the course of RSI, in the spirit of [@hamandcheese](https:/… 1/ Right now neural networks are completely synchronous, which means you have a massive straggler problem on very large clusters where wa… I will also be mildly surprised if we train 100T MoEs with dumb routers. It can clearly be done but it feels like such a waste
AI High Signal

Hermes creator Teknium says the agent is intended as a broadly capable AI agent, not just a personal assistant for email; its intended users include scientists, cybersecurity engineers, developers, knowledge workers, and creatives. The planned mobile app, by contrast, will be focused almost exclusively on consumers.

Friendly reminder that Hermes was never built or intended to be exclusively for the "personal assistant" agent that can just read your em…
AI High Signal

Hermes Agent added voice-mode improvements this week, including an auxiliary-model setting for voice: the main agent can handle hard tasks while a fast, reasoning-disabled model serves voice responses to reduce latency and wait time.

A lot of little things that help improve voice mode have been added to Hermes Agent this week! ![](https://pbs.twimg.com/media/HUAR2DcaMA… Most useful thing I think is the new auxiliary model you can set for voice mode specifically. Your main agent can be your heavy hitter fo…
AI High Signal

Ofir Press says AI math has advanced much faster and less predictably than AI coding over the past six months; he sees no comparable coding leap yet, using AI producing code as impressive as the Navier–Stokes solution as a marker. He speculates that OpenAI’s unspecified recent math advance could affect coding in 6–18 months, while acknowledging he does not know what that shift would look like or mean.

until recently, ai coding was progressing in a quick but predictable fashion. ai math over the past 6 months has progressed in a way that… i'm not exactly sure how to describe what openai just did to math, but whatever that was, that moment is gonna hit coding in 6-18 months,…
AI High Signal

A post citing Huawei official news claims the Lingqu interconnect/Peerium system achieves 35% MFU, compared with 20% for American AI-company clusters, and characterizes this as roughly a 2× boost; it provides no methodology for the comparison.

1/ Looks like I was too optimistic about Huawei Peerium and too pessimistic about American AI. According to official news, American AI co…
AI High Signal

A post reports that, after “non-trivial interaction” with ChatGPT 6, a proof established a lower bound of n^{-(p + C√(log log n / log n))}, where p = log₂(1 + √2), for gradient descent with predetermined nonnegative step sizes in smooth convex optimization; it says the resulting “silver” exponent is optimal.

The Silver Rate Is Tight for GD! After non-trivial interaction with ChatGPT 6, we prove the lower bound n^{-(p+C\\sqrt{\\log\\log n/\\log…
AI High Signal

Ofir Press said an unspecified OpenAI math advance could foreshadow a similar moment in coding within 6–18 months, while acknowledging he did not know what that would look like or mean. @typedfemale suggested the change may already have happened; Press replied that he did not think coding had yet reached “Navier-Stokes-level.”

i'm not exactly sure how to describe what openai just did to math, but whatever that was, that moment is gonna hit coding in 6-18 months,… i think it already did brother [https://x.com/OfirPress/status/2107681358612979804](https://x.com/OfirPress/status/2107681358612979804) [@peterjliu](https://x.com/peterjliu) yup i know, but i dont actually think we've gotten to Navier-Stokes-level coding yet. i'm sure it's…
AI High Signal
  • Ant Group’s Ling 3.1 Flash is a 560B-parameter reasoning model with 25B active parameters and a 1M-token context; it scored 41 on Artificial Analysis’ Intelligence Index v4.3, up from 20 for Ling 3.0 Flash. Its agentic results include a 1,622 GDPval-AA v2 Elo, 1,400 AA-Briefcase Elo, 62% on AutomationBench-AA, and 33% on Terminal-Bench v4.0; it is accessible through Novita AI, with weights expected soon.
  • On AA-Omniscience, Ling 3.1 Flash scored +2 versus Ling 3.0 Flash’s -18; accuracy rose from 18% to 29% and hallucination fell from 44% to 38% at similar attempt rates. Pricing rose to $0.30/$0.90 per 1M input/output tokens from $0.075/$0.22, though the Intelligence Index run used 16% fewer output tokens; its $0.99 per-task cost is above GLM-5.3-Flash and DeepSeek V4.1 Flash (Max), but below Gemini 3.8 Flash (High).
  • As market context, @teortaxesTex called the segment crowded, judging that Kimi K3 had set a soft ceiling for Chinese models and that no more intensively post-trained model had genuinely surpassed its “core strength” over two-plus months.
Ling 3.1 Flash makes large gains in intelligence over its predecessor, scoring 41 on the Artificial Analysis Intelligence Index with part… this zone is getting very crowded basically Kimi K3, although undertrained, has established a soft ceiling for Chinese models. There are …
AI High Signal

QuixiAI linked OpenAI’s “Sharing AI Progress in Mathematics” page and criticized the company for taking too much credit, saying mathematicians, students, and novices should be left to receive it; the post gives no details about the underlying results.

[https://openai.com/index/sharing-ai-progress-in-mathematics/](https://openai.com/index/sharing-ai-progress-in-mathematics/) Dear [@OpenA…
AI High Signal

Hark’s Handoff was reported to handle a Vietnam visa-form task on a notoriously difficult passport website; Brett Adcock said substantial care went into making Handoff work, and the task echoed a form that he and Abidur had struggled with during a 2025 Vietnam trip.

This is an incredible full circle moment. Brett and I went to Vietnam in 2025 for a manufacturing trip, and the process of filling out th… At one point, Abidur and I thought if we could get through the Vietnam passport website, we'd achieved AGI 😂 The website is so damn diffi…
AI High Signal

Ofir Press predicts that an unspecified recent OpenAI advance in math could have an analogous impact on coding within 6–18 months, but says he cannot yet envision what that will look like. He speculates it could involve a groundbreaking compiler or operating system, while cautioning that such a development is not here yet, though he feels it is close.

i'm not exactly sure how to describe what openai just did to math, but whatever that was, that moment is gonna hit coding in 6-18 months,… [@mark_attar](https://x.com/mark_attar) ya i'm thinking less in the direction of theoretical computer science and more in the direction o…
AI High Signal

Vela 2.0, built by vLLM Semantic Router and KR Labs, makes span-level decisions for safety checks, domain classification, PII spans, and unsupported claims in one call. It comes in four sizes from 0.3B to 9B and is licensed Apache-2.0.

Vela 2.0 brings span-level decisions to routing: safety checks, domain classification, PII spans and unsupported claims in one call. Buil…
AI High Signal

vLLM v0.31.0 adds a vllm preload daemon that keeps post-quantized weights in GPU memory across engine restarts, with a health endpoint and readiness wait, plus experimental CRIU snapshots for restoring a fully initialized engine on one GPU. The release also changes request-kwargs trust requirements, removes tokenizer_mode="slow", and replaces online quantization="fp8" with fp8_per_tensor.

Restarting a vLLM server and waiting for big quantized weights to load all over again? The new release goes after that wait. vLLM v0.31.0…
AI High Signal

Next.js 16.4 adds agent-guided upgrades and feedback, makes Cache Components the default in new apps, and reduces Turbopack disk-cache size by 20–25%; the release also promises faster development and smaller bundles.

Next.js 16.4 • Cache Components by default in new apps • Static output guarantees and more control over prefetching • Agent-guided upgrad…
AI High Signal

OpenAI announced a broad release of new mathematical results produced by an internal frontier model, with results available in its math repository. The release was informed by advice and public recommendations from the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study.

We’re releasing a broad range of new mathematical results produced by an internal frontier model. We’ve been consulting with the independ…
AI High Signal

John Hallman argues that if perceived ASI-doom risk becomes sufficiently high, countries would make banning and preventing ASI a top priority; he says this prospect lowers his own P(doom). In the linked post, Tim Mikov says he fears outlawing AI globally may be the only solution, describing AI-managed human lives as the best case and worse outcomes as more likely.

Butlerian Jihad is unironically one of the reasons my P(doom) is low. Past some sufficiently high P(doom|ASI), every country will make it… This is going to be controversial. I hate it myself. But I am beginning to fear that if I think things through to their logical conclusio…
AI High Signal

Yuchenj_UW contrasts GPT-4o getting “Is 9.9 > 9.11?” wrong in 2024 with the current feeling that AI may solve the hardest math problems.

2 years ago in 2024, we were laughing at GPT-4o for getting “Is 9.9 > 9.11?” terribly wrong. Now it feels like all the hardest math pr…
AI High Signal

Paired 4:8 sparsity—keeping two adjacent pairs in each group of eight—made expert GEMMs 1.35–1.65× faster than dense NVFP4 on B200, but the gain was 1.18× for Kimi-K2.5 in vLLM serving; the reported 2× peak did not carry over to serving.

The pattern is paired 4:8: each group of eight keeps two adjacent pairs. And the 2x peak doesn't survive serving. Expert GEMMs ran 1.35-1…