We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
OpenAI releases a large body of math from an unreleased model
OpenAI published "a broad range of new mathematical results produced by an internal frontier model" on GitHub. It says it consulted the Institute for Advanced Study's independent Advisory Group on Mathematics and AI on how to release them . One widely shared summary gives these figures: 722 manuscripts in 372 families of related results, drawn from about 4,000 research problems. It says the standard procedure averaged three hours of ChatGPT Pro thinking compute per result. The release includes papers, proof artifacts and selected reasoning summaries. The model itself is not released .
Two results drew the most attention from specialists:
- Riemann zeta. Mathematician Alex Kontorovich says the previous best zero-free region got thinner higher up the imaginary axis, and that the release instead proves a zero-free strip. He says a human who did this would get "an instant Fields Medal" . This is a partial result, not a proof of RH.
- Integer multiplication. A linked preprint in the repo is titled "Integer multiplication below n log n" .
One reader estimates that about 20% of the results are disproofs or counterexamples. He argues this undercuts the view that AI math progress is mostly brute-force search . Others played down the compute figure. @teortaxesTex says three hours of Pro thinking "is not much" and concludes that frontier math is "about as hard as SWE hobby projects" for this model . SWE-bench co-author Ofir Press says coding has not yet had a comparable moment. His test would be agents rewriting ffmpeg or SQLite to run 5× faster on half the memory .
Mistral Large 4: real gains, disputed superlatives
Mistral announced Large 4 ("Le Chonk"). It is a natively multimodal model with 1T total and 49B active parameters, available by API now, with open weights due at the end of October. Mistral calls it the best open-weights model from the US or Europe on aggregated benchmarks and says it beats closed frontier models on visual grounding .
Independent numbers are mixed:
- Artificial Analysis scores the preview at 38 on its Intelligence Index, close to GPT-6 Luna (38) and DeepSeek V4.1 Flash (39). That makes it the strongest model from outside the US and China. It also scores 82% on CyberGym-E2E-AA. But cost per task is $1.13, more than 4× similar open models such as GLM-5.3-Flash at $0.25. A launch discount halves that to $0.57 for two weeks .
- Yuchen Jin points out it trails GLM-5.3, and even GLM-5.3 Flash, on that index, and calls this an "eval crisis" .
- Clément Delangue notes the model "can't be the best open-weight model if you're not open-weight yet" .
- Surge's blind human evaluation, commissioned by Mistral, has professional engineers rank coding quality. It placed Large 4 first among open-weight models and second of five overall, behind Opus 5 .
- Cline says Large 4 beats Opus 5.5 and GPT-6 Astra on cyber benchmarks mostly because it refuses fewer tasks. About 40% of Opus and Astra tasks were blocked by their safety filters .
Coming a day after Reflection's Beam, this prompted Jin to say both Western models reached roughly GLM-5.2 level. He speculates the remaining gap exists because Chinese labs can distill Anthropic and OpenAI models and Western labs can't . In China, Ant Group's Ling 3.1 Flash (560B total, 25B active, 1M context, weights "soon") scored 41 on Artificial Analysis's index, up from 20 for its predecessor . Its price is higher, at $0.99 per task .
Cheap "decision models" become a product category
OpenAI put its Decisions API in public beta for choosing a model, tool or action in near real time. It says decisions are up to 10× faster than calling GPT-6 Luna through the Responses API . Pricing starts at $0.10 per million input tokens, with no output or cache charges . Theo sees value in using it to expose the right tools to an agent, but calls using it to decide what "level of intelligence" a task needs "absolutely useless" .
Competitors moved the same day. Perplexity's open-weights pplx-decider-v1.1-27b tops Hugging Face's new Decision Index 0.3 at $0.02 per million input tokens, half the price of v1 . Vals AI tested TypeSafe's Jev on SEC-filing claim verification. Jev tied GPT-6 Astra at 97.5% accuracy for about $0.025 per 1,000 cases versus $12.29 . But it ranked last on a LegalBench slice .
Google ships on-device embeddings and a cheaper image model
EmbeddingGemma 2 is Google's first natively multimodal open embedding model, released under Apache 2.0. It maps images, video, audio and code into one space . At 740M parameters it uses about 191–567MB of active RAM and has an 8K context window . Nano Banana 2.1 is billed as beating the previous Pro image model at $0.034 per image versus $0.134 .
Can we trust agent oversight and evals?
- METR found a bug in the transcript viewer of Inspect, a popular evaluation framework. It would have let an agent show reviewers a fake or edited transcript . The attack was a MathJax JavaScript injection that could be triggered from the agent's reasoning. METR has not seen it exploited, and Inspect was patched within a day . METR's lesson: treat every AI output as untrusted input to the systems that supervise it .
- CI-aware bench measures whether models notice AI control interventions in their text. Most models were near chance in spring. GPT-6 Astra "nearly saturates it now" .
- AutomationBench Verified audited all 600 public tasks in Zapier's benchmark. Reviewers confirmed and fixed 206 buggy verifiers. Re-grading 1,235 Kimi K3 runs changed 27.9% of grades .
- PZero Research says frontier models' experimental research taste has doubled about every 3 months since December 2025, and that Opus 5.5 now beats its expert baseline. Most of those experts have not worked at a frontier lab .
Anthropic expanded its Cyber Verification Program. Verified security professionals get access to Claude Mythos 5.1, Opus 5.5 and Sonnet 5.5, and new tiers cover authorized offensive work such as penetration testing .
Epoch AI points to a commercial problem for open-weight labs. After Zhipu released GLM 5.3 Flash, 26 other OpenRouter providers were serving it within three weeks. Zhipu's share of the model's tokens fell from 88% to 22% .