We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
OpenAI is presenting Astra as a low-cost research system
OpenAI says an internal version of Astra, its next major model, produced new results on ten problems that had seen no progress on their main result for at least a decade, and in most cases much longer. The problems span high-dimensional geometry, coding theory, circuit complexity, group theory, operator algebras, quantum complexity, lattice cryptography and extremal combinatorics.
The examples are unusually broad: OpenAI lists a construction establishing non-sofic groups, a disproof of Connes’s rigidity conjecture, stronger sphere-packing and coding bounds, an exponential theorem for quantum games, a lattice-cryptography result, and new Ramsey and extremal-graph results. It estimates that finding all ten solutions required roughly $2,000 worth of tokens at Sol API rates; humans prepared the manuscripts, after which the model formalized each argument in a Lean certificate.
The important shift is the proposed research workflow, not just the number ten: generate candidate mathematics with a model, formalize it, and release enough of the process for others to inspect. If the results survive scrutiny, that makes AI-assisted proof discovery a potentially inexpensive research instrument. Astra is still described here as an internal system, however, so this is a claim about an unreleased model rather than a broadly available capability.
Verification is the immediate bottleneck
OpenAI says the mathematical arguments themselves were generated by its system, while the company helped prepare the manuscripts and formalize the proofs in Lean; it says it takes responsibility for their correctness and argues that attribution should reflect the AI’s contribution. That is a clear accountability and authorship position, but it is not the same as independent mathematical validation.
Gary Marcus’s critique identifies the unresolved questions precisely: math is unusually amenable to formal verification and synthetic data, while the public still does not know how Astra works, whether it relies on tools such as Lean, how many problems it attempted or solved, or whether independent verification has occurred. He argues that success in some forms of mathematics does not establish reliability in open-ended work, pointing to hallucination, document-reading, rule-following and even proof-clarity problems as separate tests.
The scrutiny became more concrete when QualiaQuanta asserted that at least one Astra proof was wrong, a claim Marcus amplified. The material available here establishes a public challenge, not a confirmed refutation; the defensible takeaway is therefore a serious AI-assisted mathematics result awaiting wider review—not evidence that mathematics, science or AGI has been solved.
Watch: DeepSeek V4-Flash is being compared on cost per completed task
A new comparison shifts the DeepSeek V4-Flash discussion from token price to task economics. Cline relayed an Artificial Analysis report claiming that DeepSeek completed the same benchmark tasks as Fable at 105× lower cost, while a follow-on reaction described two-orders-of-magnitude improvements as rare and significant.
The caveat is material: the same post warns that lower per-token pricing can be misleading if a model needs more turns to finish a task, and the text does not specify the benchmark or evaluation protocol. Treat the 105× figure as an important market signal to verify, not yet as an independently established performance result.
| Source | Docs | Insights | Status |
|---|---|---|---|
| AI at Meta | 0 | 0 | |
| Prof. Anima Anandkumar | 0 | 0 | |
| Ian Goodfellow | 0 | 0 | |
| Chip Huyen | 0 | 0 | |
| Oriol Vinyals | 2 | 1 | |
| Nathan Lambert | 4 | 2 | |
| Ashish Vaswani | 0 | 0 | |
| Sherjil Ozair | 0 | 0 | |
| Raquel Urtasun | 0 | 0 | |
| Greg Brockman | 9 | 2 | |
| Sebastian Raschka | 0 | 0 | |
| Thomas Wolf | 0 | 0 | |
| Jeremy Howard | 7 | 1 | |
| LocalLLM | 562 | 6 | |
| The Cognitive Revolution | 0 | 0 | |
| Richard Socher | 0 | 0 | |
| John Carmack | 0 | 0 | |
| Mustafa Suleyman | 0 | 0 | |
| Emad | 8 | 2 | |
| Tim Dettmers | 0 | 0 | |
| Geoffrey Hinton | 0 | 0 | |
| Machine Learning Street Talk | 0 | 0 | |
| hardmaru | 0 | 0 | |
| swyx | 24 | 4 | |
| Lukas Biewald | 0 | 0 | |
| a16z | 0 | 0 | |
| Nando de Freitas | 1 | 1 | |
| Pieter Abbeel | 0 | 0 | |
| Rowan Cheung | 0 | 0 | |
| Percy Liang | 0 | 0 | |
| Logan Kilpatrick | 0 | 0 | |
| Interconnects | 0 | 0 | |
| Latent.Space | 0 | 0 | |
| ChinAI Newsletter | 0 | 0 | |
| Big Technology | 0 | 0 | |
| Machine Learning | 19 | 2 | |
| Import AI | 0 | 0 | |
| Latent Space | 0 | 0 | |
| Gradient | 0 | 0 | |
| Lex Fridman | 0 | 0 | |
| Arxiv Insights | 0 | 0 | |
| Aleksa Gordić - The AI Epiphany | 0 | 0 | |
| Matt Wolfe | 0 | 0 | |
| sarah guo | 7 | 2 | |
| martin_casado | 15 | 1 | |
| Marc Andreessen 🇺🇸 | 0 | 0 | |
| Elad Gil | 4 | 2 | |
| François Chollet | 0 | 0 | |
| Vinod Khosla | 0 | 0 | |
| Yann LeCun | 0 | 0 | |
| Fei-Fei Li | 0 | 0 | |
| Ilya Sutskever | 0 | 0 | |
| Jeff Dean | 0 | 0 | |
| clem 🤗 | 0 | 0 | |
| Jim Fan | 0 | 0 | |
| Sara Hooker | 0 | 0 | |
| Soumith Chintala | 0 | 0 | |
| Sebastian Ruder @ ACL | 0 | 0 | |
| Dario Amodei | 0 | 0 | |
| Google DeepMind | 0 | 0 | |
| Demis Hassabis | 0 | 0 | |
| Satya Nadella | 0 | 0 | |
| Sam Altman | 2 | 0 | |
| Elon Musk | 15 | 1 | |
| Arthur Mensch | 0 | 0 | |
| Aravind Srinivas | 2 | 1 | |
| Aidan Gomez | 0 | 0 | |
| OpenAI | 0 | 0 | |
| Yannic Kilcher | 0 | 0 | |
| Andrej Karpathy | 1 | 1 | |
| Andrew Ng | 0 | 0 | |
| Sundar Pichai | 0 | 0 | |
| Kate Crawford | 0 | 0 | |
| Gary Marcus | 74 | 25 | |
| NVIDIA Blog | 0 | 0 | |
| Jay Alammar | 0 | 0 | |
| inFERENCe | 0 | 0 | |
| arg min | 0 | 0 | |
| Jack Clark | 0 | 0 | |
| Two Minute Papers | 0 | 0 | |
| Anthropic | 0 | 0 | |
| OpenAI | 0 | 0 | |
| Google DeepMind | 0 | 0 | |
| Gary Marcus | 0 | 0 | |
| Guillaume Lample @ NeurIPS 2024 | 0 | 0 | |
| Sarah Guo | 0 | 0 | |
| Christopher Manning | 0 | 0 | |
| Pieter Abbeel | 0 | 0 | |
| Ilya Sutskever | 0 | 0 | |
| Jerry Liu | 0 | 0 | |
| Sebastian Raschka | 0 | 0 | |
| Harrison Chase | 0 | 0 | |
| Elad Gil | 0 | 0 | |
| Sara Hooker | 0 | 0 | |
| Aidan Gomez | 0 | 0 | |
| Oriol Vinyals | 0 | 0 | |
| Jeremy Howard | 0 | 0 | |
| Simon Willison | 0 | 0 | |
| Arthur Mensch | 0 | 0 | |
| Andrej Karpathy | 0 | 0 | |
| Sam Altman | 0 | 0 | |
| Yoshua Bengio | 0 | 0 | |
| Geoffrey Hinton | 0 | 0 | |
| Andrew Ng | 0 | 0 | |
| Demis Hassabis | 0 | 0 | |
| Jeff Dean | 0 | 0 | |
| Nathan Benaich | 0 | 0 | |
| Clément Delangue | 1 | 1 | |
| Fei-Fei Li | 0 | 0 | |
| Ben Thompson | 0 | 0 | |
| Dario Amodei | 0 | 0 | |
| Yann LeCun | 0 | 0 | |
| Percy Liang | 0 | 0 | |
| Jack Clark | 0 | 0 |