ZeroNoise Logo zeronoise
Post
OpenAI releases 722 AI-generated math papers as AI-driven cyber findings rise sharply
•
6 min read
• 2153 docs
OpenAI published a large set of math results from an unreleased model, Mistral previewed a 1T-parameter model it plans to open-weight, a16z data shows AI changing vulnerability discovery, and new rounds went to Valon, General Medicine and Multiply Labs.

OpenAI's math release: big claims, still under review

OpenAI published mathematical results from an internal frontier model it has not released. It says it consulted the Institute for Advanced Study's Advisory Group on Mathematics and AI on how to release them . Reported figures:

  • Size: 722 manuscripts, grouped into 372 families of related results.
  • Source: an evaluation of about 4,000 research problems.
  • Compute: roughly three hours of ChatGPT Pro thinking per result, on average .
  • What's included: papers, proof artifacts and selected reasoning summaries .

Sam Altman called it "a new era of discovery" .

Reaction is split. Mathematician Levent Alpöge called it "the most significant moment in mathematical history" . One analysis estimates about 20% of the results are disproofs or counterexamples . Will Depue expects some results not to survive scrutiny. François Chollet asks whether progress in verifiable domains like math and code carries over to domains that still depend on human data . Individual highlights, such as integer multiplication faster than n log n, have not been independently verified .

Why it matters for investors: treat this as a claim about scientific capability to watch over the coming weeks, not a settled result. The open question for AI-for-science companies is whether these gains extend beyond domains where answers can be checked automatically.

Mistral Large 4: a strong preview with disputed benchmarks

Mistral previewed Large 4:

  • Size: 1T total parameters, 49B active, natively multimodal .
  • Availability: API now, open weights promised for the end of October. Mistral says it is "forged in Europe end-to-end" and is working privately with cybersecurity partners .
  • Training: about 3,800 Grace Blackwells in Europe. The RL run is still going and shows no sign of saturation, and a larger model is already training .
  • Pricing: $1.36 per million input tokens and $4.18 per million output tokens .

Independent results are mixed:

  • Artificial Analysis: an Intelligence Index score of 38, the top score from outside the US and China. But cost per task is more than 4x that of open models with similar scores .
  • Cyber benchmarks: Cline attributes much of the cyber lead to the model refusing fewer tasks .
  • Comparison with GLM: critics note it trails GLM-5.3 on Artificial Analysis's index .
  • Open-weight label: Clem Delangue pointed out that it can't be called the best open-weight model until the weights actually ship .

For investors, the takeaway is that Europe now has a credible sovereign model built on its own compute. Its cost competitiveness against Chinese open models is not yet established.

AI is producing far more security findings

a16z data shows reported critical vulnerabilities at 21 major software companies, including Apple, AWS, Microsoft and Google. They never exceeded 100 a month in four years, but have topped 600 a month since spring . About 87% of exploited bugs are now attacked on or before the day they become public, up from 23% in 2020 .

Armadin. Kevin Mandia, who built Mandiant, says Armadin has found more than 90 zero-day vulnerabilities at customer sites since January, all in production systems. It works as an outside "black box" attacker with no source code, using models post-trained with real exploit developers . By his own account, Armadin has no sales force or go-to-market strategy yet, and needs funding to build them .

OpenAI's internal effort. Greg Brockman said OpenAI moved 25% of its production engineers onto security, using models to find holes. That work eventually "saturated" at the critical issues Astra could find. OpenAI is now building an automated "defense factory" to rerun the process with each new model .

The investable idea is continuous, model-driven offense and defense that repeats with every new capability release.

Agents meeting businesses: a protocol, and a privacy problem

Meta and Sierra announced the Personal Agent Protocol, an open standard for how personal agents interact with businesses. Partners include Genesys, Instinct, Shopify, Stripe and Walmart . Sriram Krishnan compared it to OAuth. He asked how Amazon, airlines and other aggregators will respond if they can no longer own the final user experience . This picks up the question from the last brief about where network effects end up in agents.

In the same period, Time reported that Meta's Muse keeps hourly-updated dossiers on users and the people they mention. Lessons learned are shared across Muse agent instances, even though Meta describes each user's virtual machine as isolated .

How engineering work is changing

The Pragmatic Engineer reports that the most productive engineers now run 5–10 agent sessions in parallel and have stopped writing code by hand . Agent-authored pull requests on GitHub outnumbered human-authored ones in August . At Linear, agents have created more issues than humans since July .

Costs are a counterweight. Uber's token usage rose while its costs stayed flat, thanks to open models and smart routing between models . On capacity, Google's Amin Vahdat says Google aims to double effective token-serving capacity every six months, with as much or more of that coming from software as from hardware .

Support-agent benchmark. Gorgias, a vendor that ranks fifth in its own test, benchmarked 13 AI support vendors on 212 live stores. The top five resolve 64–75% of conversations with no human involved, and the median resolves 48% .

Deals

  • Valon raised a $150M Series D at a $2.3B valuation. It signed more than $200M in deals within six months of starting software sales. It first ran its own mortgage servicer until it reached 3x industry efficiency. One of the largest US servicers is moving 4M loans onto its platform, nearly 10% of the market .
  • General Medicine: a16z led a $120M Series B for this "general store for healthcare." It combines pricing, clinical guidance and a clinician marketplace, and is also building for enterprise customers. It reunites the PillPack team .
  • Multiply Labs (YC S16) raised a $75M Series B for robots that manufacture gene therapies and mRNA drugs. Its robots run inside AstraZeneca and Legend Biotech facilities .
  • Type One Energy raised $200M for fusion. Its CEO says this covers about half of a 400-MW plant targeted for 2034, built through a project-specific supplier network .
  • Outer Space launched with an $8M pre-seed led by Upfront. It builds outdoor structures that generate and store solar power for homes .

Investor debate and signals

Seed checks as options. Leo Polovets responded to Menlo's view that seed checks are options to be sized up later: "No seed investment should be an option bet." His fund does little follow-on investing by design .

Software moats. Naval argues models are software's last moat, as AI decompiles and recodes products. He expects more software to retreat to the server . Amjad Masad predicts all software will soon be "de facto open-source" .

Robotics. Generalist says its Gen 1.5 robot model improvised with unfamiliar tools, using a banana as a brush and a dustpan with both hands. The company says it was not explicitly trained to do this .

OpenAI releases 722 AI-generated math papers as AI-driven cyber findings rise sharply
Summary
Coverage start
1 day ago
Coverage end
17 hours ago
Frequency
Daily
Published
16 hours ago
Reading time
6 min
Research time
16 hrs 26 min
Documents scanned
2153
Documents used
23
Citations
39
Sources monitored
119 / 120
Insights
296
View
Skipped contexts
Source details
Source Docs Insights Status
SETH LEVINE's VC ADVENTURE 0 0
Hunter Walk 0 0
SaaStr 1 1
andrewchen 0 0
Elad Blog | Substack 0 0
AVC 0 0
Above the Crowd 0 0
Entrepreneur Ride Along 63 7
r/SideProject - A community for sharing side projects 328 73
Future(s) Studies 332 9
Artificial Intelligence (AI) 392 27
Software As a Service Companies — The Future Of Tech Businesses 646 83
Investing In AI 0 0
Big Technology 0 0
The Gradient 0 0
Import AI 0 0
Sam Altman 0 0
The community for ventures designed to scale rapidly | Read our rules before posting ❤️ 73 6
Co-Founder: Find Your Co-Founder Here 0 0
Entrepreneur 94 2
Naval 2 1
Machine Learning 57 4
Deep Learning 18 4
Natural Language Processing 0 0
Venture capital news and articles, for the VC industry 2 1
Newcomer 1 1
Jerry Liu 3 1
Harrison Chase 0 0
Cristóbal Valenzuela 5 2
Amjad Masad 3 2
Arthur Mensch 4 2
clem 🤗 6 3
Aidan Gomez 0 0
Kanjun 🐙 0 0
Suhail 0 0
Guillaume Lample @ NeurIPS 2024 7 1
Clouded Judgement 0 0
Bindu Reddy 5 5
Parag Agrawal 0 0
Harry Stebbings 2 1
Keith Rabois 0 0
Fred Wilson 0 0
Brad Feld 0 0
Exponential View 0 0
The Pragmatic Engineer 0 0
Latent.Space 2 2
Mark Suster 2 1
Benedict Evans 0 0
Allie K. Miller 3 3
Elizabeth Yin 💛 2 2
Roelof Botha 0 0
Andrew Reed 4 1
Luciana Lixandru 0 0
The Pragmatic Engineer 1 1
Elad Gil 0 0
Nathan Benaich 0 0
sarah guo 6 3
@jason 13 4
Vinod Khosla 0 0
Daniel Gross 0 0
Ann Miura-Ko 🦖 0 0
Mike Volpi 0 0
Aravind Srinivas 9 5
Ajay Agarwal 0 0
Leo Polovets 3 2
David Sacks 0 0
Lenny's Newsletter 0 0
Interconnects 1 1
Not Boring by Packy McCormick 0 0
Marc Andreessen 🇺🇸 0 0
Chris Dixon 0 0
Sriram Krishnan 5 3
a16z 23 12
benahorowitz.eth 0 0
martin_casado 0 0
andrew chen 3 2
Scott Kupor 6 4
David Ulevitch 🇺🇸 0 0
Dalton Caldwell 0 0
Y Combinator 1 1
Jessica Livingston 0 0
Paul Graham 1 1
Invest Like The Best 1 1
Garry Tan 12 3
Michael Seibel 0 0
Sam Altman 4 1
TechCrunch 0 0
Plug and Play Tech Center 0 0
No Priors: AI, Machine Learning, Tech, & Startups 0 0
Lex Fridman 0 0
Lightspeed Venture Partners 1 1
500 Global 0 0
Google for Startups 0 0
ThisWeekinStartups 0 0
Two Minute Papers 0 0
My First Million 0 0
Lenny's Podcast 0 0
All-In Podcast 0 0
Garry Tan 0 0
Y Combinator 0 0
Acquired 0 0
Foundation Capital 0 0
20VC with Harry Stebbings 0 0
Sequoia Capital 1 1
Greylock 2 2
Stanford eCorner 0 0
a16z 1 1
Jeremy Howard 0 0
Aravind Srinivas 0 0
Cassie Kozyrkov 0 0
Andrej Karpathy 0 0
Alexandr Wang 0 0
Naval Ravikant 0 0
Clément Delangue 0 0
Elad Gil 0 0
Fei-Fei Li 0 0
Andrew Ng 0 0
Demis Hassabis 0 0
Sam Altman 2 2
Yann LeCun 0 0