We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
OpenAI's math release: big claims, still under review
OpenAI published mathematical results from an internal frontier model it has not released. It says it consulted the Institute for Advanced Study's Advisory Group on Mathematics and AI on how to release them . Reported figures:
- Size: 722 manuscripts, grouped into 372 families of related results.
- Source: an evaluation of about 4,000 research problems.
- Compute: roughly three hours of ChatGPT Pro thinking per result, on average .
- What's included: papers, proof artifacts and selected reasoning summaries .
Sam Altman called it "a new era of discovery" .
Reaction is split. Mathematician Levent Alpöge called it "the most significant moment in mathematical history" . One analysis estimates about 20% of the results are disproofs or counterexamples . Will Depue expects some results not to survive scrutiny. François Chollet asks whether progress in verifiable domains like math and code carries over to domains that still depend on human data . Individual highlights, such as integer multiplication faster than n log n, have not been independently verified .
Why it matters for investors: treat this as a claim about scientific capability to watch over the coming weeks, not a settled result. The open question for AI-for-science companies is whether these gains extend beyond domains where answers can be checked automatically.
Mistral Large 4: a strong preview with disputed benchmarks
Mistral previewed Large 4:
- Size: 1T total parameters, 49B active, natively multimodal .
- Availability: API now, open weights promised for the end of October. Mistral says it is "forged in Europe end-to-end" and is working privately with cybersecurity partners .
- Training: about 3,800 Grace Blackwells in Europe. The RL run is still going and shows no sign of saturation, and a larger model is already training .
- Pricing: $1.36 per million input tokens and $4.18 per million output tokens .
Independent results are mixed:
- Artificial Analysis: an Intelligence Index score of 38, the top score from outside the US and China. But cost per task is more than 4x that of open models with similar scores .
- Cyber benchmarks: Cline attributes much of the cyber lead to the model refusing fewer tasks .
- Comparison with GLM: critics note it trails GLM-5.3 on Artificial Analysis's index .
- Open-weight label: Clem Delangue pointed out that it can't be called the best open-weight model until the weights actually ship .
For investors, the takeaway is that Europe now has a credible sovereign model built on its own compute. Its cost competitiveness against Chinese open models is not yet established.
AI is producing far more security findings
a16z data shows reported critical vulnerabilities at 21 major software companies, including Apple, AWS, Microsoft and Google. They never exceeded 100 a month in four years, but have topped 600 a month since spring . About 87% of exploited bugs are now attacked on or before the day they become public, up from 23% in 2020 .
Armadin. Kevin Mandia, who built Mandiant, says Armadin has found more than 90 zero-day vulnerabilities at customer sites since January, all in production systems. It works as an outside "black box" attacker with no source code, using models post-trained with real exploit developers . By his own account, Armadin has no sales force or go-to-market strategy yet, and needs funding to build them .
OpenAI's internal effort. Greg Brockman said OpenAI moved 25% of its production engineers onto security, using models to find holes. That work eventually "saturated" at the critical issues Astra could find. OpenAI is now building an automated "defense factory" to rerun the process with each new model .
The investable idea is continuous, model-driven offense and defense that repeats with every new capability release.
Agents meeting businesses: a protocol, and a privacy problem
Meta and Sierra announced the Personal Agent Protocol, an open standard for how personal agents interact with businesses. Partners include Genesys, Instinct, Shopify, Stripe and Walmart . Sriram Krishnan compared it to OAuth. He asked how Amazon, airlines and other aggregators will respond if they can no longer own the final user experience . This picks up the question from the last brief about where network effects end up in agents.
In the same period, Time reported that Meta's Muse keeps hourly-updated dossiers on users and the people they mention. Lessons learned are shared across Muse agent instances, even though Meta describes each user's virtual machine as isolated .
How engineering work is changing
The Pragmatic Engineer reports that the most productive engineers now run 5–10 agent sessions in parallel and have stopped writing code by hand . Agent-authored pull requests on GitHub outnumbered human-authored ones in August . At Linear, agents have created more issues than humans since July .
Costs are a counterweight. Uber's token usage rose while its costs stayed flat, thanks to open models and smart routing between models . On capacity, Google's Amin Vahdat says Google aims to double effective token-serving capacity every six months, with as much or more of that coming from software as from hardware .
Support-agent benchmark. Gorgias, a vendor that ranks fifth in its own test, benchmarked 13 AI support vendors on 212 live stores. The top five resolve 64–75% of conversations with no human involved, and the median resolves 48% .
Deals
- Valon raised a $150M Series D at a $2.3B valuation. It signed more than $200M in deals within six months of starting software sales. It first ran its own mortgage servicer until it reached 3x industry efficiency. One of the largest US servicers is moving 4M loans onto its platform, nearly 10% of the market .
- General Medicine: a16z led a $120M Series B for this "general store for healthcare." It combines pricing, clinical guidance and a clinician marketplace, and is also building for enterprise customers. It reunites the PillPack team .
- Multiply Labs (YC S16) raised a $75M Series B for robots that manufacture gene therapies and mRNA drugs. Its robots run inside AstraZeneca and Legend Biotech facilities .
- Type One Energy raised $200M for fusion. Its CEO says this covers about half of a 400-MW plant targeted for 2034, built through a project-specific supplier network .
- Outer Space launched with an $8M pre-seed led by Upfront. It builds outdoor structures that generate and store solar power for homes .
Investor debate and signals
Seed checks as options. Leo Polovets responded to Menlo's view that seed checks are options to be sized up later: "No seed investment should be an option bet." His fund does little follow-on investing by design .
Software moats. Naval argues models are software's last moat, as AI decompiles and recodes products. He expects more software to retreat to the server . Amjad Masad predicts all software will soon be "de facto open-source" .
Robotics. Generalist says its Gen 1.5 robot model improvised with unfamiliar tools, using a banana as a brush and a dustpan with both hands. The company says it was not explicitly trained to do this .
- The interviewer cited more than $200B in Google capex expected “this year,” mostly for data centers, and called the buildout historic; Vahdat agreed its scale and pace were unprecedented. Vahdat identified power as the most fundamental long-term constraint: Google prefers utility partnerships planned years ahead, says it covers the grid-infrastructure costs it triggers, and may use local generation or batteries to bridge supply gaps.
- Google’s target is to roughly double effective serving capacity, measured by token generation, every six months—not necessarily to double FLOPs. Vahdat said software and model improvements may contribute as much or more than hardware, and most observed intelligence-per-watt gains had come from model/software improvements.
- Long-horizon agents can issue follow-up requests in milliseconds rather than waiting for human-paced interactions; their CPU-based reasoning and context gathering also increase demand for CPUs, networking, and storage alongside accelerators.
- Google evaluates workload-specific “goodput”—useful completed work—rather than theoretical FLOPs; at 100,000 accelerators, Vahdat said failures could occur multiple times daily or hourly, with detection and recovery affecting delivered performance.
- AI data centers are increasingly co-designed around their hardware: accelerator racks can draw hundreds of kilowatts versus roughly 10–40 kW for storage racks, affecting power, cooling, and network design. Google also described separate 8i inference and 8t training chips as inference demand grew; both can run the other workload, and specialization must justify its fixed cost with durable demand.
- Google says hardware engineers now use AI as much as software engineers, with productivity gains and shorter design-to-tapeout and bring-up times; AI also helps assemble information for data-center planning but is not replacing human judgment.
- Google is pursuing orbital data centers as a “moonshot”; Vahdat cited 98–100% sunlight coverage in sun-synchronous orbit versus about 28–35% on land, while noting that cooling, reliability, and repairs are harder in space.
- Kevin, who says he built Mandiant and spent 30 years in security, says he joined Armadin after meeting founders David Slater, Travis Lam, and Evan Peña, whose team and proposed product convinced him the company addressed an urgent need created by AI’s impact on security. He argues that open models are already capable enough for structured vulnerability exploitation, and predicts that cheaper, more anonymous access to compute will expand criminal AI-enabled attacks, improve less-skilled attackers’ effectiveness, and make attribution harder.
- Armadin Red is positioned as continuous AI-enabled offensive security: agent swarms map network assets and services, monitor changes, and retest them; the company says it aims to verify exploitable risk rather than return only a list of known vulnerabilities.
- The interviewee reports that Armadin found more than 90 zero-day vulnerabilities since January 2026 in customer production environments, including at Fortune 500 companies, using outside-in, black-box testing without source-code review. He says most of those findings were discovered by humans while AI performed over 90% of routine work, but that the technology itself found their latest zero-days.
- In an Armadin test of 20 human-executed attack chains, the interviewee says no tested model completed more than eight; open-weight and leading closed models reached the same result, with closed models faster but open models catching up when given more time. He describes cost as the remaining differentiator.
- Armadin Blue is intended to turn exploit findings into compensating controls through EDR and firewalls; the interviewee described its first controls as forthcoming and initially limited, and said Armadin was working with CrowdStrike and Palo Alto Networks. Commercialization remains a stated gap: he said the company had no sales force, go-to-market strategy, or international operation yet, and that funding and scalable operating leadership were needed to build them.
- Generalist CEO Pete Florence says Gen 1.5 shows one- and few-shot learning across varied tasks and zero-shot physical generalization, capabilities he says were not explicitly trained for. In demonstrations, a robot trained to sweep with a brush used a banana as one; when given a dustpan, it used its other hand to sweep a cube into it, then tilted the dustpan to dump the cube into a bowl.
- Florence describes Gen Zero as establishing predictable robotics scaling laws as compute, data, and model size increase. He says Gen One was shown achieving 99%+ success rates on several tasks; his definition of mastery also includes speed, reliability, and improvisational responses to unexpected situations.
- Florence says Gen One had zero robot data in pretraining and achieved high success on varied tasks with one hour of data. Generalist says it is scaling data across more than 9,000 physically different hand configurations, including two- and five-finger hands and specialized tools; generalizing to new hands remains not fully solved.
- Generalist works with partners to deploy models in real environments and use real-world evaluations to identify capability gaps and guide research priorities. Florence says partners also contribute knowledge of relevant machines, processes, and workflows, while this work requires substantial relationship-building and deployment effort.
- Altara targets physical-science AI in semiconductors, batteries, and advanced materials, positioning its product as a bridge from fragmented experimental and production records, SEM images, wafer maps, simulations, and spreadsheets to experiment and production decisions. Altara described a semiconductor-value-chain customer analyzing historical R&D and production data, simulating a new experiment on Tuesday, and running it at a plant on Wednesday.
- Altara’s co-founder sees AI for science moving beyond helping researchers understand science toward helping them do it, with robotics, simulation, and agentic tools expanding AI’s role; she identifies scarce high-quality data, unresolved closed-loop experimentation and scientific evaluation, and real-world deployment as open challenges, while arguing small teams can accelerate R&D and commercialization.
- Altara’s founding engineer used semiconductor etching to illustrate the verification challenge: the process is costly, irreversible, and multistage, errors can propagate, and wafers are scarce; he described physics-based digital twins and simulation as potential verifiers. In his framework, falling offline-RL loss and stabilizing rewards do not establish physical-world performance; stronger checks compare with real-world data and introduce an overetch at one stage to test whether the policy changes downstream recipe and process conditions.
- AI use is shifting from assistance to delegated work: engineers interviewed describe running roughly 5–10 agent sessions in parallel, and the author says nearly all highly productive engineers he met who had hand-written code a year earlier had stopped doing so. Linear reports agent-created issues outnumbering human-created ones since July; GitHub data is described as showing rapid growth in agent-authored PRs, which the author says exceeded human-authored PRs in August. Factory AI reports that 83% of its users use agent skills, up 2.5x from February.
- Internal agent infrastructure is becoming common: the author says most mid-sized-and-larger companies have built their own harnesses, citing examples including Stripe, Uber, Shopify, Google, Meta, and Amazon. Development increasingly starts in Slack and triggers cloud agents, while agentic software factories connect agents to systems such as CI/CD, code review, and deployment; the author expects cloud coding agents and harnesses to expand, including offerings from vendors such as Anthropic and OpenAI and new startup products.
- AI software economics have two opposing signals: companies are adopting open models and smart routing to reduce token costs by 50% or more; at Uber, token use rose while costs stayed flat from May. Meanwhile, the author reports GPU and memory shortages alongside tighter CPU supply, with spot access difficult and some CPU reservations needed months ahead.
- More generated code is straining quality controls: an engineer at a mid-sized startup describes human reviews becoming “LGTM” rubber stamps under the volume, and Linear data shows AI-only PR reviews increasing. The author identifies agentic evaluation in CI/CD and agentic observability across deployed systems and incident management as emerging infrastructure needs. As a reliability caveat, Pi creator Mario Zechner says software feels increasingly brittle, while acknowledging that some of the problems predate agents.
- Leaner project staffing is emerging, but not the disappearance of engineering teams: Anthropic’s Head of Engineering says projects often have no more than two engineers because each is already managing agents, and the author says startup teams are shrinking; she also says teams continue to own software and on-call responsibilities and remain similar in shape to two-pizza teams. One Series D startup says it screens for AI-positive engineers. The author found no company where non-engineers were shipping production code; PMs and designers were prototyping or sending changes to engineers, who retained release responsibility.
- Altman described Dots as a personal AI helper that handles unwanted email and to-do work; in one example it evaluated ideas, found material on his computer, and proactively contacted three teammates. He said privacy must be “extremely” rigorous for the product to work.
- Altman said the planned Astra 6.1 release was held back because the ready model leaned too far toward accident risk; OpenAI is pursuing a model it can still distribute broadly while making stronger alignment claims.
- Altman supports open-source models despite warning of a coming wave of cybersecurity problems and saying society may have to accept fairly severe incidents in exchange for the liberty they provide. He also favors a pro-innovation approach with guardrails, pre-release standards and liability; he said safety assessment now needs to cover training, when models might escape a sandbox or hack systems.
- Reflection launched Beam, a text-only 501B-total/23B-active MoE trained from scratch for coding, agentic and scientific work, with full Apache 2.0 weights promised later that month. Team posts cite 23.8T pretraining tokens and RL/OPD training on 10K GB300s with over 100M rollouts across about 1M tasks; Reflection claims 80.9 on SWE-bench Verified and 3–4× the inference efficiency of GLM 5.2. The newsletter says newer leading Chinese models are generally ahead, with observers placing Beam around GLM 5.2 and below DeepSeek V4 Flash on some benchmarks; it nevertheless identifies demand for US-trained alternatives and calls the launch Reflection’s arrival as a functional neolab. Axios reported Reflection pays $150M/month for Colossus compute and has a $1B Nebius deal.
- Hugging Face’s multi-harness RL proxy can turn 10 unmodified agent harnesses into RL environments. The same weights scored 62% with Mini-SWE-Agent versus 33% with Claude Code; training across four harnesses raised first-attempt solves from 42% to 54%, while a tool-call bonus cut calls by 31%. Results are limited to one task family and one seed.
- Two model releases illustrate specialization: Reka’s Rho-1 is a 19B model for text, images, video and robot actions, trained from scratch on 320 H100s in about three months; Command Code’s Agr (31B) and Agr-flash (360M) return typed values and per-option probabilities for tool calls and routing rather than generating text.
- SemiAnalysis found Claude subscriptions delivered 5×+ the API-equivalent value of OpenAI plans, but adjusting for task cost narrowed the advantage to 1.3–2.9×; its methodology uses model- and token-specific credit costs, not list API prices. Users also reported paused new $200 sign-ups and effectively halved OpenAI usage limits; the newsletter reports default speeds for GPT-6 Astra and GPT-6.1 Sol subsequently rose about 50% across subscription surfaces and partners.
- AMD reportedly bought World Labs for $8.2B.
- Robotics models are moving from action-by-action imitation learning, which transfers poorly across actions and embodiments, toward visual-language-action models; emerging world-action models predict a future world state and actions to reach it, with broader generalization still a thesis. Robotics training also lacks an internet-scale corpus: companies use teleoperation, egocentric recordings and simulation, while learning from a robot’s own real-time interactions could provide higher-quality, embodiment-specific data.
- Deployment remains a potential moat: real-world robotics is difficult and operationally intensive, and companies with customer relationships and deployment experience may be able to adopt more general models later. The investor sees room for deployment-focused companies rather than only heavily capitalized foundation-model builders.
- Humanoid cost declines and hardware commoditization are not assured: actuators reportedly account for 40–60% of build cost, and actuator and rare-earth-magnet supply chains depend heavily on China amid tariffs and restrictions.
- Robotics data-collection businesses and robot-specific inference/physical-world chips are already emerging; data-quality tooling, observability and standardized evaluations are prospective infrastructure opportunities. Current evaluations are hard to compare across companies and matter for customer trust and safety. The generalist-model quadrant is described as the most capitalized; founder differentiation could come from architectural gains, proprietary data collection or greater computational efficiency.
- The guest’s thesis is that Meta, OpenAI, and Anthropic may evolve beyond LLMs into biotech, with non-invasive brain read/write as a shared goal; he sees precise neural modulation, with AI learning to titrate individual brain states, as a major frontier.
- An unnamed MIT team has developed an eye mask that directly measures REM and uses behind-ear stimulation to induce eye movements; the guest said it helped him fall asleep faster and amplified REM, and estimated release in 7–12 months.
- An unnamed neuroscientist and biotech founder is developing a cuff to directly measure autonomic nervous-system activity and distinguish distress from positive arousal; the guest said he had tried it and that it had been brought to police and first responders, with broader availability envisioned eventually.
- Technical caveat: the guest said current TMS and ultrasound stimulation lack good spatial control, while memory and thought remain difficult to model and neurons participate in multiple behaviors—complicating precise neural targeting.
- Altman says OpenAI paused a model-training run when capabilities were advancing faster than alignment and monitorability, and wants to move to each new capability level only when safeguards keep pace. OpenAI is discussing safety with other labs and governments; he says verifiable shared U.S.–China safety thresholds would be a major step.
- For a possible IPO, Altman says OpenAI is not in a rush and its nonprofit will continue to govern it; he says the company would prioritize its mission over shareholder returns or revenue, even if that makes its stock volatile.
- Altman expects AI to advance treatment of some diseases within several years and substantially change health care within a decade. He pointed to college students building a self-driving golf cart with AI-assisted engineering as an example of more ambitious projects becoming feasible, and forecast a major boom in entrepreneurship.
Garry Tan reports using Opus 5.5 in fast mode to implement Doom in Paul Graham’s Bel Lisp variant in about 20 minutes; the result comprised 32k C++ transpiled to under 2k lines of Bel, plus a full Bel interpreter in JavaScript. In a follow-up, Tan said maybe three more prompts brought it to 35 fps at 640×480, using eight CPU cores per frame. He shared a playable demo and source code.
Garry Tan says his GBrain bug-fixing workflow is chaotic because he opens a new thread per issue, but calls Capy’s cross-thread coordination and “virtual standup” conflict avoidance on parallel threads “SOTA.” He says GStack’s autoplan and eli5 help him move faster.
Scott Kupor recommended U.S. Tech Force and EarlyCareers.com as starting points for people seeking public-sector roles, amplifying Sriram Krishnan’s call to connect more tech talent with public-sector leaders.
Garry Tan says agents can use markdown skills and cron jobs to perform “almost every useful type of knowledge work.” He amplifies Meta chief AI officer Alexandr Wang’s account that, in some internal Meta cases, agent swarms with the right agent loop and evaluation metric can outperform a team of 100 engineers.
- General-purpose agents are not inherently winner-take-all: different agents can interoperate through messaging, APIs, and other channels, while portable text-based memory and context can ease switching. Better models, UX, integrations, and distribution may still support large companies, but Chen distinguishes those advantages from network effects.
- Potential agent-specific network effects include turning generated artifacts—such as shareable plans, microsites, or group workflows—into acquisition loops, and building shared identity, trust, reputation, private context, and intent that improve multi-party matching. Such viral loops depend on access to contacts and communication channels, plus user trust.
- The strategic fault line is where those network effects reside: open protocols and specialist networks could leave agents interchangeable and allow vertical agents to thrive, while horizontal agents that retain identity, relationships, context, and distribution could make switching mean leaving a network behind. Chen expects a contest between horizontal agents seeking to internalize network functions and incumbents or startups seeking to make their networks agent-accessible.
Paul Graham’s view is that money-motivated competitors are either ineffective or, if capable, eventually “get it” and stop working—a general observation about competitor incentives, not a company-specific development.
Scott Kupor says the “SI Force” will work to reconcile the administration’s goals of extending the U.S. lead and doing so responsibly .
Sam Altman said “We are entering a new era of discovery now” and linked to OpenAI’s page titled “Sharing AI progress in mathematics,” signaling a mathematics-focused AI progress update without further details in the post.
YC S16 startup Multiply Labs raised a $75M Series B to build robots that manufacture complex drugs, including gene therapies and mRNA treatments. The company targets a manufacturing bottleneck as AI speeds drug discovery: production is still mostly manual, slow, and vulnerable to contamination, and some gene-therapy doses can cost more than $1 million to produce. Its robots run production inside facilities at AstraZeneca and Legend Biotech.
Of the software bugs hackers actually exploit, about 87% are now attacked on or before the day they become public, up from 23% in 2020—evidence that the patching window is shrinking sharply.
Kevin Mandia describes Armadin’s approach as using agent swarms to map networks and poll for changes like a heartbeat; the company’s stated goal is to determine whether a known vulnerability is exploitable before attackers do.
Sam Altman on Donald Trump, Xi Jinping, and the Risk of Extinction (Part Two)
- Altman described Dots as a personal AI helper that handles unwanted email and to-do work; in one example it evaluated ideas, found material on his computer, and proactively contacted three teammates. He said privacy must be “extremely” rigorous for the product to work.
- Altman said the planned Astra 6.1 release was held back because the ready model leaned too far toward accident risk; OpenAI is pursuing a model it can still distribute broadly while making stronger alignment claims.
- Altman supports open-source models despite warning of a coming wave of cybersecurity problems and saying society may have to accept fairly severe incidents in exchange for the liberty they provide. He also favors a pro-innovation approach with guardrails, pre-release standards and liability; he said safety assessment now needs to cover training, when models might escape a sandbox or hack systems.