ZeroNoise Logo zeronoise
Post
Self-sustaining AI worms sharpen the case for pacing frontier AI
17 hours ago
4 min read
920 docs
A research prototype used stolen compute to adapt and replicate cyberattacks, while 1,346 frontier-lab employees called for international tools to pace automated AI development. The period’s capability signals are similarly uneven: rapid progress in formal mathematics and post-training, but persistent weakness in open-ended research.

Safety and control

A prototype AI worm uses stolen compute to adapt and spread

Researchers from the University of Toronto, Vector Institute, University of Cambridge and ServiceNow describe a worm that runs an open-weight LLM on compromised machines, generates attack strategies for each target and propagates across Linux, Windows and IoT devices by exploiting common corporate-network vulnerabilities. The paper says stolen compute gives the attacker zero marginal cost per infection and makes safeguards tied to commercial AI APIs structurally irrelevant.

Import AI reports roughly 80% success in vulnerability detection, 53% in exploitation and 88% in self-replication—about 37% for a complete attack—using a harness with separate planning, judging, action, summary and progress nodes. It remains a proof of concept rather than a live-incident report, but the paper’s decentralized swarm design means the important shift is from fixed exploit code toward an agent that can adapt and retry across targets.

Frontier-lab employees ask for pacing tools

A statement signed by 1,346 employees of frontier AI companies says leading labs may be close to automating AI research and asks the US government to support an international effort to develop technical and governance tools to “deliberately pace” frontier automated AI development. Its rationale is a coordination problem: companies and countries face pressure not to slow unilaterally, while the world lacks tools to buy time for security and oversight.

Import AI says the signatories include chief scientists and cofounders from OpenAI, Anthropic, Google DeepMind, Meta and Safe Superintelligence, among others. The proposal is framed as a way to create room for safeguards, not as a unilateral halt to research.

Capability is advancing unevenly

OpenAI puts Astra’s ten mathematics results into an audit workflow

OpenAI’s official account says an internal version of Astra resolved or substantially advanced ten long-standing problems across geometry, coding theory, circuit complexity, group theory, operator algebras, quantum complexity, lattice cryptography and extremal combinatorics. The company estimates roughly $2,000 in token cost at Sol API rates; humans prepared the manuscripts, Astra formalized each argument in Lean, and OpenAI is releasing the manuscripts, certificates and reasoning walkthroughs for examination.

OpenAI says it takes responsibility for correctness and that the mathematical arguments themselves were generated by its system. That makes expert scrutiny the next meaningful test: the release supplies material for mathematicians to examine, but the announcement is not a substitute for wider validation.

Shadow evaluation finds the open-ended research bottleneck

A separate arXiv study introduces “shadow evaluations,” in which an agent tackles the central question of an unpublished paper and the original authors grade the result. In two unpublished NeurIPS 2026 case studies, frontier agents received six days and thousands of dollars of compute, completed the engineering without human help, but made no substantial progress on the research questions; both papers were rejected, with recurring failures in research judgment, creativity, backtracking, resource awareness and instruction-following.

The contrast is useful rather than contradictory: human-defined, verification-friendly problems can now produce striking results, while choosing worthwhile questions and abandoning bad approaches remain unresolved. The study calls this early evidence from only two case studies, so it narrows claims about research automation rather than settling them.

Products and economics

DeepSeek’s flash update makes post-training the battleground

A Two Minute Papers review reports that DeepSeek’s updated flash model more than doubled many benchmark results, improved one result sevenfold, and beat both its previous flash version and a pro model roughly five times larger. The review says the architecture and size did not change; the gain came from post-training that taught the model when to plan, check its work and recover from mistakes.

The same review says the weights can be downloaded for local use or accessed through an API without session caps. It also notes that benchmarks are not everything, so the precise comparison still needs independent checking; if the reported gains hold, the competitive boundary is shifting toward post-training efficiency and open distribution, not parameter count alone.

Inference engineering is becoming a market layer

The financing number needs correcting: a Latent Space episode described Baseten as having raised a “$13B round,” while Baseten’s own announcement says its Series F raised $1.5B. Baseten reports 20× revenue growth and 40× inference-volume growth over the last year, and says the financing will fund compute, software and talent for production inference.

The broader signal is that serving weights quickly, reliably and affordably is becoming its own discipline, alongside model training. In the accompanying discussion, Baseten engineers describe 20–200% optimization gains and an aggressive path from roughly 30–40 to 300–400 tokens per second, while warning that speed comparisons depend heavily on hardware, load and prompt characteristics.

GPT-Live rebuilds voice around continuous media

OpenAI describes GPT-Live as a third-generation voice system whose full-duplex model listens and speaks simultaneously; deeper reasoning and tool use run on a separate asynchronous path so slow application work does not stall the audio stream.

The supporting systems redesign is substantial: OpenAI says its WARP protocol cuts media and data startup from six network round trips to one, while Instant Connect can let a client start a session with a single UDP packet. The product shift is therefore architectural, aimed at decoupling conversational responsiveness from the latency of deeper model calls.

Self-sustaining AI worms sharpen the case for pacing frontier AI
Summary
Coverage start
1 day ago
Coverage end
17 hours ago
Frequency
Daily
Published
16 hours ago
Reading time
4 min
Research time
10 hrs 2 min
Documents scanned
920
Documents used
10
Citations
18
Sources monitored
114 / 114
Insights
Skipped contexts
140
View
Source details
Source Docs Insights Status
AI at Meta 0 0
Prof. Anima Anandkumar 2 1
Ian Goodfellow 0 0
Chip Huyen 0 0
Oriol Vinyals 0 0
Nathan Lambert 5 0
Ashish Vaswani 0 0
Sherjil Ozair 0 0
Raquel Urtasun 0 0
Greg Brockman 2 1
Sebastian Raschka 0 0
Thomas Wolf 0 0
Jeremy Howard 0 0
LocalLLM 639 5
The Cognitive Revolution 0 0
Richard Socher 0 0
John Carmack 0 0
Mustafa Suleyman 0 0
Emad 0 0
Tim Dettmers 0 0
Geoffrey Hinton 0 0
Machine Learning Street Talk 0 0
hardmaru 3 1
swyx 6 2
Lukas Biewald 0 0
a16z 0 0
Nando de Freitas 0 0
Pieter Abbeel 0 0
Rowan Cheung 3 0
Percy Liang 0 0
Logan Kilpatrick 1 0
Interconnects 1 0
Latent.Space 1 1
ChinAI Newsletter 1 1
Big Technology 0 0
Machine Learning 158 0
Import AI 1 1
Latent Space 1 1
Gradient 0 0
Lex Fridman 0 0
Arxiv Insights 0 0
Aleksa Gordić - The AI Epiphany 0 0
Matt Wolfe 0 0
sarah guo 3 0
martin_casado 7 0
Marc Andreessen 🇺🇸 2 0
Elad Gil 2 1
François Chollet 0 0
Vinod Khosla 0 0
Yann LeCun 6 2
Fei-Fei Li 0 0
Ilya Sutskever 0 0
Jeff Dean 0 0
clem 🤗 1 0
Jim Fan 0 0
Sara Hooker 0 0
Soumith Chintala 0 0
Sebastian Ruder @ ACL 0 0
Dario Amodei 0 0
Google DeepMind 0 0
Demis Hassabis 0 0
Satya Nadella 0 0
Sam Altman 0 0
Elon Musk 21 1
Arthur Mensch 0 0
Aravind Srinivas 0 0
Aidan Gomez 2 0
OpenAI 0 0
Yannic Kilcher 0 0
Andrej Karpathy 0 0
Andrew Ng 0 0
Sundar Pichai 0 0
Kate Crawford 0 0
Gary Marcus 39 12
NVIDIA Blog 0 0
Jay Alammar 0 0
inFERENCe 0 0
arg min 0 0
Jack Clark 0 0
Two Minute Papers 1 1
Anthropic 0 0
OpenAI 7 2
Google DeepMind 0 0
Gary Marcus 0 0
Guillaume Lample @ NeurIPS 2024 0 0
Sarah Guo 0 0
Christopher Manning 0 0
Pieter Abbeel 0 0
Ilya Sutskever 0 0
Jerry Liu 0 0
Sebastian Raschka 0 0
Harrison Chase 0 0
Elad Gil 0 0
Sara Hooker 0 0
Aidan Gomez 0 0
Oriol Vinyals 0 0
Jeremy Howard 0 0
Simon Willison 0 0
Arthur Mensch 0 0
Andrej Karpathy 0 0
Sam Altman 0 0
Yoshua Bengio 0 0
Geoffrey Hinton 0 0
Andrew Ng 0 0
Demis Hassabis 0 0
Jeff Dean 0 0
Nathan Benaich 0 0
Clément Delangue 5 5
Fei-Fei Li 0 0
Ben Thompson 0 0
Dario Amodei 0 0
Yann LeCun 0 0
Percy Liang 0 0
Jack Clark 0 0