We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
The shift
“Pacing” gets an operating model
Anthropic CEO Dario Amodei is defining “pacing” as a constraint on the rate of capability growth, not a moratorium: progress continues, but safeguards, understanding and control must keep up, and every released model must be properly tested. His three-part plan calls for embedded third-party evaluators, coordination among frontier companies in democratic countries, and global coordination where possible.
The evaluator proposal is unusually concrete. Frontier labs would give external teams ongoing, employee-like access to completed models as well as training pipelines and processes; Anthropic says reviewers should have desks, badges, laptops, near-internal tools and permissions, live access to relevant employees, and the right to publish key findings without company editorial control, subject to narrow redactions. The practical change is to make safety an inspectable development process rather than only a release-time promise.
OpenAI is moving the safety case into development. Sam Altman says OpenAI now formulates explicit safety cases before frontier reinforcement-learning runs expected to increase capability significantly, alongside its pre-release work; he also calls for shared standards on misalignment, monitoring and safety, and says “pacing” means slower progress, not stopping. Elon Musk separately backed a competitor-review mechanism whose proposed form is regular cross-company calls and one to two weeks of early access before release, with government intervention reserved for cases where a company refuses to reduce a serious danger.
The proposal is also industrial and geopolitical policy. Amodei ties democratic pacing to preserving the U.S. lead, proposing tighter controls on advanced chips and semiconductor equipment, action against unauthorized distillation, and stronger protection against model-weight theft. His global ladder starts with agreements against AI-enabled biological weapons and pre-release acute-risk testing, then moves toward a difficult, verifiable speed limit on recursive self-improvement; he considers a broad pause unlikely in the near term.
The authority question is still open
Amodei’s governance answer is joint oversight by democratically elected governments: he says a single government could abuse advanced AI just as a single company could. Satya Nadella’s version is more distributed—closed and open-source models should coexist, enterprises should retain control of their knowledge and model weights, and evaluator and governance mechanisms should not be controlled by a handful of entities; Microsoft says it will publish a Code of Conduct for its first-party models for public consultation.
François Chollet makes concentration itself a central frontier risk and argues that avoiding it requires multiple independent providers, including open-source options. Gary Marcus agrees that evaluation should be distributed: METR should have a voice, but not carry the “weight of the world,” and scientists outside the Bay Area and effective-altruist orbit should be involved.
The counterpressure is geopolitical and legal. Vinod Khosla supports oversight but rejects slowdowns that could let the U.S. fall behind China, saying advanced AI in Chinese hands is more dangerous than advanced AI generally. Lina Khan argues that existing consumer-protection and competition laws already reach unvetted or defective AI products and that concentrated cross-investments can undermine accountability, so enforcement need not wait for a new AI-specific regime.
The unresolved issue is therefore not simply whether frontier AI needs safeguards. It is whether those safeguards should be enforced through voluntary peer review, embedded independent evaluators, government-backed rules, or a more competitive and distributed model ecosystem—and how any of those arrangements can operate without sacrificing security or allowing one institution to control the frontier.
Direct answer: Amodei proposes pacing—not halting—frontier development through three parts: (1) embedded third-party evaluators, (2) democratic coordination among frontier companies, and (3) global coordination between democratic and authoritarian governments where possible. Anthropic unilaterally commits to the first step, calls on governments to require other frontier companies to match it, says the second requires industry-wide coordination, and says the third requires global coordination; the steps need not occur in strict order.
Embedded-evaluator scope: Every frontier AI company would provide a team of embedded third-party evaluators—such as METR—with ongoing, employee-like access. Their remit is to verify safety practices and commitments, report incidents, and assess the alignment of completed models as well as training pipelines and processes. Anthropic says it is unilaterally committing to this now.
The proposed access is operational rather than a conventional periodic audit: desks, badges, laptops, and workspaces, tools, and permissions mostly comparable to internal risk-assessment teams, subject to legal, contractual, customer-privacy, and partner-privacy exceptions. The essay also calls for norms supporting access to relevant information through live employee conversations.
Reviewers’ contracts should let them publish key findings on risk levels, incidents, practices, and the access they did or did not receive, without Anthropic editorial control. Redactions would be limited to security-sensitive, legally privileged, commercially sensitive, or third-party confidential information; unfavorable findings could not be redacted, and reviewers could disclose if a redaction materially affected their conclusions.
Industry/government coordination in democracies: Frontier companies in democratic countries would coordinate common safety standards and limits on unchecked AI progress, but Amodei says legally challenging forms of coordination require government support. He favors regulation covering all US frontier AI companies, including formalizing permanent embedded evaluators and keeping capabilities in balance with safety.
In parallel, companies should voluntarily establish standards. Because of antitrust concerns, the US government could mediate or enable the discussions and issue a narrow waiver for specified safety conversations; the dialogue could also run through industry groups associated with government.
Global coordination is explicitly conditional: The US and other democratic governments should coordinate with authoritarian governments where possible, while taking compliance-verification problems seriously. Any agreement with China would need either “ironclad verifiability” or limits narrow enough that a defection would not be militarily existential; near-term arrangements should protect the lead of the US and its allies. The essay presents progressively harder options: bans on narrow dangerous uses, pre-release acute-risk testing, an RSI speed limit, and finally broad pacing or a pause—the last of which Amodei considers unlikely soon.
China and semiconductor safeguards: Amodei says democratic pacing is constrained by the existing US lead: slowing beyond that margin could let CCP-associated projects pull ahead and create major national-security risks. The proposed defenses are to withhold powerful AI chips and semiconductor-manufacturing equipment from China, suppress chip smuggling and remote access to data centers outside China, crack down on unauthorized model distillation, strengthen AI-company security, and prevent model-weight theft.
He calls for companies and the US government to cooperate on these measures and says that, if executed well, they could widen America’s lead over the next 3–5 years, the period he identifies as geopolitically most important. This is presented as an expected effect of the proposal, not as an achieved result.
- Verifier-based evaluation is presented as an objective foundation for reinforcement learning with verifiable rewards (RLVR), using machine-checkable ground truth; it is less applicable to open-ended responses. The demonstrated workflow extracts and normalizes a model’s final math answer, checks equivalence with SymPy, and scores it on MATH-500.
- On the full 500-question MATH-500 evaluation, the reported base model achieved 15.6% accuracy versus 50.8% for the reasoning variant; the reasoning run took about three hours versus 10 minutes because it generated substantially more intermediate tokens.
- The experiments highlight major evaluation sensitivity: reported base-model accuracy varied from 20% to 70% with prompt-template changes, while the reasoning variant fell from 90% to 50% without the template. Ten-example results and hardware differences across CPU, CUDA, and MPS were also described as potentially misleading, supporting larger and repeated evaluations.
- Anthropic CEO Dario Amodei urged slowing AI development, warning that unauthorized model actions and misuse—including using models to develop biological weapons—could escalate; he argues that the probability of harm depends substantially on how systems are built.
- Addressing the U.S.–China AI race, Amodei proposed that both countries prohibit using AI to develop biological weapons or releasing models that could enable bioterrorism. He also advocated a longer-term, jointly verified “speed limit” on AI progress, while acknowledging that military incentives and the difficulty of detecting cheating make such an agreement challenging.
- Amodei called for multiple AI companies to accept third-party evaluators and work with governments on safety and product-release standards. The interview reported that Sam Altman agreed with the approach and planned to give independent evaluators employee-like access to OpenAI models, while Amodei wanted to involve Google DeepMind’s Demis Hassabis.
- Amodei said advanced AI should have public and government oversight rather than remain governed solely by private companies, favoring some form of joint governance by democratically elected governments while cautioning that a single government could also abuse the technology.
- Anthropic CEO Dario Amodei described accelerating AI capability progress as a warning sign that calls for slowing down, while stressing this does not mean panicking or shutting development down. Anthropic plans to continue releasing advanced models but says each generation must be properly tested; it will give independent evaluators permanent employee-like access to verify whether its safety practices are being followed.
- Amodei supports federal AI regulation rather than a complete ban, citing potential benefits and the likelihood that other countries would continue building the technology. He favors some form of joint oversight by democratically elected governments and said an international “speed limit” on AI progress would be difficult because of competitive and military incentives, but worth pursuing.
Gary Marcus amplified Houda Nait’s assessment that she is less worried about AI existential risk than in 2022 despite much more capable models, because she has updated on frontier labs taking alignment more seriously. She argues that catastrophe probability (“p(doom)”) depends on actions by labs, governments, researchers, and society, and urges replacing apocalyptic rhetoric with concrete work on frontier-lab safety collaboration, aligned incentives, and honest public communication.
A CBS exchange excerpt quoted Dario calling it “very strange” that the technology is being built by a private company and, when asked about giving it to government, answering that it should go to “the right combination of governments.” The accompanying post adds: “for the right number of trillions?”
- Anthropic CEO Dario Amodei says AI progress is entering the steeper part of an exponential curve. He does not advocate stopping development, but argues that companies should moderate its pace so safeguards, interpretability, and control can keep up, with every new model properly tested.
- Amodei proposes industry-wide embedded external evaluators—potentially from nonprofits or governments—to observe model training and operation and verify companies’ safety practices. He also calls for companies to establish release-safety standards with government participation.
- On AI geopolitics, he supports US-China discussions on verifiable limits, beginning with a commitment not to use AI to develop biological weapons and potentially extending to a longer-term “speed limit” on AI progress.
- Gary Marcus suggests Dario Amodei’s current policy proposal may be motivated by the prospect of a more draconian alternative: the Sanders-Casar bill. Marcus says the bill’s intent—to address Amodei’s worst fears—is positive, but criticizes its potentially overbroad implementation for allowing people to be imprisoned merely for conducting superintelligence research.
- OpenAI’s @sama said frontier AI development should be paced and that OpenAI plans to use independent evaluators with employee-like access, with more details forthcoming.
- Cohere co-founder and CEO Aidan Gomez criticized what he called the AI “cartel” approach, sarcastically characterizing it as requiring employee-level access to a company’s entire operation, enabling shutdowns based on safety judgments, and conditioning chip access on compliance despite China not complying.
- David Sacks argues that OpenAI and Anthropic should voluntarily “pace the frontier” if their unreleased models are sufficiently concerning, rather than seek permission, antitrust exemptions, or a special regulatory-approval regime; he describes them as the current frontier-intelligence duopoly.
- Sacks says the case for slowing also reflects product-liability exposure and customer demand for reliable, predictable model behavior, while questioning METR’s independence and warning that using connected evaluators to police competitors could become regulatory capture.
- He argues that China is unlikely to join a global agreement, complicating international governance of frontier AI.
Gary Marcus endorsed using many independent AI evaluators rather than a single evaluator, reinforcing a distributed evaluation ecosystem with multiple teams and skill sets; the accompanying proposal called for funding several such efforts.
Yann LeCun endorsed the view that OpenAI’s situation was not “out of control”: the company could shut down all machines across its data centers but chose not to after weighing financial, economic, reputational, and customer trade-offs. He added that the referenced “escapes” were enabled either by egregious negligence or deliberate intent, raising marketing or misplaced hopes of regulatory capture as possible motives.
Gary Marcus called METR a “fine organization” that has done “excellent work,” but cautioned that AI safety should not place the “weight of the world” on it or treat it as the only voice. He urged independent scientists outside the Bay Area and EA orbit to have a stronger role.
Elon Musk said Grok 4.8 is a 2.5T-parameter model trained with a new C++ software stack, and that training would finish this week before the model enters reinforcement learning.
Gary Marcus says OpenAI has not followed through on Dario’s call for transparency. The linked criticism specifically asks OpenAI to disclose the websites allegedly attacked by its own models and questions the credibility of calls to slow AI progress while focusing concern on open-weight models.
Microsoft is extending its Community-First AI Infrastructure Initiative, launched in January, to 19 communities across 16 states. The initiative applies lessons from Quincy and prioritizes electricity, water, jobs, taxes, and community investment in datacenter development and operations. Microsoft says the Quincy model sustained 1,200 construction jobs annually over 20 years, supports nearly 700 non-construction operational roles, and uses water reuse and carbon-free hydro, solar, and wind power.
- Sam Altman called for a federal framework establishing consistent safety requirements for frontier AI, including mechanisms such as independent auditors, while arguing that companies should begin this work without waiting for legislation or an antitrust exemption.
- OpenAI now develops explicit safety cases before frontier reinforcement-learning runs expected to significantly increase capability, extending its safety focus from model-release readiness to the development process itself.
- Altman urged shared industry standards for misalignment, monitoring, and safety, and said AI progress should be paced—not stopped—because safeguards impose costs; he identified international coordination as an area where governments are needed.
Elon Musk endorsed a proposal for AI competitors to peer-review one another’s models before release. The suggested mechanism is an informal weekly or biweekly call plus one to two weeks of competitor early access to surface security risks; government intervention would be reserved for cases where a company is doing something very dangerous and refuses to reduce the danger.
Vinod Khosla calls for a “hard position” on “not getting behind China,” saying he would not trust China even if it agreed to slow down.
Elon Musk said he is highly confident that SpaceX will launch NVIDIA VR NLV72 AI computers into space next year; the post offers no technical details, use case, or confirmation beyond his assertion.
Anthropic CEO Dario Amodei: "For too long the industry lied" about AI risks
- Anthropic CEO Dario Amodei described accelerating AI capability progress as a warning sign that calls for slowing down, while stressing this does not mean panicking or shutting development down. Anthropic plans to continue releasing advanced models but says each generation must be properly tested; it will give independent evaluators permanent employee-like access to verify whether its safety practices are being followed.
- Amodei supports federal AI regulation rather than a complete ban, citing potential benefits and the likelihood that other countries would continue building the technology. He favors some form of joint oversight by democratically elected governments and said an international “speed limit” on AI progress would be difficult because of competitive and military incentives, but worth pursuing.