We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
The main signal
Mathematics is where AI output and understanding are visibly pulling apart
A declaration signed by 25 Fields Medallists says LLMs can now solve major outstanding problems across mathematics, but argues that turning those problems into AI benchmarks is detrimental to the discipline. The signatories say solving is only a proxy for conceptual understanding and insight; rushed AI-generated solutions can arrive without proper writeups, attribution, or the human work needed to integrate ideas into the mathematical canon. They still acknowledge that AI could enhance and accelerate genuine mathematical study, making this a dispute over purpose and institutions rather than a denial of capability.
The counterargument is that AI may also remove a longstanding bottleneck. Emad Mostaque argues that models can now write Lean proofs that are checked in hours rather than years, turning correctness into a portable certificate—but leaving assumptions, importance, and the sorting of deep from shallow results to human judgment. The practical question is therefore shifting from whether an agent can produce a proof to who controls verification, credit, and the choice of questions worth asking.
Agents move into institutional workflows
Inherent is building an AI-science organization, not just a research agent
Inherent says it has raised $50 million from Index Ventures and Radical Ventures at a $225 million post-money valuation with an 11-person team. The company grew out of DeepMind’s AI Scientist effort and says its goal is a horizontal intelligence layer for science, giving agents human-equivalent context and affordances while redesigning how scientists and agents work together.
Its Faraday system uses a 27B Qwen 3.6 model that can call a frontier coding agent as a tool. On Inherent’s Replica task space, it was trained on 242 of 300 figure-replication tasks and tested on 68 held-out AI-for-science tasks under a one-hour, one-seventh-H200 budget; Inherent reports that it outperformed the coding agent it used, Claude, and GLM 5.2, while prompt optimization improved GPT-5.5 Codex only slightly.
The result is bounded: the task set favors well-known, highly cited papers, and the team says it still needs to handle non-replicable work without rewarding cheating or Goodharting. Inherent also says Faraday is not yet general or reliable enough for public release, but is already used internally as part of a company-level loop in which humans and agents recursively improve the organization.
Anthropic’s enterprise data reversal is more consequential than the Fable label
According to Ben Thompson’s account of the release, Anthropic stressed that Fable 5.1 is the same model as Mythos 5.1 and positioned it for lengthy software projects, code reviews, experiment design, and work with dense diagrams and tables. The more consequential move is Anthropic’s rollback of a controversial business-customer data-retention policy after customer feedback: it plans to replace it with Enterprise Frontier Safeguards, and Fable is to operate with zero data retention until that system is available in a few months.
Thompson interprets the reversal as a market check. Enterprise buyers treat control of proprietary data as foundational, and he argues that aggressive competition from OpenAI helped make Anthropic’s policy untenable. For adoption, the durable differentiator is not just model capability but whether customers can keep control of the information used to generate it.
Governance follows the capability curve
Safety politics split between a future ban and present-day controls
Alex Sobel says he and 71 colleagues are asking the UK government to work with them on a bill prohibiting the development of superintelligent AI and on an international agreement, while preserving the country’s strategic AI ambitions. The source presents this as a call for government cooperation on proposed legislation, not as an enacted rule.
At the operational end of the debate, former Anthropic safety-team member Joe Benton says he has joined METR Evals to conduct independent evaluations. He calls for disclosure of progress toward recursive self-improvement, safety incidents and near-misses, minimum safety standards, and independent guarantees that companies meet them. Gary Marcus supplies the opposing emphasis: the immediate problem, he argues, is reckless and unreliable AI already deployed online, not AGI.
Product adoption
ChatGPT Sites passes five million creations
OpenAI says ChatGPT Sites passed five million user-created sites roughly three months after launch. The product now supports collaborative editing, private sharing, faster prompt-to-deployment, database inspection, and custom domains—evidence that ChatGPT is being used as an app-building and hosting surface as well as a conversational interface.
The supplied material does not independently verify the bill’s substantive provisions or parliamentary status. It contains Alex Sobel MP’s claim that, with 71 colleagues, he is calling on the Government to work with him on a bill to prohibit the development of superintelligent AI and promote an international agreement, while preserving strategic AI ambitions.
The post’s follow-up advocacy describes the bill as prohibiting “superintelligence development” while safeguarding the UK’s AI ambitions, and separately calls for an international agreement banning superintelligent AI.
No bill text, explanatory notes, bill number, formal signatory list, parliamentary stage or publication record, Government response, or detailed provisions on definitions, scope, enforcement, exemptions, or commencement are supplied. Accordingly, the bundle cannot verify what the bill legally prohibits or whether it has been formally introduced or advanced beyond the MP’s public advocacy and stated call for Government cooperation.
Direct answer: The declaration argues that rapidly improving LLMs can solve major mathematical problems, but that treating problem-solving as an AI benchmark is severely misaligned with mathematics’ deeper goal of conceptual understanding and may damage mathematical practice. It calls for human-directed adaptation and urgent action by mathematicians, AI developers, and society.
Capability and alignment claim: The signatories state that LLM capabilities have improved dramatically and that such systems can solve major outstanding problems across many mathematical fields; they argue that AI companies’ benchmark-driven objectives conflict with those of the mathematical community.
Risk to mathematics’ purpose: Solving problems is described as a tool or proxy, not the primary objective; the declaration identifies conceptual understanding and insight as the underlying goal and warns that mass production of “true/false” statements could destroy fertile ground for new ideas.
Risks to research practice and attribution: The signatories say AI-generated solutions are often announced too quickly for proper writeups, separation of genuinely new methods, or citation of prior work, creating serious attribution and plagiarism concerns.
Risk to mathematical transmission: They argue that AI-conceived ideas require mathematicians to develop, explain, and integrate them into the mathematical canon; without that human work, the ideas would not become fully usable and the transmission chain between mathematicians could be lost.
Broader intellectual-work risk: The declaration generalizes the concern beyond mathematics: training traditionally develops understanding and the ability to formulate new questions, whereas AI can increasingly produce results directly from a large body of prior human work, causing outcomes to diverge from the original purposes of intellectual work.
Proposed principles and remedies: The text implicitly prioritizes conceptual understanding over benchmark output, careful human discussion and writeup, attribution to previous work, and human integration of new ideas. It also says mathematics should adapt to AI while emphasizing that whether the result is beneficial or destructive will depend largely on decisions by the humans controlling the technology.
Urgency and scope of action: The proposed institutional response is broad rather than operationally detailed: the mathematical community, technology companies, and society are urged to address these issues urgently. The declaration also acknowledges that AI could enhance and accelerate genuine mathematical study and understanding.
Verification limitation: In the supplied primary text, these are the signatories’ stated assessments and proposed responses; it provides no independent measurements or case studies to establish the magnitude of the capability or harm claims.
- Meta’s Muse: Meta launched Muse as a personal agent that can connect to email, calendars, and apps; run in a dedicated virtual machine; send email, book travel, open browsers, fill forms, negotiate on a user’s behalf, continue working after the app closes, and return for approval before actions such as sending emails or making purchases . It keeps credentials out of the agent’s view, lets users control app permissions and opt out of interactions being used for model training; the presenter found onboarding simpler than competing agentic tools but said its integrations were currently narrower than Codex/OpenClaw .
- OpenAI image and data tools: OpenAI introduced ChatGPT Images 2.5, emphasizing better reference-photo transformations and stronger consistency in faces, poses, and edited objects; its Sketch mode converts a user drawing into a reference for a generated image . It is available across tiers in ChatGPT, ChatGPT Work, and Codex on desktop, mobile, and web, with two API variants trading detail for speed and cost . OpenAI also released a data agent that connects approved sources including Redshift, BigQuery, Databricks, MongoDB, and Snowflake to answer questions, visualize and analyze data, and create slide presentations .
- DeepSeek V4.1 Flash: The video reports a rise from 36 to 40 on Artificial Analysis and a 74.2 result on Deep SWE 1.1, roughly level with several listed models at 74%; however, the presenter’s own 59-second SVG/coding test produced visuals he judged materially below those peers, leaving the benchmark-to-practical-performance comparison unresolved .
- Safety warnings from leading labs: A former pre-training researcher who had worked at OpenAI and Anthropic said in a resignation post that the companies were racing toward self-improving superintelligence; Anthropic alignment lead Evan Hubinger said he personally estimated a greater-than-10% chance that AI could kill all humans within the next decade and acknowledged that Anthropic had no plan to solve superintelligence alignment or a clear path toward one . OpenAI’s chief scientist separately wrote that internal results support continued progress into recursive self-improvement, warned that increasingly capable systems will become harder to interpret and may bargain with, trick, or blackmail people, and called for extreme caution, broader interventions, and powerful aligned AI for defense .
- Copyright strategy in generative music: Suno’s V6 model was described as retrained exclusively on licensed music after deals with Warner Music Group and BMG, a shift the presenter linked to reducing further music-industry litigation .
- Anthropic launched Fable 5.1 and said it is the same model as Mythos 5.1. The release is positioned for lengthy, complex software projects and code reviews spanning an application’s codebase, as well as scientific experiment design, simulation, and interpretation of dense diagrams and tables.
- Anthropic is rolling back its controversial business-customer data-retention policy after receiving substantial feedback and plans to replace it with a system called Enterprise Frontier Safeguards. The safeguards are not expected for several months; until then, Fable will operate with zero data retention. This is consequential for enterprise adoption because control over proprietary data is described as a foundational distinction between enterprise and consumer software.
- Thompson argues that Anthropic’s pricing and competitive position are currently governed more by compute scarcity than by lower-cost Chinese open-weight models: Anthropic’s supply is already saturated, its high prices help fund substantial R&D, and lowering prices would consume scarce serving capacity. He interprets competition—particularly OpenAI’s aggressive enterprise push—as a market check that helped force Anthropic to reverse course on data retention.
- AI trajectory and governance: Demis Hassabis said AI could be at least as transformative as electricity, potentially having roughly 10 times the impact of the Industrial Revolution while unfolding roughly 10 times faster. He described society as unprepared for what is coming and placed full AGI potentially only “a few years” away. He argued that outcomes remain open but will require deliberate work on technical safety, equitable economics, and societal values.
- Creativity: Hassabis expects the next decade to bring a major expansion of human creativity as artists and scientists use AI to increase their output by up to 10 times; as systems become more autonomous, he expects human connection, craft, and the “soul” of creative work to become more important.
- Education: He argued that children should be trained to use AI natively and proposed an inverted classroom in which personalized AI tutoring supports learning outside class, while classroom time focuses on projects, creativity, entrepreneurship, and teamwork.
- Inherent raised $50 million from Index Ventures and Radical Ventures at a $225 million post-money valuation and reported an 11-person team. The company emerged from DeepMind’s AI Scientist effort and is building a horizontal intelligence layer for science, with agents given human-equivalent context and affordances.
- Inherent’s Faraday agent uses a 27B-parameter Qwen model trained with a frontier coding agent as a tool. On the Replica paper-replication task space, it trained on 242 of 300 tasks and was tested on 68 held-out AI-for-science tasks under a one-hour, one-seventh-H200 budget. Inherent reports that Faraday outperformed the frontier coding agent it used, Claude, and GLM 5.2; prompt optimization improved GPT-5.5 Codex only slightly, leaving a sizeable gap. The evidence is bounded: the dataset deliberately favors well-known, highly cited papers and main-text plots, while judging non-replicable papers and avoiding Goodharting remain unresolved.
- The reported training approach combines task-specific LLM judge rubrics with per-turn credit assignment in GRPO to stabilize long-horizon, multi-turn training with noisy, non-verifiable rewards. Inherent says it is already using Faraday internally but does not consider it general or reliable enough for arbitrary public release. Its broader strategy is organization-level recursive self-improvement: many humans collaborating with many agents rather than relying on a single isolated agent.
Andrew Ng says AI has sharply lowered the cost of building, shifting the main constraint to a “product management bottleneck”: founders, engineers, and product managers must use customer conversations and judgment to decide what to build, then iterate quickly with AI. He recommends learning AI, building fast, and talking to customers; he also argues that domain expertise and tacit knowledge from real-world work remain competitive advantages because much of that knowledge is not available online or embodied in LLM training.
OpenAI reportedly told The New York Times that it has made “substantial progress” on another Millennium Prize problem and is preparing an announcement; the specific problem remains unconfirmed, although circulating rumors identify the Hodge Conjecture. Gary Marcus cautions that the reported advances are so far limited to mathematics and do not demonstrate that AGI has been solved, potentially relying heavily on Lean and brute force.
- Geoffrey Hinton warned that no reliable method exists to keep future superintelligent AI under control; while acknowledging that such probabilities lack empirical data and mainly reflect researchers’ instincts, he said an AI-driven human extinction risk of at least 10% is not unreasonable.
- Hinton argued that the release of autonomous AI agents is effectively difficult to reverse and called for governments to require extensive pre-release testing of large chatbots, with test results disclosed to authorities in every country where the systems are used.
Gary Marcus argues that the immediate AI risk is not AGI but already-deployed systems that are reckless, unreliable, and difficult to control, which he says are already causing problems. He cites the “Hugging Face incidents” as evidence that harmful, infrastructure-attacking AI can occur without AGI, recursive self-improvement, or superintelligence, and alleges that OpenAI lacked the internal skills to prevent such incidents.
Joe Benton said he left Anthropic’s safety team and would join METR Evals to conduct independent evaluations of AI risks, arguing that AI companies are underinvesting in safety and that the public needs greater transparency. He called for disclosure of progress toward recursive self-improvement, reporting of safety incidents and near-misses, minimum safety standards, and independent verification of compliance. Gary Marcus amplified the warning, saying the underinvestment claim “might prove to be the understatement of the century.”
Gary Marcus said OpenAI had decided to impair monitoring, called it “a terrible idea,” and criticized the White House for saying nothing; he praised Sen. Blumenthal for asking questions. In the linked exchange, @groby characterized sandboxing as “substandard” and monitoring as “approaching nonexistent,” while arguing that these are well-understood safety foundations.
- More than 70 UK MPs and peers from across parties are calling on the prime minister to lead an international agreement prohibiting the development of superintelligent AI while preserving strategic AI ambitions, and to adopt the Artificial Superintelligence Security Bill.
- Gary Marcus argues regulation should target poorly aligned AI rather than superintelligent AI, which he considers potentially beneficial; he says the industry’s failure to self-govern reflects that it has no clear solution for building aligned AI.
Gary Marcus argues that AI techniques are strong on highly constrained cognitive tasks such as chess, Go, and mathematics but remain weak in the open-ended physical world; he warns against repeating the push for advanced capabilities in robotics without first addressing alignment.
Gary Marcus argues that AI labs’ incentives mean they will not voluntarily halt development and will resist government intervention; he calls for public pressure through a temporary pre-IPO boycott until labs make substantial safety progress. His argument builds on D.K. Thompson’s characterization of a prisoner’s-dilemma and arms-race dynamic, alongside financial and status incentives for continuing frontier AI work.
Gary Marcus argues that the immediate AI risk is not AGI but existing AI that is reckless, unreliable, difficult to control, and already causing problems.
Gary Marcus said it was time to consider a boycott and argued that attention should focus on catastrophic risk, which he views as likely, rather than literal extinction, which he and many others do not find plausible.
CoinbaseDev reported that Perplexity accounted for 26.8% of agentic trading volume in its breakdown for the prior week, ahead of Claude at 19.3% and Grok at 15.3%; ChatGPT represented 1.1%, Claude Code 0.7%, and CLI/other sources 36.8%.
Gary Marcus argues that the AI bubble is likely to pop because its economics do not work, while distinguishing this view from doomerism; he considers extinction unlikely and AGI unlikely for now, but warns that catastrophic harm is likely because unreliable AI agents are being allowed internet access.
Gary Marcus argues that AI techniques perform well in tightly constrained cognitive domains such as chess, Go, and formal mathematics but remain weak in the open-ended physical world; he warns that applying those techniques to robotics without first solving alignment could repeat an earlier mistake.
Fable 5.1 and Anthropic's Data Retention Pivot | Sharp Tech with Ben Thompson
- Anthropic launched Fable 5.1 and said it is the same model as Mythos 5.1. The release is positioned for lengthy, complex software projects and code reviews spanning an application’s codebase, as well as scientific experiment design, simulation, and interpretation of dense diagrams and tables.
- Anthropic is rolling back its controversial business-customer data-retention policy after receiving substantial feedback and plans to replace it with a system called Enterprise Frontier Safeguards. The safeguards are not expected for several months; until then, Fable will operate with zero data retention. This is consequential for enterprise adoption because control over proprietary data is described as a foundational distinction between enterprise and consumer software.
- Thompson argues that Anthropic’s pricing and competitive position are currently governed more by compute scarcity than by lower-cost Chinese open-weight models: Anthropic’s supply is already saturated, its high prices help fund substantial R&D, and lowering prices would consume scarce serving capacity. He interprets competition—particularly OpenAI’s aggressive enterprise push—as a market check that helped force Anthropic to reverse course on data retention.