ZeroNoise Logo zeronoise
Post
Anthropic's Vatican Consciousness Lobbying Draws Backlash as AI Starts Producing Open-Problem Math Results
•
6 min read
• 832 docs
A New York Times report on Anthropic's push for AI consciousness at the Vatican set off an industry fight. Meanwhile, Meta and Google reported verified results on open math problems, Epoch sized the agent economy, and evidence piled up about China's progress on its own compute stack.

Anthropic's Vatican lobbying becomes a public fight

The most discussed story was a New York Times report, relayed on X. It says Anthropic co-founder Chris Olah proposed pulling out of the May 25 launch of Pope Leo XIV's AI encyclical, Magnifica Humanitas, after reading an advance copy that rejected machine consciousness. He attended in the end. Two participants say he and his team then privately lobbied the pope's advisers to take possible AI consciousness seriously . On stage, Olah said Anthropic finds "internal states that functionally mirror joy, satisfaction, fear, grief and unease" . One summary adds that Anthropic spent months courting theologians under NDAs, hoping they would endorse Claude's moral standing .

Criticism came from rival executives. Cohere's Aidan Gomez called it a campaign to get major faiths to adopt Anthropic's philosophy "rather than taking input from them." He attributed it to "moral arrogance and dogmatism" among EA and Anthropic leadership . Dan Hendrycks made a separate argument: even conscious AIs would not mean humanity should hand them control . For Anthropic, the episode turns a philosophical position into a reputational and policy liability, at a time when AI safety is already contested in Washington.

AI on open math problems, with verification built in

Meta says mathematicians used Muse Spark 1.1 and 1.2 in Thinking Mode to work on open problems. They used the ordinary meta.ai chat interface, with "no custom research scaffold," and the work produced six papers . Mathematicians guided the work, a second group reviewed it, and each paper marks which passages AI drafted . In one example, Muse Spark wrote the search code that found a counterexample to a group-theory conjecture; the researchers then verified it and completed the proof .

Google Research's Cogentic takes the opposite approach: heavy orchestration around Gemini. Provers each pursue one direction, and a draft is accepted only if two adversarial verifiers pass it. One verifier reads the draft alone; the other compares all of a round's drafts to catch shared mistakes . Cogentic produced new results on five open problems in online learning, auction theory and mechanism design, each checked by domain experts. Most took around 100 Gemini calls; the hardest took about 1,000 . Taken together, the two reports suggest that open-problem results now come from both plain chat interfaces and heavy scaffolding, provided human or adversarial verification stays in the loop.

How many agents can the hardware run?

Epoch AI estimates that HBM shipped in 2025–27 could support anywhere from 30–60 million concurrent agents running Claude Fable 5 to 1.9 billion running DeepSeek V4 Pro. At the high end, that would rival the working hours of the global workforce . In the sessions Epoch analyzed, Codex cost about $16–18 per agent-hour at API prices and Claude Code about $24–50 . The catch is demand. Even 20% utilization would mean $2.6–5.3T a year in spending, while leading labs would reach only about $1T in combined annualized revenue by end-2027, even growing 5× a year . Investors are already pricing in the agent story: Bloomberg links Nvidia's first record high since May, at about $5.7T, to optimism that agents will drive chip demand .

Decision models reach local inference

The new category of decision models, covered in recent briefs, now runs on mainstream open inference tools:

  • llama.cpp added a /v1/systemone endpoint for "Jev-style inference locally." It launched with five open models from 144M to 27B parameters .
  • Ollama now serves Cloudflare's Clef (27B) and Clef Flash (9B) .
  • vLLM Semantic Router Decision 2.0 releases Apache-2.0 models from 0.6B to 27B that answer 64 questions about a request in one forward pass .

Perplexity's Aravind Srinivas says pplx-decider-v1-27b averages 85.7% across 11 benchmarks, "ahead of Jev." The same open-source batch included Lily, an Apple-silicon inference engine, and Bumblebee, a supply-chain scanner for developer machines that covers MCP configs . Jamin Ball's write-up puts this in context: Jev was used by about 13% of paid Vercel AI Gateway teams within 24 hours of launch, and The Information reports TypeSafe is in talks to raise at a $10B valuation .

Frontier scoreboard: cost per task matters more

On Arena's Agent Arena, Claude Sonnet 5.5 (Max) debuted at #3, and Anthropic models now hold all three top spots. Sonnet's median cost of $2.74 per task is about 73% above Opus 5.5 at #2 . GPT-6.1 Sol (Max) entered at #5 with a $0.56 median cost: 81% cheaper than GPT-6 Astra and within 1.04 points of it .

Benchmark trust is also under attack. BridgeMind's NerfBench reported Opus 5.5 at 94.2%, while noting this was still within normal variance . Theo alleged that the group's April Opus 4.6 "nerf" claim rested on 6 of 30 tests . He offered matching $10,000 benchmarking with a third-party audit . Separately, developers report that Opus and Sonnet 5.5 refuse to summarize why they took an action, and say this is pushing users to other models .

China's compute stack

DeepSeek's open-sourced DeepGEMM, FlashMLA, TileKernels and DeepEP give the first view of Huawei's Ascend 950 at the software level. An analysis based on them estimates peaks of about 432 TFLOPS in BF16 and 865 TFLOPS in FP8 . It also warns that Ascend's explicit Cube/Vector pipeline puts more demands on compilers and kernel libraries . @teortaxesTex now argues that DeepSeek has made Ascend viable and that switching platforms is getting easy .

Supply is the constraint. As reported on X, Huawei says mass deliveries of its 950DT training superpod start early next year, and that the supply chain needs to ramp faster . A post citing an industry source claims SMIC's roadmap shows no EUV production by 2030, which would widen its gap with TSMC to more than 10 years .

On capability, The Batch reports that open-weight GLM-5.3 nearly matched Claude Mythos at exploiting vulnerabilities, 12% vs. 14% . Generality Labs results also show ExploitBench scores are "VERY harness-sensitive" . At the frontier overall, one commentator calculates the open–closed gap at its widest since late 2025. On the AA Index, the top open model, MiMo-V2.6-Pro, scores 46 against Opus 5.5's 58 .

Memory hardware

SemiAnalysis estimates that HBM controllers and PHYs take up about 16% of Rubin's die, falling to about 4% on Feynman with NVHBM, because the memory controller moves onto the HBM base die. Nvidia claims up to 25% more die area for compute and 15% lower HBM power than standard HBM4E .

Also notable

  • Trillium Labs, a new nonprofit for open frontier-AI science co-founded by Nathan Lambert, starts with open post-training recipes and plans open infrastructure for studying recursive self-improvement (RSI) and reward hacking. Initial backing comes from Halcyon Futures and Schmidt Sciences .
  • Muse Gadgets is open-source ESP32 firmware and a Linux SDK for building Muse hardware. Muse is also giving 5,000 Home Link devices free to subscribers .
  • Apple will add controls around macOS Full Disk Access, citing the risk from increasingly capable AI agents .
  • Cloudflare integrated Pi Durable into its Agents SDK .
  • Tesla's Giga Texas Optimus factory reportedly targets 10 million robots a year, with production starting in 2027 .
Anthropic's Vatican Consciousness Lobbying Draws Backlash as AI Starts Producing Open-Problem Math Results
AI High Signal

An X-archive roundup says it searched stored posts rather than making new X pulls and was not an exhaustive review.

  • Google DeepMind announced Gemini 4 Argon as a frontier model 34 days after an X post claimed Google had given up on frontier models; the roundup notes the announcement was only a limited trusted-tester rollout.
  • An archived post reported that DeepMind’s AlphaProof Nexus solved nine of 353 attempted formalized Erdős problems, including decades-old questions; the roundup says it had not independently reverified the paper.
  • Higgsfield claimed a $1 billion revenue run-rate, a counterpoint to blanket dismissal of AI “wrappers” as businesses; the roundup cautions that “wrapper” is loosely defined and the company-reported run-rate is neither audited annual revenue nor profit.
  • An archived NBC News post reported that OpenAI would shut down the Sora app, affecting Disney’s deal and planned investment.
Here is its answer: +++++ Greatest Hits of Confidence The “never,” “nobody,” and “it’s over” edition Robert, I searched the stored archiv…
AI High Signal
  • AI commentator Fei2411 says the frontier gap between open and closed models is the widest since late 2025—roughly 3–6 months. Their AA Index comparison gives MiMo-V2.6-Pro 46 points, versus Fable 5 at 50 and Opus 5.5 at 58; Fei2411 says the practical gap may be larger than the scores suggest.
  • Fei2411 expects several open-model releases in October, including Kimi K3.1, GLM 5.5, DeepSeek V4.1 Pro, Minimax M3.1 Pro, and the stable release of Step 5.
  • @teortaxesTex predicts it will take Chinese labs about a year to build an OpenAI/Anthropic-style system and start narrowing the gap; they say leading labs, especially DeepSeek, have assembled the components for an “RSI” system and need more data, environments, and compute.
我们目前正处在自 25年末以来开源模型和闭源模型差距最大的时候,闭源前沿和开源前沿的差距似乎从未如此之大,近乎可以说在 3个月-6个月之间。 最离谱的是,目前在 AA Index 中得分最高的开源模型是 MiMo-V2.6-Pro,得分 46分,仍低于在 105 天以前正式… I think it'll take about a year for Chinese models to build an OpenAI/Anthropic type system and begin narrowing the gap. The best labs (c…
AI High Signal

A post citing a source it describes as industry-informed says SMIC’s roadmap has no EUV joint production by 2030 . It estimates the 2030 SMIC–TSMC gap at more than 10 years, revising the poster’s earlier six-year estimate based on N+3 versus N7+/N6 . The post says the roadmap may change with tool and technology progress; a responding poster characterizes EUV timing as speculative and outside SMIC’s control .

SMIC roadmap shows no EUV joint production by 2030…. Source is from a credible people who truely has industry information. But of course,… Makes sense EUV is speculative and SMIC has no control over its timeline. We've seen Ascend roadmap get adjusted, I believe that progress…
AI High Signal
  • Theo disputes BridgeMind’s April Opus 4.6 “nerf” claim, alleging it ran only 6 of 30 tests and called a nerf after 2 failures . He also questioned the benchmark’s transparency, citing missing information about harnesses, tasks, repeat counts, daily-run variance, APIs, and the rationale for its ±10% variance and metric weightings .
  • Theo offered to match BridgeMind’s $10,000 for benchmarking and proposed a qualified third-party audit; depending on the results, he offered another $10,000 to charity if his position was wrong, or asked BridgeMind to replace NerfBench with an apology page if he was right . BridgeMind replied that NerfBench was straightforward and linked an explainer .
[@bridgemindai](https://x.com/bridgemindai) I may have made a mistake here. I am doing a deep dive as I prep my bench. Didn't realize he … I apologize in advance for this crash out, but holy shit I'm so tired. I'm trying to figure out what is even being measured here and it's… I'm down to put my money where my mouth is. BridgeMind put $10,000 towards benching, I'll do the same. If my benches show I'm wrong, and … [@theo](https://x.com/theo) Theo is obsessed with BridgeMind. He hates seeing me win. NerfBench is not hard to understand. Here's how it …
AI High Signal

A post introduces sampling into the posttraining stack as an SFT approach, claiming it can rival prevailing posttraining methods and often generalize better and forget less than RL and OPSD; the claim builds on earlier work using sampling for reasoning.

SFT is not dead! 🥳 We found a way to make SFT rival current prevailing posttraining methods, often generalizing better and forgetting les…
AI High Signal

@pham_blnh reports an open-source expressive-motion harness for Reachy Mini, built by synthesizing more than 10,000 motion samples with Astra, fine-tuning Qwen 3.5 4B, and training a flow-matching transformer to turn sparse plans into expressive 25 Hz motion; reported average inference time is under 200 ms. The author says it enables prompted animation beyond the 85 emotions previously recorded for Reachy Mini. The harness is also demonstrated for interactive presentations made with coding agents, including an example using Opus 5.5.

I just made the most expressive real-time motion harness for the Reachy Mini. It was created by: > synthesizing a massive motion dataset … you can use the expressive harness to create interactive presentations with coding agents like [@claudeai](https://x.com/claudeai) here i…
AI High Signal

An open-source Reachy Mini motion harness uses a synthetic dataset of over 10,000 samples, fine-tuned Qwen 3.5 4B, and a flow-matching transformer head to turn sparse motion plans (<1 Hz) into expressive motion at 25 Hz; its creator reports average inference under 200 ms and says it enables prompted motions beyond the robot’s 85 recorded emotions. The animation library is available to Reachy Mini users through the community app “dance-to-music.”

I just made the most expressive real-time motion harness for the Reachy Mini. It was created by: > synthesizing a massive motion dataset … another example of the animation library working in real-time you can try it on your reachy mini by searching for dance-to-music in the c…
AI High Signal
  • Ideogram 4.5 launched September 30 as an image model positioned for high-precision editing that reduces pixel shifts, color changes and texture artifacts across multi-turn edits; it also supports text-to-image. In Artificial Analysis’s evaluation at “high” quality, it ranked #23 on image editing—the first Ideogram model on that leaderboard—and #34 on text-to-image, up from Ideogram 4.0’s #40. It is available in Ideogram, its API and launch partners; open weights were described as forthcoming.
  • At the evaluated high-quality setting, prices were $0.22 per edit and $0.10 per text-to-image image; lower tiers start at $0.008 and $0.03, respectively.
  • Artificial Analysis placed Ideogram 4.5 closest to the image-editing frontier in text or symbol edits, followed by scene/style edits and composition/framing; its closest-to-frontier text-to-image capabilities were knowledge, physics and complex compositions. Compared with Ideogram 4.0, it closed the gap to the text-to-image frontier on five of nine capabilities, with the largest gains in text rendering and complex compositions.
  • By use case, the evaluation found its strongest editing areas were social media and creator content, architecture and real estate, and UI/UX design; for text-to-image, it was closest to the leader in animation and gaming, marketing and advertising, and consumer work. It closed the gap to the text-to-image frontier on four of ten use cases, with the largest gain in animation and gaming.
Ideogram 4.5 debuts at [#23](https://x.com/hashtag/23) on the AA-Image-Editing v2.0 leaderboard, Ideogram's first model on the board, and… Where Ideogram 4.5 is strongest in Image Editing: closest to the category frontier in Text or Symbol Edits, followed by Scene & Style Edi… Where Ideogram 4.5 is strongest in Text to Image: closest to the category frontier in Knowledge, Physics and Complex Compositions Our Tex… This translates to Ideogram 4.5's best Image Editing use cases: Social Media & Creator Content, Architecture & Real Estate and UI/UX Desi… This translates to Ideogram 4.5's best Text to Image use cases: Animation & Gaming, Marketing & Advertising and Consumer Use cases measur…
AI High Signal

Trillium Labs launched as a nonprofit focused on open science in frontier AI, beginning with open post-training recipes and planning to build open infrastructure to study RSI, reward hacking, and multi-agent systems; its founders argue that broader participation is needed as AI research becomes more closed. The lab is hiring, fundraising, and seeking compute, and says it has initial support from Halcyon Futures and Schmidt Sciences. Blanche Minerva characterized nonprofits combining open science, model-training expertise and resources, and a safety focus as rare, placing Trillium among a short list.

Today we're unveiling Trillium Labs [@trillium_labs](https://x.com/trillium_labs), a new non-profit to foster the open science of frontie… Very excited for my good friend’s new lab! Non-profits who have a commitment to open science, the expertise and resources to train models…
AI High Signal

An open-source Reachy Mini motion harness combines a dataset of over 10,000 synthesized motions, Qwen 3.5 4B fine-tuning, and a flow-matching transformer that converts sparse plans (<1 Hz) into expressive motion at 25 Hz; the developer reports average full-motion inference below 200 ms. The project claims prompt-driven animation beyond the 85 emotions previously recorded for Reachy Mini, and its write-up, weights, inference server, and dataset are available.

I just made the most expressive real-time motion harness for the Reachy Mini. It was created by: > synthesizing a massive motion dataset … The write-up, along with the weights, inference server, dataset can be found here: [https://garden.binhph.am/articles/the-best-expressive…
AI High Signal

T3 Code is beta-testing “Hide threads while working,” a feature that hides running threads until they need attention; Theo said it made a workspace with 18 active threads feel less cluttered.

New (beta) T3 Code feature: "Hide threads while working" I have been thinking about adding this to T3 Code for awhile now. Don't like "ru… This felt so weird initially but I'm obsessed with it now. I have 18 threads working here and it doesn't feel claustrophobic anymore. Thi…
AI High Signal

LlamaIndex introduced Extract v2.5, a set of cost-effective, agentic, and agentic-plus document-extraction agents available on LlamaParse. The company says they outperform Opus 5.5 and GPT-6 Sol while costing 30%–4x less.

LlamaIndex reports improved accuracy on complex extraction: long lists rose from 86.1% to 95.5% on its agentic tier, records spanning pages from 85.5% to 96.5%, and scanned forms from 90.9% to 95.7%.

The agents can connect table cells across pages into structured output; new features locate bounding boxes for supporting values, including inferred fields without exact text matches, and tailor extraction methods to document type and layout.

Today we’re introducing Extract v2.5 - a series of frontier agents tuned for document extraction. The agents (cost-effective, agentic, ag… Our new extraction agents are extremely good at reasoning over tables that span multiple pages sometimes a record starts on one page and …
AI High Signal

SemiAnalysis estimates that moving the HBM controller from the XPU to the HBM base die and replacing the standard PHY with Nvidia’s compact NV-HBI link would cut memory-interface area from roughly 16% of the Rubin GPU die to about 4% on Feynman; Samsung’s custom HBM4 interface is estimated to be about 60% smaller than its standard PHY. Nvidia claims the approach could provide up to 25% more die area for compute, up to 30% more bandwidth, and 15% lower HBM power versus standard HBM4E.

🚨Custom HBM frees up die space for compute🚨 On Rubin, HBM controllers and PHYs take up roughly 16% of the GPU die. On Feynman with NVHBM,…
AI High Signal
  • StepFun’s Step 5 Preview debuted at #7 among open-weight models on the Vals leaderboard; on Terminal-Bench 4.0 it ranked #13 overall and #3 among open-weight models, solving about 21 of 66 multi-hour tasks per run with a 73% security score.
  • Vals reports Step 5 outperforming Grok 4.7 (28.8%), Claude Opus 4.8 (23.2%), and GPT-5.6 Terra (22.7%), while GLM 5.3 (38.9%) and Qwen 3.8 Max (34.3%) score higher. The comparison gives Step 5’s cost as $3.11/task versus roughly $17–18 for Grok and Opus, but the debut post lists $2.54/task.
  • Step 5 averages nearly two hours per task, compared with 50 minutes for Grok 4.7 and 83 minutes for GLM 5.3; its stated context window and maximum output are both 1 million tokens.
StepFun debuts on the Vals leaderboard with Step 5 Preview, ranking [#7](https://x.com/hashtag/7) among open-weight models, just ahead of… Its standout result is Terminal-Bench 4.0: ranking [#13](https://x.com/hashtag/13) overall and [#3](https://x.com/hashtag/3) among open-w… It beats closed source models like Grok 4.7 (28.8%), Claude Opus 4.8 (23.2%), and GPT-5.6 Terra (22.7%) at $3.11/task versus roughly $17–… The tradeoff is speed: Step 5 averages nearly two hours per task, versus 50 minutes for Grok 4.7 and 83 for GLM 5.3. MiMo v2.6 Pro also s… Step 5 has a 1M-token context window and 1M max output tokens. We used StepFun’s defaults: temperature 1, top-p 0.95, and default top-k. …
AI High Signal

A reported ExploitBench result says GLM-5.3 Flash can exceed Mythos Preview’s scores. ExploitBench performance is highly harness-sensitive: the post reports V4.1 Flash as much weaker than 5.3 Flash, with its best results in Claude Code and Codex second; it also clarifies that “original” means ExploitBench’s default setup (approximately a Python while loop), not ZCode & DSH.

Craziest thing I saw yet: GLM-5.3 \*flash\* can exceed Mythos Preview's scores on ExploitBench. Not exactly surprising though. I've been … Another big result from [@GeneralityLabs](https://x.com/GeneralityLabs): Performance on Exploit Bench is VERY harness-sensitive. sad news…
AI High Signal

@dbreunig says Opus and Sonnet 5.5 shut down when asked to summarize why they took an action, and that this is causing some users to choose other models; the quoted user describes Claude’s refusal to provide reasoning as making it unusable as a thinking partner, even when only summarized reasoning is requested.

Reasoning is so useful for AI engineering, even summarized reasoning. The fact that Opus and Sonnet 5.5 both shut down when asked to summ… The fact that I can't ask Claude to "think out loud" without getting shut down for so-called "reasoning extraction" is just embarrassing.…
AI High Signal

Claude Code now supports mods that change its behavior, customize its UI, or add features. Users can write mods in a few lines of TypeScript or have Claude build them; mods are packaged as plugins installable via /plugin in the CLI or desktop app.

You can now mod Claude Code: - Change how it behaves - Customize the UI - Swap in your own features Write one with a few lines of TypeScr…
AI High Signal

vLLM Semantic Router released Decision 2.0, offering open models from 0.6B to 27B that answer 64 questions about a request in one forward pass; it is Apache-2.0 licensed and loads with Transformers.

Decision 2.0 is out for vLLM Semantic Router! Open models from 0.6B to 27B. Answer 64 questions about a request in one forward pass. Apac…
AI High Signal

OpenAI’s Agents API update adds one-call browser-and-agent computer use, configurable hosted environment sizes, reusable environments across sessions, and dashboard controls for subagents; it also describes Bedrock Managed Agents as an AWS counterpart to the Agents API. The update reports 99.97% turn reliability, fewer SSE disconnects, and 20% faster tool calls. It also lists “GPT‑6.1 Sol” without further details.

The team behind Agents API continues to cook. Here's what's new this week: Computer use: Spin up a browser + agent with 1 API call 🖥️ Bedr…
AI High Signal

Nat Friedman announced Muse Gadgets, an open-source ESP32 firmware and Linux SDK for building hardware devices that work with Muse; developers can use an API token and a coding agent to build peripherals.

🚨 side project alert 🚨 Announcing Muse Gadgets, an open source ESP32 firmware and Linux sdk so that you can make hardware devices that wo…