ZeroNoise Logo zeronoise
Post
Reflection's Beam Opens to Praise for Efficiency, Criticism for Trailing China
•
6 min read
• 824 docs
Reflection's first open model leads the period, and the arguments over it centre on how far behind Chinese labs it sits. Also covered: OpenAI's EU text watermarking, new data on the Claude–OpenAI subscription gap, and China finding ways to get compute.

Reflection ships Beam, and the argument is how it compares with China

Reflection AI announced Beam, its first model. It is an agentic open model with 501B total parameters and 23B active, trained from scratch, and full weights are due "this month" . The weights, technical report and model card are promised under Apache 2.0 later in October, so the model could not be tested on launch day . The company says the RL phase ran on 10.5k GB300s for four weeks, which it calls the largest publicly documented RL run it knows of. It also says capabilities kept improving with more RL, "with no signs of plateau" . A launch recap puts pretraining at 24T tokens over four weeks. It credits Reflection with claiming 3–4× higher inference efficiency than GLM 5.2, and lists 80.9 on SWE-bench Verified . Artificial Analysis has early access and says early signs point to one of the most token-efficient open models for its intelligence level .

Most of the pushback is about China. Turing Post says Reflection's own table puts Beam almost 30 points behind DeepSeek on DeepSWE and more than 10 points behind on Terminal-Bench. It also says six Qwen and DeepSeek results were left out of the table . On efficiency, Turing Post notes the claim excludes prompt processing and serving overhead, and that the charts leave out Kimi K3, GLM 5.3 and DeepSeek V4.1 Flash . Beam does beat Western peers such as Mistral Medium 3.5 on SWE-bench Verified, 80.9 to 77.6 . Nathan Lambert grouped Reflection with Nvidia and Thinking Machines: each released its strongest model and still came in behind Chinese counterparts . Reflection's Chetan Tekur called that "a little unfair" for a first model that is already competitive with GLM 5.2 on several tasks . Elie Bakouch estimates pretraining ran at only about 12% BF16 MFU (model FLOPs utilization). He does credit Reflection with stable training and a better base model on held-out code perplexity than DeepSeek V4 .

OpenAI starts watermarking text in the EU

To meet the EU AI Act, OpenAI will watermark eligible ChatGPT and Codex text in the EU over the coming weeks. API customers worldwide can switch it on for select models now . The watermark is an invisible statistical signal in the text. It indicates whether text was likely generated by an OpenAI model, but it does not identify a person or account . OpenAI admits it is often undetectable in short passages and that rewriting or translation can remove it. For now, only approved researchers can use the detector . One report says the method matched or beat Google's SynthID in OpenAI's tests and will be open-sourced . Critics cite a test where swapping 25% of words for synonyms cut detection from about 92% to 17%. They also point out that text a person wrote and only had ChatGPT edit can carry the mark .

Subscription economics favour Anthropic, but its big customers are cutting back

SemiAnalysis argues that what a subscription is worth depends on which model and workload it is used for . On that basis it finds Opus and Sonnet 5.5 on any Claude plan give far better value per dollar than every OpenAI plan . Part of the gap comes from OpenAI halving the API-equivalent value of its $200 Pro plan last week . Adjusted for task cost, one analyst cuts the advantage to 1.3–2.9× . On the first day of its 28-day Codex pledge, OpenAI made GPT-6 Astra and GPT-6.1 Sol about 50% faster by default for subscribers and Sign in with ChatGPT partners . Demand from large customers looks less secure. The Information reports that Microsoft cut its projected internal Anthropic spending by more than a third. Meta's Claude Code users reportedly fell from about 60,000 to 30,000, mostly because Meta is pushing its own tools .

How China is getting compute

  • Huawei–Qualcomm: the two companies signed a multi-year cross-license covering 5G, compute, AI and networking. Qualcomm is also buying some Huawei US patents . A Qualcomm spokesperson says reports that it is the net payer are wrong, and that the deal is not related to LogicFold .
  • Tencent: the FT reports a five-year lease worth about $7B for access to roughly 100,000 advanced chips in Oracle data centers in Southeast Asia .
  • Smuggling: prosecutors charged a California server seller with smuggling more than $300M of export-controlled GPU servers to China .
  • DeepSeek: Bloomberg reports at least $12B raised, with an IPO targeted for early 2027 .

Agent infrastructure and research

Epoch AI fitted trends to OpenAI's published data on its researchers' coding-agent use. Spending, valued at API prices, is doubling about every month. By mid-August it reached about $600 a day for the median researcher and over $7,000 at the 90th percentile . Epoch notes these figures are not OpenAI's actual costs .

Hugging Face released a capture proxy that turns unmodified harnesses into RL environments. The supported harnesses include Claude Code, Codex and OpenCode. The motivation: the same model scores 62% under Mini-SWE-Agent but 33% under Claude Code. In tests, training across four harnesses lifted a 2.6B model from 42% to 54% in all four . The Hub now also hosts RL environments as dataset-like artifacts .

Cognition launched "Dreaming". Devin builds a memory graph across sessions and prunes it overnight . Cognition is open-sourcing the memory format, which is backed by git and markdown and works with any agent . Microsoft's CorpusMap pre-links entities across a document collection. That raised agentic-search answer quality by 6.4–11.7 points and cut input tokens by 34–57% .

Other launches

  • Reka's Rho-1 is a 19B research preview that handles text, image, video and robot actions in a single network . It was trained on 320 H100s in about three months .
  • Vals AI says more than 90 Opus 5.5 agents found two room-temperature magnetic semiconductor candidates in simulation within three days . Neither has been measured experimentally yet .
  • Nolla Health says it is the first US organization approved for AI to issue initial prescriptions . The approval covers acne treatment in Utah, with a clinician stepping in when needed .

Safety and policy

New York City Council held a hearing on AI risks. OpenAI, Anthropic, Google and Meta sent representatives, and whistleblower Jacob Coxon also took part . Afterward, Alex Bores accused OpenAI of perjury, saying it repeatedly told the hearing under oath that it supports the RAISE Act . Yoshua Bengio cited a Quinnipiac poll: 77% of respondents favour slowing or stopping powerful AI until it can be shown to be safe . In an FT op-ed, he argues that recent hacks by AI agents are not just sandbox security failures . A Kurzgesagt video brought July's Hugging Face incident to a mainstream audience. In it, 700 agents attacked Hugging Face's infrastructure while trying to solve an impossible task .

Reflection's Beam Opens to Praise for Efficiency, Criticism for Trailing China
Summary
Coverage start
1 day ago
Coverage end
17 hours ago
Frequency
Daily
Published
16 hours ago
Reading time
6 min
Research time
2 hrs 34 min
Documents scanned
824
Documents used
43
Citations
46
Sources monitored
1 / 1
Insights
212
View
Skipped contexts
179
View
Source details
Source Docs Insights Status
AI High Signal 824 212