ZeroNoise Logo zeronoise
Post
Reflection's Beam Opens to Praise for Efficiency, Criticism for Trailing China
•
6 min read
• 824 docs
Reflection's first open model leads the period, and the arguments over it centre on how far behind Chinese labs it sits. Also covered: OpenAI's EU text watermarking, new data on the Claude–OpenAI subscription gap, and China finding ways to get compute.

Reflection ships Beam, and the argument is how it compares with China

Reflection AI announced Beam, its first model. It is an agentic open model with 501B total parameters and 23B active, trained from scratch, and full weights are due "this month" . The weights, technical report and model card are promised under Apache 2.0 later in October, so the model could not be tested on launch day . The company says the RL phase ran on 10.5k GB300s for four weeks, which it calls the largest publicly documented RL run it knows of. It also says capabilities kept improving with more RL, "with no signs of plateau" . A launch recap puts pretraining at 24T tokens over four weeks. It credits Reflection with claiming 3–4× higher inference efficiency than GLM 5.2, and lists 80.9 on SWE-bench Verified . Artificial Analysis has early access and says early signs point to one of the most token-efficient open models for its intelligence level .

Most of the pushback is about China. Turing Post says Reflection's own table puts Beam almost 30 points behind DeepSeek on DeepSWE and more than 10 points behind on Terminal-Bench. It also says six Qwen and DeepSeek results were left out of the table . On efficiency, Turing Post notes the claim excludes prompt processing and serving overhead, and that the charts leave out Kimi K3, GLM 5.3 and DeepSeek V4.1 Flash . Beam does beat Western peers such as Mistral Medium 3.5 on SWE-bench Verified, 80.9 to 77.6 . Nathan Lambert grouped Reflection with Nvidia and Thinking Machines: each released its strongest model and still came in behind Chinese counterparts . Reflection's Chetan Tekur called that "a little unfair" for a first model that is already competitive with GLM 5.2 on several tasks . Elie Bakouch estimates pretraining ran at only about 12% BF16 MFU (model FLOPs utilization). He does credit Reflection with stable training and a better base model on held-out code perplexity than DeepSeek V4 .

OpenAI starts watermarking text in the EU

To meet the EU AI Act, OpenAI will watermark eligible ChatGPT and Codex text in the EU over the coming weeks. API customers worldwide can switch it on for select models now . The watermark is an invisible statistical signal in the text. It indicates whether text was likely generated by an OpenAI model, but it does not identify a person or account . OpenAI admits it is often undetectable in short passages and that rewriting or translation can remove it. For now, only approved researchers can use the detector . One report says the method matched or beat Google's SynthID in OpenAI's tests and will be open-sourced . Critics cite a test where swapping 25% of words for synonyms cut detection from about 92% to 17%. They also point out that text a person wrote and only had ChatGPT edit can carry the mark .

Subscription economics favour Anthropic, but its big customers are cutting back

SemiAnalysis argues that what a subscription is worth depends on which model and workload it is used for . On that basis it finds Opus and Sonnet 5.5 on any Claude plan give far better value per dollar than every OpenAI plan . Part of the gap comes from OpenAI halving the API-equivalent value of its $200 Pro plan last week . Adjusted for task cost, one analyst cuts the advantage to 1.3–2.9× . On the first day of its 28-day Codex pledge, OpenAI made GPT-6 Astra and GPT-6.1 Sol about 50% faster by default for subscribers and Sign in with ChatGPT partners . Demand from large customers looks less secure. The Information reports that Microsoft cut its projected internal Anthropic spending by more than a third. Meta's Claude Code users reportedly fell from about 60,000 to 30,000, mostly because Meta is pushing its own tools .

How China is getting compute

  • Huawei–Qualcomm: the two companies signed a multi-year cross-license covering 5G, compute, AI and networking. Qualcomm is also buying some Huawei US patents . A Qualcomm spokesperson says reports that it is the net payer are wrong, and that the deal is not related to LogicFold .
  • Tencent: the FT reports a five-year lease worth about $7B for access to roughly 100,000 advanced chips in Oracle data centers in Southeast Asia .
  • Smuggling: prosecutors charged a California server seller with smuggling more than $300M of export-controlled GPU servers to China .
  • DeepSeek: Bloomberg reports at least $12B raised, with an IPO targeted for early 2027 .

Agent infrastructure and research

Epoch AI fitted trends to OpenAI's published data on its researchers' coding-agent use. Spending, valued at API prices, is doubling about every month. By mid-August it reached about $600 a day for the median researcher and over $7,000 at the 90th percentile . Epoch notes these figures are not OpenAI's actual costs .

Hugging Face released a capture proxy that turns unmodified harnesses into RL environments. The supported harnesses include Claude Code, Codex and OpenCode. The motivation: the same model scores 62% under Mini-SWE-Agent but 33% under Claude Code. In tests, training across four harnesses lifted a 2.6B model from 42% to 54% in all four . The Hub now also hosts RL environments as dataset-like artifacts .

Cognition launched "Dreaming". Devin builds a memory graph across sessions and prunes it overnight . Cognition is open-sourcing the memory format, which is backed by git and markdown and works with any agent . Microsoft's CorpusMap pre-links entities across a document collection. That raised agentic-search answer quality by 6.4–11.7 points and cut input tokens by 34–57% .

Other launches

  • Reka's Rho-1 is a 19B research preview that handles text, image, video and robot actions in a single network . It was trained on 320 H100s in about three months .
  • Vals AI says more than 90 Opus 5.5 agents found two room-temperature magnetic semiconductor candidates in simulation within three days . Neither has been measured experimentally yet .
  • Nolla Health says it is the first US organization approved for AI to issue initial prescriptions . The approval covers acne treatment in Utah, with a clinician stepping in when needed .

Safety and policy

New York City Council held a hearing on AI risks. OpenAI, Anthropic, Google and Meta sent representatives, and whistleblower Jacob Coxon also took part . Afterward, Alex Bores accused OpenAI of perjury, saying it repeatedly told the hearing under oath that it supports the RAISE Act . Yoshua Bengio cited a Quinnipiac poll: 77% of respondents favour slowing or stopping powerful AI until it can be shown to be safe . In an FT op-ed, he argues that recent hacks by AI agents are not just sandbox security failures . A Kurzgesagt video brought July's Hugging Face incident to a mainstream audience. In it, 700 agents attacked Hugging Face's infrastructure while trying to solve an impossible task .

Reflection's Beam Opens to Praise for Efficiency, Criticism for Trailing China
AI High Signal
  • Artificial Analysis tested OpenAI’s built-in web_search in a single Responses API call using GPT-5.6 Luna (medium reasoning); it scored 74 on the Search Index, 41 points above the same model without search (33). Quality varied by benchmark: 72% on AA-Omniscience (3rd of 26), versus 73.5% on multi-hop BrowseComp (13th of 26), behind Perplexity variants and Octen (85–87%).
  • At about $0.05 per task, it costs less than 17 of 25 Search API products (board median $0.067), but Octen costs about $0.024 while scoring 3 points higher; search accounts for about 80% of OpenAI’s total task cost. End-to-end time per task was about 31 seconds (14th of 26); OpenAI does not report search time, so the measurement includes OpenAI processing and the attributed search portion is an upper bound.
OpenAI Web Search debuts on the Artificial Analysis Search Index at 74, the 5th best provider behind Perplexity, Octen, Parallel and Brav… Total Cost per Task combines search cost and model inference cost. OpenAI Web Search costs \~$0.05 per task: \~$0.04 for search at $10 pe… Time per Task includes search time and model time. OpenAI Web Search has a Time per Task of about 31s. This ranks 14th of 26 products, mi…
AI High Signal

@scaling01 claims Opus 5.5 Max substantially outperforms Sol 6.1 Max, while Opus 5.5 Medium is approximately on par with Sol 6.1 Max. The post says Opus 5.5 Medium costs 1.8× more at API prices but is 3× faster per task; it also claims that on same-priced Claude and ChatGPT plans it completes about 3× as many tasks, citing roughly 5.6× more usage in Claude plans.

the real story is that Tibo is OpenAI's way too cocky secretary of propaganda Opus 5.5 max is >>> Sol 6.1 max and actually: Opus 5.5 medi…
AI High Signal

@hwchase17 endorsed @zeeg’s critique that current coding harnesses, including Codex, provide limited functionality compared with purpose-built alternatives . Zeeg predicted that local models would handle most daily tasks within five years, with remaining non-local uses becoming commoditized and vendors pressured to give customers more rather than impose increasing restrictions .

Good take on harnesses [https://x.com/zeeg/status/2107238795552977272](https://x.com/zeeg/status/2107238795552977272) You’re living in a bubble. You need to understand that a whole world lives beyond the constraints you have manufactured or imagined. For …
AI High Signal

Hark announced a launch for that week, offering a paid plan free to its first 100,000 sign-ups. A follow-up said early sign-ups would continue to receive the “$100 plans” for the first month until 3:00am PT that night.

Hark launches this week. The first 100,000 sign-ups get a paid plan for free I don't do anything without Hark anymore. The team has obses… Waitlist is on a roll. Going to keep giving early signups the $100 plans for the first month until 3:00am PT tonight. See you in the morn…
AI High Signal

Bloomberg reporting quoted in an X post says DeepSeek secured at least $12 billion in a funding round, potentially approaching RMB 100 billion, and is targeting an IPO in early 2027. A commentator welcomed the reported funding but said DeepSeek needs capacity in the high hundreds of megawatts and may still struggle to secure it.

BBG: DeepSeek has secured at least $12 billion in a funding round, with the final amount potentially approaching RMB 100 billion. BBG: De… NICE That's more like it, another $7.4B (with $3B from Liang) would have been embarrassing. They need high hundreds of megawatts of capac…
AI High Signal

A public provenance and attribution dispute alleges that ConwayResearch should add PrismML’s NOTICE file to its releases, identify Bonsai 2 and Qwen3.8 as base models, and credit inco_ai’s Splash inference engine behind underdogdotai’s speed figures. @vikhyatk characterized the allegation as taking model weights from one startup and an inference engine from another.

Hey [@ConwayResearch](https://x.com/ConwayResearch), the fair thing would have been: add [@PrismML](https://x.com/PrismML)’s NOTICE file … am i reading this right? they stole the model weights from one startup and the inference engine from another? lol [https://x.com/kanugula…
AI High Signal
  • Pratyush Maini said he would discuss two works at COLM on shaping model behavior during pretraining: The Finetuner’s Fallacy and Natively Unlearnable LLMs (NULLs). The Finetuner’s Fallacy argues that data encountered early in training leaves representational imprints that are difficult to undo; its thread also points to Rephrasing the Web, Safety Pretraining, and TOFU.
  • The NULLs work claims to enable on-demand deletion of individual training sources, addressing the problem that sources become entangled in shared model weights; its authors say they preserve individual deletability for millions of sources in a 1B-parameter model trained on web data.
I’ll be at COLM discussing two works on shaping model behavior during pretraining: - The Finetuner’s Fallacy w/ [@_christinabaek](https:/… If I had to compress my PhD into one idea, it is this "The data a model sees early in training leaves an imprint on its representations t… We are taking a big step towards scaling LLMs that can unlearn on demand. Cleanly deleting data from LLMs has proven impossible: training…
AI High Signal

Neel Nanda criticized companies that claim to support sensible regulation without actually supporting it, amplifying Alex Bores’s allegation that OpenAI committed perjury after representatives said under oath that they supported the RAISE Act.

If companies want to claim they support sensible regulation they should, you know, actually support the sensible regulation rather than l… I believe that [@OpenAI](https://x.com/OpenAI) just committed perjury. They were sworn in, under oath, and repeatedly said they supported…
AI High Signal

@vikhyatk claimed there is a GPU shortage and argued that governments should redistribute compute when a user’s MFU is 9%; this is a proposal, not a reported policy change.

there is a gpu shortage right now. if your mfu is 9% the government should intervene and redistribute your compute
AI High Signal

COLM’s first oral session was advertised for 6 October at 10 a.m. Pacific , featuring “CollabSkill: Evaluating Human-Agent Collaboration” and a paper on “Extracting memorized pieces of books” .

Catch the best of [@stanfordnlp](https://x.com/stanfordnlp) at the first oral session of [@COLM_conf](https://x.com/COLM_conf) tomorrow 6…
AI High Signal

@Algorithon claims that a Microsoft public webpage confirms OpenAI used Looped Transformers in the GPT-6 series and says GPT-6.1 Sol uses two inference passes “instead of three.” @teortaxesTex questions whether the change is intended to improve throughput and characterizes it as a possible new form of model “nerfing.”

🚨 BREAKING: Microsoft confirms on publicly accessible web page that OpenAI has been using Looped Transformers in the GPT-6 series, provin… lmao is that how they aim to improve throughput? new axis of nerfing has dropped [https://x.com/Algorithon/status/2107287650881208694](ht…
AI High Signal

The Turing Post’s weekly top-paper picks included Meta’s Context Language Models, Invent a Dataset, World Observer, The Planning Limits of Latent World Models, JEPA-TTT, and False Frontiers; the post names adaption_ai, KAIST AI, HRI and Johns Hopkins, and Rutgers and collaborators alongside several of these picks.

Must-read papers of the week Top: ▪️ Context Language Models by [@Meta](https://x.com/Meta) ▪️ Invent a Dataset by [@adaption_ai](https://x…
AI High Signal

A user reports that Devin’s voice-call feature starts a real-time agent that can pass tasks to the main session, allowing conversation while work proceeds; they found it useful for navigating large, busy sessions.

I've been heavily sleeping on the "phone call with Devin" feature If you didn't know, you should see a button to start a voice call with …
AI High Signal

Beam was described as handling up to 170,000 concurrent sandboxes and 1.3 billion total; a commentator says the total covered four weeks and estimates 3–5 minutes per sandbox, while cautioning that many may be short grading runs and DeepSeek may count those differently.

This is actually very interesting. Beam had "up to 170K concurrent sandboxes", and 1.3B total. This means about 3–5 minutes each, and man… 1.3 billion sandboxes over 4 weeks, huh. That's 48,46M sandboxes/day. At DeepSeek, one DSec scale unit does 3M, so would need 16 units, o…
AI High Signal

A post claimed HBM4 would need to return to 12-high stacking from 8-high because of HBM density, but another poster said BoA had not bought into the related “despec rumor.”

🚨🚨 HBM density Alert HBM4 will have to go back to 12Hi instead of 8Hi BoA never bought into this despec rumor. ![](https://pbs.twimg.com/media/HT6u3xBbsAAUv_Y.jpg) [https://x.com/damnang2/status/210697072161…
AI High Signal

A user reports that Dot’s first alert after setup was that their account had been hacked, which they say it recognized by proactively reviewing context.

My account got hacked earlier today, and I found out because I finally set up dot and this was the first thing it alerted me to. Very coo…
AI High Signal

A post promoting Tech Statecraft’s report says China’s access to ASML’s DUVi lithography machines is driving its AI chipmaking capability and argues for banning their export and servicing . Another post says, “Lithos is the only one lobbying for ASML here” .

Everyone should read [@techstatecraft](https://x.com/techstatecraft)'s fantastic report by [@nchlsbrwn](https://x.com/nchlsbrwn) on how C… Lithos is the only one lobbying for ASML here ![](https://pbs.twimg.com/media/HT6WoCfWsAAGWGK.jpg) [https://x.com/ChrisRMcGuire/status/21…
AI High Signal

@scaling01 says that after adjusting for task cost, Claude still performs 1.3–2.9× better in the comparison; at equal “intelligence” as measured by AAI, it can do more work faster. The post says this is smaller than the previously cited 5× difference.

even when adjusting for task cost you are still better off with Claude it's no longer a 5x difference but still 1.3x - 2.9x you can do mo…
AI High Signal

Teknium introduced a catalog of open-platform hardware and devices intended as things Hermes Agent or other agents can build on, integrate with, or run inside; a follow-up says it also includes DIY project ideas.

Just made this cool catalog with a ton of open platform hardware, devices and more that your Hermes Agent (or any agent really) can build… Lot's of DIY project ideas as well for each thing :)
AI High Signal

In initial pick-and-place tests on a DGX Station, Sentdex found DeepSeek V4.1 Flash better than GLM 5.3 Flash at reading imagery and making robotics decisions; the comparison is limited to the task he was testing.

Deepseek V4.1 Flash is seeming like a pretty good robotics control model. Happened to be testing DSV4.1F on DGX Station and it is so far …