ZeroNoise Logo zeronoise
Post
AI’s New Battleground Is the Agent Control Plane
4 min read
723 docs
Meta is extending Muse into a connector and developer platform as Jev, SIFT, and GAVEL show how cheap control loops, search, and explicit state tracking may matter as much as larger base models.

Top Stories

Why it matters: The frontier is moving above model weights—toward agents that can access tools, verify progress, and make cheap intermediate decisions.

Muse is expanding from an assistant into an agent platform. Meta lists Muse for Mac, Canada expansion on iOS and web, Granola and Notion connectors, and a developer platform. Levie’s thesis is that personal agents must complete work end to end through MCP/CLI, websites, and transactions; services optimized for agents rather than only human users will capture demand. The product question is therefore shifting from “which chatbot?” to whether services expose reliable, agent-usable paths to action.

Jev makes routine control decisions a product category. The Turing Post describes TypeSafe’s Jev as its first public “System One Model” and RLCD as the named training approach; its timing reflects agent workflows that repeatedly ask a large model whether to retrieve more context, call a tool, enforce a rule, or stop. Omar Sar reports using Jev to check whether an agent’s goal is complete after each turn, making frequent verification cheaper, but says the experiment is preliminary and still needs benchmarking. The counter-signal is important: NousResearch’s Teknium says a Jev compaction strategy simply removed tool calls, broke the cache, and raised input-token costs; he clarifies that the criticism targets that repository and strategy, not Jev’s valid use cases generally. The near-term test is whether specialized control models improve reliability, rather than merely moving failure modes into the harness.

Research & Innovation

Why it matters: New results suggest that search, verification, and explicit state management can produce large gains without changing the underlying model.

SIFT makes self-improving coding agents cheaper. A report on MIT and Sakana AI work says Self-Improvement via Fast Tree-search reached 35.1% on Polyglot with o3-mini after 30 expansions, versus DGM’s 30.7% after 80 nodes, using under 50 CPU-hours and under five hours of wall time. An LLM judge ranks candidate modifications before expensive benchmark evaluation; on TerminalBench, gpt-5.4-high improved a starting agent from 29.2% to 36.7%.

GAVEL shows the leverage of an external world model. The reported harness lifted Qwen3-8B from 41.2% to 91.8% on long-horizon robot tasks and from 19.9% to 92.6% on BEHAVIOR-1K across 500 multi-task instructions. It tracks object relations and action preconditions, repairs directly resolvable violations without another model call, and sends only semantically difficult errors back to the LLM.

Products & Launches

Why it matters: Releases are competing on long-horizon execution, modality, latency, and inference cost—not only peak benchmark scores.

StepFun released Step 5 Preview, a 600B-total/27B-active MoE with vision and a 1M-token context window. StepFun claims lower task cost, frontier-level performance in software engineering and professional knowledge work, and sustained execution over long horizons; it says open weights will arrive October 15.

Qwen launched Qwen3.8-LiveTranslate, an Interleave-based simultaneous-interpretation model covering 60 languages. Qwen reports average lagging falling from 2.8 to 2.3 seconds and adds speaker diarization with voice preservation, synchronized bilingual display, and long-context disambiguation.

Industry Moves

Why it matters: AI companies are organizing around browser access and independent evaluation as deployment moves into real workflows.

Meta is staffing browser use as a core capability. Shuyan Zhu says he left academia to work on Meta’s personal-superintelligence effort and is focused on making its models better at browser use; Edward Sun identifies him as the browser-use lead.

ValsAI is building an evaluation business around real work. It says models are advancing faster than legacy benchmarks and is developing independent evaluations designed to measure both capability and risk.

Policy & Regulation

Why it matters: Government AI organization is becoming more explicit, even before its mandate is clear.

Andrew Curran reports that President Trump announced a U.S. AI Force to oversee AI development and would announce an AI Czar in the near future.

Quick Takes

Why it matters: Research throughput, model economics, and security norms are all being reset at once.

  • Review capacity: Denny Zhou reports that ICLR 2027 received more submissions than all previous ICLR years combined.
  • Frontier inference: DL Weekly reports that DeepSeek shipped a 552-billion-parameter MoE scoring 74.2 on DeepSWE v1.1, narrowly ahead of Opus 5 at 74.0.
  • Disclosure repair: LiveOverflow says OpenAI’s CISO apologized after a public dispute over vulnerability handling, while noting that the critical thread was personal and outside the disclosure plan.
AI’s New Battleground Is the Agent Control Plane
AI High Signal

An AI-alignment-relevant debate: Tracewoodgrains argues that Matthew Adelstein’s moral framework—treating insects as potentially more important than humans and favoring the destruction of “net-negative” life—could make human extinction a relatively minor loss if a successor intelligence might create more positive value. The critic says this conflicts with human-focused existential-risk alignment and urges much greater confidence that human extinction is bad than in the speculative premises about insect experience. @teortaxesTex amplified the response, arguing that criticism should engage the framework’s core premises rather than peripheral objections.

Matthew Adelstein's argument against human value is morally monstrous and timed extraordinarily badly, and effective altruists should rej… And \*this\*, kids, is how you actually dunk on utilitarians. You need to know them to strike them. stuff about "eww bugs" or "aella weir…
AI High Signal
  • StepFun introduced Step 5 Preview as a flagship model for agentic work, claiming frontier-level performance in software engineering and professional knowledge work, with particular strength in finance.
  • Step 5 Preview is described as a 600B-total/27B-active MoE with 1M context and vision, sustained long-horizon software-engineering capabilities, and substantially lower task cost at comparable intelligence. StepFun plans to release its open weights on Oct 15.
  • Commentary speculates that an unreleased Step 4 may have preceded the launch and that Step 5 could be approximately GLM 5.3-level; conditionally, its architecture might support faster reinforcement learning than Kimi, ZAI, or Xiaomi, although the model needs to move beyond “Preview.”
Introducing Step 5 Preview: Advancing the Pareto Frontier. Step 5 Preview is our new flagship model for agentic work, delivering frontier… The most interesting thing about this release, to me, is the implication that there had been Step 4 which never made it. it's plausible t…
AI High Signal
  • AI development should prioritize empowering engineers and scientists, natural-language coding with verifiable and secure underlying code, broader access to knowledge, people with disabilities and health problems, artists, researchers, farmers, and public- and private-sector decision-makers.
  • Agents capable of dangerous tool use—including autonomous weapons or website hacking—should face legal scrutiny and accountability under sovereign law; the commentary identifies AI weaponization as a serious human risk and says safety and human dignity must not be compromised.
We should be developing AI as a tool (1) to empower engineers and scientists to solve the big engineering challenges we face such as ener…
AI High Signal
  • Open models accounted for 78.4% of token volume on Vercel AI Gateway versus 21.6% for closed models on a day described as potentially record-setting; Moonshot AI and DeepSeek ranked third and fourth in inference spend, and their combined spend surpassed OpenAI’s second-place total. This spend reflects model inference across providers, mostly in the US, rather than revenue paid directly to open-weight labs.
Looks like today may be a record day for token volume % of open models on Vercel AI Gateway: 🟦 Open 78.4% 🟨 Closed 21.6% While spend 💲 us…
AI High Signal

Google DeepMind Tokyo is led by Heiga Zen, who co-founded Google Brain’s Tokyo team with the post’s author in 2018; the author now runs Sakana AI and says Tokyo’s AI ecosystem has grown substantially.

日経新聞による、グーグルディープマインド東京を率いる全炳河(@heiga\_zen)氏の素晴らしい特集記事。 2018年、Heigaさんと2人でGoogle Brain東京チームを立ち上げた日々を懐かしく思います。当時は時差の厳しい深夜の会議をこなしながら、日本のAI研究の…
AI High Signal

Muse is gaining early traction as a broad personal assistant in an individual user report. The user says it optimized credit-card rewards across five cards and saved “no less than 1000 bucks,” handled restaurant reservations, email cleanup and follow-up updates, marketing-email unsubscribes, email drafting and sending, hotel planning, and laptop research and purchasing. After one week, the user called Muse their “new assistant” and said they were willing to give it access to mostly everything, including logins.

Getting a ton of utility out of Muse. This may take a minute to catch on because it’s truly a modality shift but here is what I used it f…
AI High Signal

Jev can be used for evaluations; LLM judges should be treated as classifiers, tested against human labels, and monitored to avoid overfitting.

Can you use Jev for Evals? Yes! Remember that a LLM Judge is also classifier\*. Make sure to test your classifiers against human labels a…
AI High Signal
  • An AI-security commentator warned that public cloud coding agents connected directly to private GitHub repositories create a concentrated attack surface; they singled out Codex Cloud and, conditionally, Claude, arguing that browser-side exploits, stolen credentials, phishing, or server compromise could let an attacker compromise the agent without first pivoting through the network.
  • The commentator also argued that ordinary SaaS tools and third-party libraries in the path of advanced AI work can become attack routes, citing a JFrog registry in a Hugging Face agent swarm as an example. They qualified the broader risk assessment: one serious OpenAI exploit does not demonstrate how to secure an organization of its complexity, while their loosely held view is that models already have serious offensive cyber capabilities and could produce a period of turbulence if offense outweighs defense.
I want to be very clear about this. just because we found one serious exploit on openai does not mean we know how to secure an organizati…
AI High Signal

A commentary shared by @sapinker argues that AI-extinction scenarios are preposterous and that treating them as inevitable can fuel fatalism, panic, and distraction from more mundane, realistic AI-safety challenges.

Is AI really going to kill us all? The scenarios are preposterous, and the presumption of inevitability encourages fatalism, panic, and d…
AI High Signal
  • An unverified rumor claims Anthropic trails OpenAI in training-compute intensity and inference economics. A follow-up argues the gap may be evident in Anthropic’s Haiku and Sonnet tiers, but questions whether it also applies to the higher-end Fable/Mythos models versus Astra and whether Anthropic accepts lower margins on Fable; the comparison remains unresolved.
According to rumors, Anthropic has a disadvantage to OpenAI in training compute intensity and inference economics. Interesting... ![](htt… huh It's obviously true for garbage tier models like Haiku and Sonnet, which Anthropic neglects. OpenAI cares a lot about the economy tie…
AI High Signal

The supplied posts express strong enthusiasm for Muse: one says its reception has been “beyond our biggest dreams” , while a quoted reaction calls it “the next ChatGPT moment we were waiting for” .

🥰 the muse reception has honestly been beyond our biggest dreams 🥰 [https://x.com/bidhan/status/2101408110942294411](https://x.com/bidhan… muse is the next chatgpt moment we were waiting for
AI High Signal

Trump has said AI could contribute “possibly as much as 25%” of the country’s GDP; a related analysis characterizes this as viewing AI as a powerful technology like the internet rather than expecting an imminent takeoff, a stance it says is shared in some ways by China’s leaders and supports continued development without a slowdown. A contrasting estimate attributed to Wenfeng puts AI’s impact at 10–20% of global GDP while predicting a singularity and broadly superhuman AI, linking AI’s value to its ability to multiply the wider economy.

Trump is not AGI-pilled. He boasts that AI could be “possibly as much as 25% of our Country’s GDP.” That’s all. In some ways, he shares t… Trump is correct. If anything he's a bit more bullish than Wenfeng who says 10-20% of world's GDP even as he predicts the singularity and…
AI High Signal
  • Mid-training appears to be a missing lever in open LLM development. The COLM-accepted MidTool work uses web, PDF, code, and synthetic tool data to shift a base model toward downstream distributions and teach general tool use; with identical SFT and RL, mid-trained models consistently performed better even though SFT loss barely changed.
  • Tool calling and effective use of tool outputs are distinct capabilities: text-only training improved tool-use success in a multimodal model but not final task performance, indicating that coding and deep-search agents need targeted mid-training data. Scaling the approach remains resource-intensive, and the authors argue mid-training should be co-designed with post-training.
Mid-Training Is the Missing Stage in Open LLM Research Frontier labs treat mid-training as essential, but most open research starts from …
AI High Signal

Hermes Desktop now supports pressing the left and right ⌘ keys together to attach a screenshot of the focused window directly to the active Hermes Desktop session.

You can now press left + right ⌘ together to attach a screenshot of the focused window straight to your active Hermes Desktop session [![…
AI High Signal
  • GAVEL, an external graph-world-model harness, reportedly improved Qwen3-8B performance on long-horizon robot tasks from 41.2% to 91.8% without changing the model itself.
  • The harness maintains object relations, action preconditions and effects, and probabilistic beliefs about unobserved object locations; it predicts action outcomes, catches violations, repairs directly resolvable errors without another LLM call, and routes only semantically difficult errors back to the model.
  • On BEHAVIOR-1K across 500 multi-task instructions, reported success increased from 19.9% to 92.6%; reasoning over possible object locations also reordered remaining subtasks and reduced travel distance by about 5.4%.
Impressive paper showing the impact of a good harness. Improves Qwen3-8B from 41.2% to 91.8% on long-horizon robot tasks without any chan…
AI High Signal

ICLR submission volume has surged: the source reports more than 60,000 submissions for ICLR 2027, compared with more than 500 for ICLR 2017, characterizing this as a greater-than-120x increase over 10 years.

10 years later, there are more than 60k submissions to ICLR 2027, a >120x increase 🤯 [https://x.com/ylecun/status/795046028143689728](… Looks like we have over 500 submissions to ICLR this year. [http://openreview.net/group?id=ICLR.cc/2017/conference](http://openreview.net…
AI High Signal

Muse received exceptionally strong praise: Alexandr Wang called it “as groundbreaking of a product as the iPhone,” while the original post said it exceeded expectations for what Meta could build and achieve and declared, “We are in the endgame now.”

“Muse, to me, is as groundbreaking of a product as the iPhone” [https://x.com/evrgn11112231/status/2101356115048972336](https://x.com/evr… They actually did it. Muse, to me, is as groundbreaking of a product as the iPhone and has exceeded my wildest expectations for what Meta…
AI High Signal

Muse’s Spotify connector can find and create playlists around users’ interests, update them daily on a recurring basis, and generate new podcasts that it adds to Spotify playlists. The example signals a broader shift toward consumer actions taking place directly inside LLM chats, with businesses exposed through connectors and reusable UI components.

The Spotify connector in Muse feels like the start of Cursor for music and podcasts. It can find create playlists for content you’re inte…
AI High Signal
  • @scaling01 predicts that sufficiently capable AI could reshape entertainment, envisioning a future where people can “watch and play anything.”
  • A linked post labeled “GPT-6 Astra + Unreal + Blender” describes porting an existing Godot concept to Unreal and modifying its trees and lighting.
one of the things I was always most excited about is how the entertainment industry would change with strong AI you can watch and play an… GPT-6 Astra + Unreal + Blender I got a new MacBook, so I decided to try switching from Godot to Unreal. I moved the concept I had already…
AI High Signal

An anecdotal AI reliability signal: the author reports that DeepSeek in DSH failed a very local search for a specific antidepressant in Argentina, replying that it was not imported, and continued to fail after being prompted to check local brand names; Astra reportedly failed the same test. The author contrasts this with strong—but unquantified—HLE performance.

Did a very local search (a specific antidepressant in Argentina) and DeepSeek in DSH totally flopped, "neeh it's not imported". I told it…