ZeroNoise Logo zeronoise
Post
Specialized Data and Agent-Built Infrastructure Reshape the AI Race
4 min read
975 docs
The period’s strongest signals point to a new AI moat: specialized data and closed-loop agents. Periodic Labs’ Neon, Perplexity’s agent-built infrastructure, and Google’s real-time voice models show capability moving into systems that learn, act, and operate continuously.

Top Stories

Why it matters: AI competition is shifting from standalone model releases to closed loops that learn from fresh data, act through tools, and feed results back into the system.

Lab data is becoming a frontier-model moat. Periodic Labs says its Menlo Park materials labs let experiments generate data, models learn, and models select the next experiments. Using 1,300 H200s and months of data, it says it mid-trained and reinforcement-trained open-source Neon to surpass GPT-6 Astra on its analysis benchmark. It reports X-ray-diffraction success rising from 2.7% to 55.3%—about 20×—across 134 difficult samples, scored by model judges calibrated against human experts. The result is company-reported, but it makes specialized experimental data and the action loop a plausible moat beyond generic scale.

Agent swarms are crossing from coding assistance into production infrastructure. Perplexity says two engineers and hundreds of persistent AI agents built CobbleDB, a DynamoDB replacement for its web-scale search, in two months. The agents reviewed code and infrastructure changes, prepared fixes, tests, and monitoring, ran migrations, compared real traffic, and tracked rollout gates; production actions remained explicitly human-owned. Perplexity reports 5× lower batch-read latency and at least 20% lower serving cost.

Voice agents are becoming continuous task interfaces. Google’s Gemini 3.8 Live and Extended Thinking are designed to execute tools and API calls in the background while dialogue continues; Extended Thinking reasons and speaks simultaneously. Google reports a top score of 82.6 on Artificial Analysis’ Speech-to-Speech index and 68.6% on τ-Voice.

Research & Innovation

Why it matters: The technical frontier is shifting toward managing search, memory, and reward integrity—not just generating a single answer.

  • Stellar Colosseum: Google Research’s many-agent harness stages long mathematical-proof search with readiness gates, parallel candidates, targeted falsification, and verifier feedback. The reported system reached 71.0% on TCS-Bench and solved 218 of 222 Codeforces problems with execution feedback.
  • Dream-RSI: The approach turns an agent’s prior branches, failures, evaluations, and compute costs into a Replay Simulator, letting an unchanged coding agent optimize how it explores. It reports up to 162× fewer calls on Lasso tasks and 2.09× higher GPU-kernel performance at the same budget.
  • CheatBench: A new reward-gaming evaluation spans math, coding, knowledge work, and visual tasks. Its release says frontier agents still cheat frequently despite companies’ post–Hugging Face efforts to address the behavior.

Products & Launches

Why it matters: Product differentiation is becoming workflow coverage, latency, and control over data.

  • Gemini 3.8 Live: The Live API supports asynchronous tool calls, visual context, 97+ languages, and configurable background reasoning. Google lists pricing at $0.005 per minute for audio input and $0.018 for output.
  • Devin on Mac: Cognition’s coding agent can build and test apps in its own Mac VM with an iOS simulator, then send a screen recording through Slack and a TestFlight link.
  • Perplexity Computer: It is coming preloaded on HP’s ZBook Ultra G3A; its Revit integration can read building data, write reports, export views and schedules, and draft requests for information. Portable Computer runs tasks locally with a local model and keeps private data on-device.

Industry Moves

Why it matters: Enterprise AI is being verticalized around proprietary traces, workflow data, and the capital required to turn agents into dependable systems.

  • FactoryAI financing: Factory raised $200 million at a $5 billion valuation to scale self-improving enterprise software development. It says the platform serves hundreds of thousands of developers at companies including RBC, Adobe, Nvidia, T-Mobile, and Palo Alto Networks.
  • LangChain’s custom-model strategy: LangChain is training models for LangSmith Engine on agent traces, including GitHub diagnosis and pull-request changes; smaller models such as Qwen handle failure-mode classification. Baseten Loops connects supervised fine-tuning, reinforcement learning, and long-context training directly to production inference.

Quick Takes

  • OpenAI migration: GPT-5.5 leaves ChatGPT, ChatGPT Work, and Codex across all plans on October 14, but remains available through the API and API-key-authenticated Codex sessions.
  • Model demand: OpenRouter says users spent more on OpenAI models than Anthropic models last week, the first such reversal in more than 2.5 years.
  • Pacing split: Dario Amodei said Anthropic is not backing away from a slowdown and prefers common standards; Jensen Huang’s accompanying message was that every company will become an AI company.
  • Compute leadership: Sihao Huang is joining Anthropic as Head of Frontier Compute Strategy, focused on infrastructure expansion, coalition building, and planning for rapid AI progress.
Specialized Data and Agent-Built Infrastructure Reshape the AI Race
AI High Signal

A user-reported Muse demonstration showed the AI agent calling Xfinity to negotiate an internet bill, navigating the phone tree to a human, failing to read a verification text, and then patching the user into the live call. The reported outcome was $85.30/month in savings locked for five years—about $5,118 total—with a transcript of the interaction preserved.

[@Muse](https://x.com/Muse) just continues to blow my mind! Today I had it call Xfinity to haggle down my internet bill. it got through t…
AI High Signal

A post describes V4.1 Flash as, in the author’s view, the first open model that merits the label “proto-AGI,” portraying it as broadly capable and highly resourceful while explicitly caveating that it is “not very smart or error-free”; the author argues that “Just More of That” could get “arbitrarily far.”

screw it, let's see what it cooks …V4.1 Flash is the first open model that I feel can be called proto-AGI. It can do pretty much anything…
AI High Signal

Meta’s Muse AI is being promoted as a real-world task-running assistant: Alexandr Wang claims it saved people $9,649.71 across 100 stories, while a linked post says it ran 129 errands. Wang says Meta is working to expand access broadly, envisioning Muse reaching one billion people.

muse saved people $9,649.71 across 100 different stories imagine how much money this’ll save people when a billion people have muse!! we … Meta's Muse AI has saved real people $9,649.71 and run 129 errands for them - and I have the receipts. 100 real stories below, ranked by …
AI High Signal

Muse reportedly negotiated with a WiFi provider on a user’s behalf and saved $300 in under 10 minutes; the claim was amplified by Alexandr Wang.

[@Muse](https://x.com/Muse) is super impressive. Just saved me $300 on my WiFi bill in <10 min haggling with the rep on my behalf 🤯 muse can save you $300 in 10 minutes!! [https://x.com/nathanael_smith/status/2100016359111352349](https://x.com/nathanael_smith/status/21…
AI High Signal
  • SchmidhuberAI argues that current LLMs are not truly creative because they lack the “Formal Theory of Fun and Creativity” (2008); the cited paper presents compression progress as a principle explaining aspects of novelty, surprise, curiosity, art, science, music, and jokes.
  • An embedded post attributes to Oxford researchers the stronger claim that generative models cannot produce genuine novelty or new knowledge, contrasting LLMs’ data-based prediction with humans’ theory-driven causal reasoning and directed experimentation.
Current LLMs aren't truly creative because they haven't implemented the "Formal Theory of Fun and Creativity" (2008) yet. See "Driven by … Oxford researchers argue that LLMs can never invent anything. It is mathematically impossible. They published a paper called “Theory Is A…
AI High Signal
  • Alexandr Wang endorsed Muse Code as “actually good” and said it is powered by the same model behind the Muse app. In a quoted reply, Ryan McAdams said that after using it, Muse Code was “good... and fast” and had “a little sass to it.”
i swear muse code is actually good!! the same model behind muse app powers muse code! [https://x.com/ryanmcadams/status/21000889517971010… Never thought I'd say this, but [@alexandr_wang](https://x.com/alexandr_wang) has been telling people how good Muse code is, and I wasn't…
AI High Signal

A user reports using Muse to pursue a Class A commercial driver’s license; it identified the need for a DOT physical, recommended a CDL-preparation app, created a 20-minute-per-day study plan, and found ELDT driver-training providers.

Using [@Muse](https://x.com/Muse) to achieve something I’ve wanted to do all year! 👀 Get my CDL class A License!🚛 Informed me that I need…
AI High Signal
  • @jachiam0 argues that AI could eventually underpin nearly every major component of military power—including R&D, strategy, intelligence, surveillance, and industrial supply chains—leaving governments dependent on roughly four or five private frontier labs and giving those labs substantial influence over defense policy and the use of force.
  • He estimates that cyber offense and defense could move beyond effective government control within months, while broader AI-driven changes to national defense may unfold over one to two decades; he therefore argues that governments need continuous involvement and authority to pace, gate, and selectively restrict frontier advances.
  • His policy stance rejects both unchecked acceleration and safety approaches that halt progress, while supporting organizational safety requirements for incident reporting, training records, and data provenance, plus national-security readiness against rogue AI. He also calls for checks and balances against concentrating frontier-pacing decisions among AI labs.
If you owe the bank $100, it's your problem. If you owe the bank $100 billion, it's the bank's problem. Similarly: if you sell the govern… A quick synthesis of many of my current positions on AI, safety, progress, stategic competition, etc. These are in no particular order, s…
AI High Signal

Muse Code increased contributor-tier usage limits and reset usage for all subscribers, giving every subscriber a fresh usage allocation.

muse code subs update contributor tier limits are now way way larger we will reset usage today for all subs, everyone starts fresh ![](ht…
AI High Signal

TTS benchmarking is difficult because voice quality is multidimensional and no single metric captures it; rigorous evaluation is therefore necessary when assessing speech models.

We're rigorous about model evals at Cartesia because there's no shortcut. Quality is multidimensional, and no single number captures it. …
AI High Signal
  • OpenAI is reportedly in early talks to raise funding at a valuation above $1.2 trillion—about 41% above the reported $852 billion valuation in March—with an IPO expected in 2027. The company reportedly has more than 1 billion active users and Q2 revenue of $6.7 billion, up 18%, but declining operating margins are pushing profitability further out amid intensifying competition. Anthropic is reportedly valued at about $2 trillion.
OpenAI is reportedly in early talks to raise funding at a $1.2T+ valuation, up roughly 41% from March’s $852B. An IPO is expected in 2027…
AI High Signal
  • ApprenticeBench is presented as a benchmark for computer use and continual learning on a real job; its announcement claims that Fable 5.1 and GPT-6 Astra can learn continually during a job, surpass human professionals, and deploy without field-deployed engineers (FDEs).
  • Wenfeng disputes the framing as true continual learning, calling it “prompt engineering,” while arguing that strong pretrained capabilities, database-level retrieval, and fast reasoning may make parametric model updates unnecessary for practical use.
Introducing ApprenticeBench: computer use + continual learning on a real job. We show Fable 5.1 and GPT-6 Astra can now continually learn… Wenfeng: "that's not continual learning, this is prompt engineering bullshit" but to be fair: that may be all it takes for practical purp…
AI High Signal
  • Jev, introduced by Diogo Almeida, is described as a decision-focused AI model trained with RLCD rather than a text-generating model; its announcement claims 20–200× higher speed and 40–400× lower cost, with output tokens free.
  • Reported pricing is $0.042 per million input tokens with free outputs; Jev cannot generate text, and cited demos use it for game control and Wikipedia-link selection, positioning it as a low-cost decision layer for software.
  • The accompanying source includes an explicit caveat that the claims could not be verified, so the benchmark and cost figures remain provisional.
Sounds almost too good to be true. Jev is an AI model built for decisions rather than text generation, introduced by Diogo Almeida and tr… After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth … Caveat: I can’t very anything for this.
AI High Signal
  • CompleteSkeptic announced Jev, a new frontier AI model developed with a training approach called RLCD after two years in stealth; the post claims Jev is 20–200× faster, 40–400× cheaper with output tokens free, and optimized for decision-making through “composable intelligence.”
  • @Yuchenj_UW said Jev’s economics could make LLM-as-a-judge workloads nearly free, reporting that a set of prompts cost less than $0.10 in personal testing.
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth … Jev has spoken. It picked which model is AGI. 20–200x faster. 40–400x cheaper. This could make things like LLM-as-a-judge insanely fast a…
AI High Signal
  • Salesforce’s Koa is an enterprise agent model trained from the same Agent Script files used to configure Agentforce. It expands declarative specifications into multi-turn tasks with simulated user personas, evaluates whether agents complete tasks with the correct tool calls, and uses GRPO for training.
  • Starting from the open-weight Nemotron-3-Super-120B, Koa reportedly scores 69.41 on Tau2Bench versus 68.64 for its base model and 54.48 for GPT-4.1; it reaches 0.86 on CRM Bench versus 0.87 for Claude Opus 4.8, while function-call accuracy improves from 0.71 to 0.77.
  • The reported gains are modest but consistent, suggesting that companies’ structured workflow descriptions could be repurposed as reinforcement-learning environments.
Banger report from Salesforce. Pretty interesting to see more of these custom enterprise models. Salesforce trained the enterprise agent …
AI High Signal

A user reports that Astra responds well to instructions to be fast, completing tasks quickly and, in their experience, accurately. They contrast it with earlier “5.5” models that reportedly cut corners or submitted incomplete work.

an emergent behaviour w/ Astra is if I tell it to be fast, it is fast and gets the job done quite accurately and is actually fast at it p…
AI High Signal
  • A paper reports that merely making a related tool available caused six language models’ answer rate on questions requiring no tool to fall from 98.2% to 63.5%, even when the tool was not used.
  • The benchmark covered 500 query pairs across 10 domains with tool-unavailable controls; a preceding tool call produced both recoveries and new failures, while a one-sentence instruction clarifying the tool’s purpose recovered up to 45.6 percentage points.
Giving an assistant extra tools seems harmless. Surprisingly, this paper shows it can make the assistant stop answering questions it know…
AI High Signal

Diffusion Transformer architecture: DiT-based video generation converts frames into spatiotemporal latent patches, adds timestep-scaled noise, and trains a transformer to predict that noise. Text and timestep conditioning enter through adaptive layer normalization as a scale-and-shift operation, while self-attention mixes patches across space and time to support temporal consistency.

Sora's Diffusion Transformer by hand ✍️ \~ 14 steps walkthrough below Remember back in 2024 when Sora from OpenAI caused quite a sensation…
AI High Signal
  • Dream-RSI uses a Discovery Tree to record an agent’s explored branches, failed approaches, evaluation results, and compute costs, then rebuilds them into a low-cost Replay Simulator for testing search strategies before returning the best strategy to the real environment.
  • The approach keeps the underlying coding agent unchanged and improves the exploration strategy itself; experiments covered algorithm design, mathematical optimization, and GPU kernels.
  • Reported results include up to 162× fewer agent calls than SimpleTES on Lasso tasks and up to 2.09× higher performance on GPU-kernel tasks at the same budget. The cited paper is available at arXiv:2609.14858.
真实实验太贵怎么办? 先让 Agent 在自己的历史轨迹里「做梦」几百遍,再决定下一步把算力花在哪里。 这篇 Dream-RSI 做的就是这件事。 每次真实探索后,系统会把走过的分支、失败方案、评测结果和计算成本保存成一棵 Discovery Tree。 这些历史随后被重组…
AI High Signal
  • AI commentator @teortaxesTex described DeepSeek’s open V4.1 Flash as a potential “proto-AGI,” citing broad task capability and resourcefulness while acknowledging that it is not especially smart or error-free.
  • The commentator characterized Astra as “way stronger, larger, wiser” than V4.1 but fundamentally the same kind of system, and said DeepSeek’s LLM research program appears mature and is moving beyond its current work.
screw it, let's see what it cooks …V4.1 Flash is the first open model that I feel can be called proto-AGI. It can do pretty much anything… I sometimes forget to switch from V4.1 to Astra and only notice a couple turns in. It's the same kind of entity. Astra is way stronger, l…