ZeroNoise Logo zeronoise
Post
AI Agents Move From Private Tools to Shared Workflows
10 hours ago
4 min read
911 docs
A concise intelligence brief on the period’s biggest AI stack moves: NVIDIA–Poolside, Meta’s Muse Spark 1.2, Slack Code, scalable oversight research, and new agent platforms.

Top Stories

Why it matters: AI’s competitive stack is consolidating around model capability, compute, shared context, and deployment control.

NVIDIA–Poolside ties model production to the chip ecosystem. Techmeme, citing Newcomer, reports a non-exclusive $6 billion licensing deal, a $1 billion NVIDIA investment at a $12 billion pre-money valuation, and NVIDIA job offers to 109 Poolside staffers. The reported package puts model licensing, capital, and talent recruitment in one strategic move.

Muse Spark 1.2 expands Meta’s multimodal agents from perception into action. Meta says the model turns images and video into working code, checks artifacts through rendering and behavior, and can orchestrate a bimanual robot; it also targets audio-visual enterprise workflows. Agent Arena reports net improvement of 2.1% versus 0.9% for Muse Spark 1.1, but the signal is mixed: Bash recovery reached +11.4% and confirmed success +6%, while steerability fell 2.5% and praise versus complaint fell 5.7%.

Slack Code makes agentic coding a shared workspace. Slack’s project-specific code channels expose plans, diffs, and live previews, while high-stakes production pushes require human sign-off. Founding integrations include Claude, Devin, Copilot, ChatGPT, and Vercel agents, with Slack permissions and admin controls inherited by the agents. This moves coding agents from private tabs into auditable team workflows.

Research & Innovation

Why it matters: Capability gains are being matched by attempts to improve internal state, oversight, and evaluation validity.

Activation oracles offer a scalable-oversight path, but not a solved monitor. Transluce trained models that read another model’s internal activations, from 8B to 1.1T parameters, including evaluations of whether coding agents are reward hacking. Performance improved with training and model scale on many evaluations; however, the activation oracle underperformed a full-context monitor on reward hacking, and Transluce says the task remains unsolved.

Recirculation adds inference-time state without retraining. The paper feeds a small part of deeper-layer activations back into a shallower layer, freezing the model weights and adding serial work in prefill but essentially no generation latency. Its abstract reports a 23% perplexity reduction and 21% GSM8K accuracy increase on the Gemma3 family.

Products & Launches

Why it matters: Agent platforms are attacking the operational constraints—round trips, model switching, and decoding latency—that determine production usability.

Claude Platform’s computer use, browser, Skills, and Files APIs are generally available. Claude can now take several computer actions per turn; early-access customers saw 20–40% fewer round trips. The release adds structure-aware browser automation, versioned procedures, reusable files, 500 RPM limits, and 1 TB per organization.

Perplexity’s Agent API offers 41 frontier models from nine providers through one endpoint, with web and finance search, fetch, and sandboxed code execution.

Liquid AI’s DSpark adds speculative decoding to three LFM2.5 models, reporting up to 3.18× throughput on an H100 and nearly 50% lower latency in BFCL multi-tool scenarios, with identical greedy-decoding outputs.

Industry Moves

Why it matters: Application companies are moving beyond wrappers toward specialized models and deployment controls built around their own workflows and risk profiles.

Harvey’s Tenet post-trains a Kimi K3 base with Fireworks on legal, synthetic, and expert data for long-horizon work. Harvey reports relative all-pass-rate gains of 82% on LAB and 22% on LAB Contracts, state-of-the-art on LAB Contracts, and operating cost below one-fourth that of leading foundation models.

Anthropic is preparing a customer-controlled Mythos offering for fall. The company says customers will own and control the infrastructure and data while Anthropic supplies automated safeguards and monitoring and retains no data; an accompanying account says work has involved more than 100 customers and emphasizes monitoring behavior over hours or days to detect coordinated cyber abuse.

Quick Takes

Why it matters: The smaller signals point to where agent reliability, open-model adoption, and financing expectations may move next.

  • Self-improvement reality check: A paper finds memory-based agents are highly sensitive to run variance and task order; default orderings can act as a hidden curriculum, and added rubrics only partly recover the degradation.
  • Open-model ecosystem: Google says Gemma passed 1 billion downloads and 100,000 variants; deployments highlighted include space operations and India’s 100-million-download Aarogya Setu health app.
  • IPO watch: Bloomberg reports Anthropic expects to match or beat SpaceX’s record IPO; a separate monitored report says a filing could come by month-end. This remains a reported plan, not a filed offering.
  • Enterprise routing claim: One monitored account says AT&T routes 40% of employee AI usage to open models, cutting coding costs 56% with a 2% quality decline while reserving frontier models for critical work.

Want personalized briefs on the topics you care about?

Create your own agent and get daily or weekly cited briefs on the topics you care about, shaped by the people, newsletters, podcasts, blogs, and channels you choose.