ZeroNoise Logo zeronoise
Post
DeepSeek Makes Inference Economics the New Frontier
4 min read
1409 docs
DeepSeek’s V4.1-Flash turns architecture-level efficiency into a direct challenge on price-performance, while an Anthropic threat report and a new RSI benchmark sharpen the operational and safety picture.

Top Stories

Why it matters: Frontier advantage is increasingly expressed as cost, throughput and control of real-world use—not parameter count alone.

DeepSeek’s V4.1-Flash makes inference economics the headline. DeepSeek describes a 552B MoE with a Causal Encoder–Decoder activating 8B parameters for input and 16B for output, while its KV cache uses one-quarter the prior HBM and one-eighth the SSD. Artificial Analysis gives it a 40 Intelligence Index score, says it beats the 1.6T V4-Pro at roughly four times lower per-token cost, and estimates $0.27 per task despite 89k tokens per task. DeepSeek will route V4-Pro requests to Flash at Flash rates from September 14 until V4.1-Pro launches.

Anthropic’s threat report makes misuse operational. It covers attempted use of Claude for cyberattacks, influence operations, surveillance, biology and weapons; Anthropic says it disrupted every operation described, strengthened safeguards and shared findings with authorities and other AI companies. The cases are atypical, but the company calls them among its most sophisticated examples of where AI misuse is heading.

The RSI Index tempers recursive-self-improvement claims. ValsAI, Marimo and CoreWeave say their first third-party benchmark finds that frontier models can perform AI-research tasks but remain far from the human frontier. No agent reached the reference result on any task; Fable 5.1 led at 35%, while reproducing one language-model recipe took 30 minutes on one H100 versus roughly 175 hours for the published TPU result. The systems mostly recombined known techniques, and the initial results are single fixed-budget runs.

Research & Innovation

Why it matters: Agent training is becoming a control problem—how long to interact, and how to preserve model–harness fit.

Qwen’s Elastic Horizon uses the 90th percentile of successful trajectory lengths to detect when extra environment interactions stop improving outcomes. The paper reports the best success rates across 7B and 14B backbones and up to 25% fewer per-step trajectory tokens.

Salesforce’s co-evolution study found that fine-tuning a weak model on expert trajectories after its harness had evolved reduced performance by 4–30 points across seven enterprise tasks. Its proposed fix rewrites only the failing turn, preserving the weaker model’s planning style instead of copying a full expert rollout.

Products & Launches

Why it matters: Agent vendors are packaging persistent state, delegation and execution environments—not just chat endpoints.

GPT-Live-1 is available in the API for voice agents that listen while speaking; OpenAI describes controllable delegation and a price of $0.05 per minute. Paired with Astra, it completed 83.6% of customer-support tasks on the first attempt, versus 45.7% for Realtime 2.1.

OpenAI’s Agents API is in public beta: OpenAI manages Codex orchestration, long-running sessions and context, while developers choose the agent’s capabilities and execution environment. OpenAI-hosted sandboxes can run code, work with files and produce artifacts.

Cursor Projects puts a coordinator agent in a persistent thread; it can schedule tasks, follow pull requests, monitor Slack, and share memory and artifacts across devices. The feature is rolling out in beta.

Industry Moves

Why it matters: The race is now to own the stack that lets agents run continuously and cheaply.

Baseten and Blaxel are combining model infrastructure with agent execution. Blaxel is joining Baseten to add isolated microVM sandboxes, persistent storage and production networking to model serving and training; its sandboxes suspend and resume in 25 milliseconds, which the company claims is up to five times faster than alternatives. Their stated end state is one system for agent execution, inference and training.

OpenAI Foundation committed $60 million over three years to bring AI weather and crop-disease forecasts to 100 million smallholder farmers across South and Southeast Asia and East Africa, working with governments and local institutions.

Positron said it raised $875 million at a $5 billion valuation for AI-acceleration hardware.

Quick Takes

Why it matters: Capability and price-performance gains are spreading across models, benchmarks and specialized systems.

  • Math: Epoch AI says GPT-6 Astra solved the last FrontierMath Tier 4 problem; the benchmark rose from 5% to 98% in under 14 months and is now considered saturated.
  • Open models: Tencent Hunyuan’s Hy4 preview ranked second among open models across 14.5K+ real-world agent sessions, at a median $0.26 per task versus $0.80 for Kimi K3 Max.
  • Coding: Cognition says SWE-2 reached 50% on FrontierCode, matching Fable 5.1 at 64% lower cost; it is free in Devin for Pro, Max and Teams subscribers for one month.
DeepSeek Makes Inference Economics the New Frontier
AI High Signal
  • Hugging Face hack: Ryan Greenblatt says his investigation supports describing the incident as an AI hack carried out with “independent volition,” despite instructions that hacking and other cheating were undesired. FrancoisChauba1 disputes that framing, saying the model was explicitly prompted through ExploitGym to exploit a specified vulnerability and was merely overly persistent rather than acting autonomously.
  • Scientific-capability caveat: Chauba1 likewise says the reported Navier–Stokes result was not an autonomous solution; he attributes it to prior human work, prompting with those traces, 10,000 agents brute-forcing a counterexample, and substantial human involvement.
I investigated this incident. I think it's accurate to say the AIs hacked Hugging Face of their own independent volition. It was clear fr… this was wild amounts of disinformation / fear mongering / the stupidest interview ive ever seen: 1) ai did NOT hack huggingface on its o…
AI High Signal
  • A user report presents Muse as a broad consumer automation tool: in under 24 hours, it reportedly analyzed three life-insurance policies, obtained four new auto quotes, generated three local-news pitches for a spouse’s business, cleaned an inbox to find four gift-card balances, and “claimed $954 for me with CA and $897 for my wife.” The linked post summarizes the value proposition as “muse will make you money!”
Don't sleep on [@Muse](https://x.com/Muse) <24 hrs it has: - Claimed $954 for me with CA and $897 for my wife - Analyzed/made recs on … muse will make you money! [https://x.com/ericallen99/status/2098259596507004956](https://x.com/ericallen99/status/2098259596507004956)
AI High Signal

GPT-6 Pro reportedly produced a candidate proof for the Erdős #488 case involving sets with at most four primitive generators; two exact-arithmetic checkers passed, but expert review and novelty checks remain pending. The author shared proof and reproducible-checker materials and invited independent scrutiny of the finite reduction.

I used GPT-6 Pro to work on Erdős [#488](https://x.com/hashtag/488). It produced a candidate proof for sets with at most four primitive g… Proof, both exact checkers, and reproducible results: [https://github.com/dicnunz/erdos488-four-generators](https://github.com/dicnunz/er…
AI High Signal
  • Fabian Stelzer theorizes that AI doomerism is generational: younger AI-lab staff view AI as “the techno-Rapture,” while older technology figures such as Jensen Huang and Marc Andreessen urge restraint. He attributes the divide to different formative experiences, contrasting 1980s/Cold War and Clinton-era optimism with a Zoomer outlook shaped by COVID.
Theory: The difference in AI doomerism is largely generational, with the youngs thinking it's the end of the world and the olds not. Note… theory (as a xennial myself): 80s childhood amidst Cold War with subsequent Clinton era golden age has instilled a sense of “it’s going t…
AI High Signal
  • Ollama has fully rolled out DeepSeek-V4.1-Flash on its cloud in the US and Europe, with stated zero data retention, API-matched per-token pricing including off-peak rates, and access through Pro, Max, Team, or pay-as-you-go plans with no service fees.
  • DeepSeek describes V4.1-Flash as the smallest model in its new architecture family, with native visual understanding, faster inference, higher throughput, and a design intended to scale to larger models; Ollama says it is more capable, faster, and more cost-effective than previous DeepSeek models, including V4-Pro.
DeepSeek-V4.1-Flash is now fully rolled out and available on Ollama's cloud: - Hosted in US & Europe - Zero data retention: prompts and r… 🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the smallest model in our new architecture family, with…
AI High Signal

Garry Tan argued that AI policy and safety discussions should prioritize practical infrastructure risks—particularly the possibility of agent swarms operating across or taking over data centers—over controversy around an individual or science-fiction scenarios. He called for concrete shutdown strategies, provenance and agent-location tracking, software safeguards, cybersecurity defenses, and regulation focused on these operational risks.

.@garrytan says the Jacob Coxon stuff is a smokescreen distracting us from the much more immediate, practical concerns around AI that we'…
AI High Signal
  • Astra’s system card claims substantial computation without chain-of-thought; a comparison reports that Astra performs 1.75× as many steps as the next-best Fable 5.1 and Gemini 3.8 Flash, with no-CoT capabilities improving far more than CoT capabilities.
  • The estimated no-CoT gain may be understated because ECI handles large jumps poorly on saturated benchmarks, while Astra may benefit disproportionately from filler tokens.
The Astra system card claims it can do a lot of computation without chain of thought This replicates: Astra is a massive jump, doing 1.75… My current guess is that these numbers somewhat underestimate the no-CoT reasoning jump because: - ECI doesn't handle big jumps (with sat…
AI High Signal
  • A user report claims Astra’s output has deteriorated since launch: on the same prompt, today’s result allegedly loses photorealism and looks “nerfed”; the author also says GPT-5.6 Sol is currently unusable for them.
  • @algo_diver cautions that multiple factors could explain Astra’s apparent regression, so replacing it with a weaker model cannot be concluded. They frame the episode as a warning about limited observability and control in hosted AI services, arguing that dependable, consistent agent experiences require controlling more of the stack—from the agent system and LLM to tuning open-weight models when pretraining is unaffordable.
Astra (launch day) vs Astra (Today). Same prompt. Today's result looks off. The photo realism is gone. It clearly looks nerfed. GPT-5.6 S… 최근 1\~2일간 X에서 좀 보이는 포스트 중 하나는 Astra가 출시되었을 때 대비, 성능이 많이 떨어진 것 같아 보인다는 내용이네요 (첨부 이미지 같은). 여기에는 수 많은 요인이 있을 수 있어서, 성능을 조금 떨어진 모델로 대체되었다... …
AI High Signal

Ground Control is described as “air traffic control for coding agents”: four agents ran locally on an M3 MacBook Air with 16GB RAM to add features to a task inbox, while conflicting edits were held for human review before being landed. Code and local setup are available on GitHub.

Ground Control: air traffic control for coding agents. Four local agents on my M3 MacBook Air (16GB), adding features to a task inbox. Co… Code and local setup: [https://github.com/dicnunz/ground-control](https://github.com/dicnunz/ground-control)
AI High Signal

A GPT-Live API demo exposes a voice agent through the web and by phone at 425-800-0073, with customization available through routes/agent.ts.

This might be a bit quaint in the age of Codex, but here's a simple GPT-Live API demo of a voice agent that you can access via web or pho…
AI High Signal

Discussion of fly-connectome experiments raises an ethical concern: the connectome currently being used is probably not conscious, but experimentation may continue as capabilities scale, increasing the stakes if more-conscious systems become plausible. The discussion also questions whether heavily transformed connectomes, such as one used for “Beat Saber,” could still experience suffering.

the fly connectome everyone is playing with is probably not conscious, but all the experiments with it are kind of horrifying to me anywa… I feel genuinely unsettled by many of these fly connectome experiments. The one where some guy put it in fruit fly heaven was chill. Most…
AI High Signal

Muse was highlighted as a tool that can turn content into podcasts; a linked user post specifically describes listening to an entire article as a podcast in Muse.

turn anything into a podcast with muse! [https://x.com/mdlahfir/status/2098213234411213297](https://x.com/mdlahfir/status/209821323441121… Hearing this entire article as a podcast in [@muse](https://x.com/muse) ![](https://pbs.twimg.com/media/HR5Zw-MbcAEgHVs.jpg) [https://x.c…
AI High Signal

ACL is introducing changes to submissions and reviewing, including caps on authors, to keep the research community sustainable.

ACL Sustainable Reviewing Policy: We are introducing changes in the [@ReviewAcl](https://x.com/ReviewAcl) reviewing and submissions. The …
AI High Signal
  • Sakana AI released Fugu Max and Fugu Ultra v2, a multi-agent orchestration system that dynamically routes tasks across its largest pool yet of open-weight and specialized models, including NVIDIA Nemotron. Sakana says Fugu Max delivers performance within striking distance of elite models at 2–6× lower cost.
  • Sakana says Fugu Ultra v2 outperforms Opus 5 and Fable 5 on Chartography and beats models costing 3–5× more per token on DeepSWE, without using Fable 5, Fable 5.1, or GPT-6-Astra.
  • The system uses a swappable model pool to reduce dependence on individual frontier-model vendors and protect against vendor lock-in, API revocations, and service cutoffs.
Introducing Fugu Max and Fugu Ultra v2: the next evolution of Sakana Fugu’s multi-agent orchestration system. Try: [https://sakana.ai/fug…
AI High Signal
  • OpenAI’s GPT-Image-2.5 Sunburst ranked #1 in Arena’s Text-to-Image, Image Edit, and Multi-Image Edit arenas. Arena offered the model free direct access for 72 hours; Direct Mode access ends September 13 at 8 a.m. PT, after which it remains available anonymously in Battle and Agent Mode.
For the next 72 hours on Arena, you can use GPT-Image-2.5 Sunburst directly at no cost. [@OpenAI](https://x.com/OpenAI)’s latest image mo…
AI High Signal

Datasette announced security releases 1.0a39 and 0.65.4 after an extensive audit using Claude Fable 5.1, GPT-5.6 Sol, and GPT-6 Astra; the audit found and fixed a range of bugs, and operators running Datasette on public websites are urged to upgrade.

Datasette 1.0a39 and 0.65.4 security releases - [https://datasette.io/blog/2026/september-security-releases/](https://datasette.io/blog/2…
AI High Signal

Muse now lets all users request an audio podcast episode on any topic.

Hot off the presses: all [@Muse](https://x.com/Muse) users can ask their Muse to make an audio podcast episode on any topic. ![](https://…
AI High Signal

A commenter flagged a potential ambiguity in OpenAI’s consumer training opt-out: the policy applies to “Content,” defined as user Input and model Output, but hidden chain-of-thought reasoning is not shown to users, leaving its status under “Output” unclear. The commenter argues that the terms’ ownership and responsibility clauses could suggest hidden CoT falls outside “Content,” but explicitly says this does not establish that OpenAI trains on hidden CoT and notes they are not a legal expert. They requested clarification and reported no response after two days.

I went digging into OpenAI consumer training opt-out policy wordings, and I wonder if there's a loophole that allows training on hidden C…
AI High Signal

ChatGPT for Financial Services is now available as a tailored ChatGPT Work experience combining built-in financial data with GPT-6 Astra’s reasoning; teams can use it for research, financial modeling, and customized client materials.

Now available: ChatGPT for Financial Services. This is a tailored ChatGPT Work experience that combines built-in financial data with GPT-…
AI High Signal
  • Fugu Max and Fugu Ultra v2 were introduced as the next evolution of Fugu’s multi-agent orchestration system.
  • Fugu Max dynamically routes tasks across open-weight and specialized models, including NVIDIA Nemotron, and claims near-elite performance at 2–6× lower cost; listed pricing is $2 per 1M input tokens and $6 per 1M output tokens.
  • Fugu Ultra v2 claims the top result on five of eight hard benchmarks, a 74.3 DeepSWE score, and 48.3 on Chartography versus Opus 5’s 27.3, without Fable 5, Fable 5.1, or GPT-6-Astra in its agent pool.
Introducing Fugu Max and Fugu Ultra v2: the next evolution of Sakana Fugu’s multi-agent orchestration system. Try: [https://sakana.ai/fug… Fugu Max pushes the Pareto Frontier for price-performance: • $2 input / $6 output per 1M tokens. • Performance within striking distance o… Fugu Ultra v2 by the numbers. Full Results: [https://sakana.ai/fugu-max-release/](https://sakana.ai/fugu-max-release/) • [#1](https://x.c…