ZeroNoise Logo zeronoise
Post
AI’s New Battleground: Inference Economics and the Agent Harness
1 day ago
4 min read
674 docs
OpenAI’s price cut, NVIDIA’s high-throughput inference, and controlled harness tests point to a systems race, alongside new physical-AI and enterprise open-model signals.

Top Stories

Why it matters: Agents are increasingly competed on cost, latency, and scaffolding—not only model capability.

  • Inference economics is moving into the product surface. OpenAI cut GPT-5.6 Sol API prices 20% for input and 33% for output, to $4/$20 per million tokens through at least Nov. 21. Arena reports that the cut moved Sol’s frontier position in Code and Work, while Luna reached the overall and all three category frontiers at $0.04–$0.08 per task. Artificial Analysis also measured 3,431 output tokens/s for Gemma 4 31B on a private Groq 3 LPX endpoint, stable from 10K to 100K input tokens; NVIDIA says the rack is in full-scale production.

  • Harness choice is now a measurable competitive variable. In a same-model Opus 4.6 task—cloning Excalidraw from browser inspection through verification—Factory Droid finished in 8 minutes with 20 tool calls at $1.60, versus Claude Code’s 19 minutes, 40 calls, and $1.87. The tester says the harness was the only difference; it is one controlled comparison, but a clear cost and execution lever.

  • Physical AI posted a sharp benchmark jump. Team Tianjiao’s robot won the World Humanoid Robot Games long-jump final at 7.97 meters, 0.98 meters below Mike Powell’s record and well above the 1.25-meter robot best reported for 2025. It is a concrete athletic-control result, not evidence of general robotics.

Research & Innovation

Why it matters: The highest-leverage work is moving into the agent loop—serving, evaluation, and recoverable execution state.

  • AgentX 1.0 makes agentic inference measurable. SemiAnalysis released an open-source multi-turn coding benchmark built from roughly $3 million of real traces across 1,000-plus chips and about 2 MW of compute. vLLM reports sparse KV retention above a 95% cache-hit rate for 14 concurrent requests with contexts up to 1M tokens, while rate-matched prefill/decode reached 4.45× the throughput at 60 tok/s on GB300 Dynamo versus B300.

  • Video serving is being redesigned, not merely scaled. NVIDIA SANA’s Sol Engine work on MiniMax H3 combines a four-step low-resolution draft with a three-step LTX refinement pass and reduced 10-second 768p generation on one GB200 from 414 seconds to 14.93 seconds—a reported 27.7× speedup.

  • ACES questions static agent-skill gates. Across 145 real skills, structural scan scores correlated with LLM-judge quality at only Spearman ρ=0.14. The proposed Skill Lift instead compares the same task with and without a skill under identical conditions; its evaluation covered 947 paired cases, 58 production skills, and four harnesses.

Products & Launches

Why it matters: Agent products are becoming cross-provider control planes and enterprise-ready connector layers.

  • AgentSky launched an “OpenRouter for Agents.” Its API connects Claude Code, Codex, DeepSeek, Kimi, OpenCode, and other cloud agents; Agent Playground runs identical tasks with real tools such as GitHub and Gmail while comparing time, cost, and tokens side by side. The launch reports a $150-versus-$2 same-task gap but leaves the cheaper system for readers to guess.

  • Claude’s enterprise-managed MCP authentication is generally available. Admins centralize authorization through an identity provider, users connect tools without individual OAuth, and developers can apply the system to third-party connectors in Claude’s directory.

Industry Moves

Why it matters: Organizations and governments are pairing open models with proprietary data and dedicated compute rather than relying entirely on frontier APIs.

  • Thomson Reuters is pursuing model ownership. DatologyAI says Thomson-1.0-Large used its domain-relevant proprietary data, cost $450,000 in compute, and is competitive with closed frontier models at a fraction of deployment cost. A current report says Thomson Reuters built the model on Alibaba’s Qwen to reduce reliance on Claude.

  • South Korea is funding a narrowed sovereign-model race. Its government-backed competition reduced Round 2 from four teams to Upstage, SK Telecom, and LG AI Research. Each advancing team is expected to receive roughly 1,000 NVIDIA B200 GPUs for six months—about ₩40 billion per team—with the field planned to narrow to two in early 2027.

Policy & Regulation

Why it matters: Regulatory scrutiny is reaching the financial infrastructure funding AI bets.

  • The SEC is examining an AI hedge fund’s financing and leverage. MTSlive reports, citing the New York Times, that the agency subpoenaed banks that lent to and traded for Situational Awareness and ordered them to preserve related records.

Quick Takes

Why it matters: Practical deployments continue to spread across climate response, defense, and everyday AI operations.

  • Flood forecasting: Google says Flood Hub and Groundsource forecast riverine floods up to seven days ahead and urban flash floods up to 24 hours, with alerts intended for 2 billion people across 150 countries.
  • Defense autonomy: A team building autonomous interceptors in Ukraine reports 45 units sold to a special-forces unit and $500 million in letters of intent, all within eight weeks.
  • Usage controls: ChatGPT Work and Codex will restore a five-hour Plus-account limit to smooth compute demand and prevent accidental exhaustion of weekly usage; the $100 and $200 Pro tiers remain exempt for now.
AI’s New Battleground: Inference Economics and the Agent Harness
AI High Signal
  • Team Tianjiao’s humanoid robot won gold in the long-jump final at the 2nd World Humanoid Robot Games with a 7.97-meter jump . A related commentary post says this remains 98 centimeters below the male human record and predicts that robots clearing 9 meters could no longer be surpassed by baseline humans; the latter is speculative rather than an observed result .
In the long jump final at the 2nd World Humanoid Robot Games, Team Tianjiao's robot cleared 7.97 meters to win gold. [![Video](https://pb… This is 98 centimeters worse than the male human record which has stood for 35 years. The previous record was 5 cm worse and stood for 23…
AI High Signal
  • Apodex introduced the Apodex 1.1 model family, claiming frontier-level agentic performance for complex professional work, scientific research, financial analysis, and deep search.
  • The release adds an asynchronous Agent Team that decomposes tasks, coordinates agents in parallel, continuously integrates findings, and allows user guidance; Apodex also open-sourced the locally deployable FrontierAgent research workbench. Apodex 1.1 is available through its online workbench, while 1.1 mini is offered as an open-weight model for local execution.
  • A separate post says Apodex 1.1 mini is based on Qwen 3.5 35B A3B MoE.
Meet Apodex 1.1: Scaling Agentic Intelligence for Complex Work Open Source Harness: [https://github.com/ApodexAI/FrontierAgent](https://g… Apodex 1.1 mini is based on Qwen 3.5 35B A3B MoE btw [https://x.com/Apodex_AI/status/2091916791308313018](https://x.com/Apodex_AI/status/…
AI High Signal
  • MiniMax H3 is supported in vLLM-Omni for text-to-video with synchronized sound, frame-conditioned video generation with audio, image-plus-audio lip-synced clips, and automatic relighting/compositing of green-screen footage; vLLM published a deployment recipe.
  • The surrounding H3 ecosystem includes 24GB-VRAM local ComfyUI setups, enterprise SGLang and vLLM-Omni deployments, quantization guidance down to 8GB VRAM, serving optimizations, and developer tooling.
One model, four tricks: 🎬 text → video with synced sound 🖼️ first/first+last frame → video with audio 🎤 image + audio → lip-synced clip 🎥 … The MiniMax H3 ecosystem is growing faster than ever! ☺️ Check out Awesome MiniMax H3 Integrations - an index tracking everything built ar…
AI High Signal
  • Unlimited-token AI coding workflows are described as enabling unusually high productivity, including work across “8–10” projects at once—characterized as consistent with the baseline expectations of Staff+ engineers even before AI.
  • The discussion highlights a growing gap between unlimited-token frontier workflows and the economics of token usage, while framing agentic coding as a rapidly advancing frontier.
It's intersting (and slightly enjoyable?) to see the amount of doubt about my workflow on Reddit: [https://www.reddit.com/r/ClaudeAI/comm… One of my favorite takes is "anyone in the business knows that there's no universe where someone works '8-10' projects at a time." Even b… I 100% agree that my quote in the developer newsletter is indicative of the growing gap between unlimited token workflows and economic us… But today's Fable is tomorrow's Sonnet, and you have to remember that my job is to push the boundaries of agentic coding. I'm incredibly …
AI High Signal
  • The SEC subpoenaed banks that lent to and traded for Situational Awareness, an AI hedge fund, seeking details on its trades and leverage and ordering them to preserve all related records, according to the New York Times.
SITUATION DETECTED: The SEC has subpoenaed banks that lent to and traded for the AI hedge fund Situational Awareness, demanding details o…
AI High Signal
  • Alex Zhang introduced Speculative Programmatic Tool Calling (sPTC), a general technique that speculates on tool calls during code generation and queues them early in a harness to overlap token generation with REPL execution time.
Introducing Speculative Programmatic Tool Calling (sPTC)! A general class of technique for speculating on tool calls during code generati…
AI High Signal

China’s Galileo robotic dog is described as changing shape like a Transformer to adapt to tough environments and varied tasks. The post’s author characterizes the demonstration as a “brutal flex of industrial capability” and evidence of who is a “live player.”

✨🇨🇳China’s Galileo robotic dog! It changes shape just like a Transformer, adapting to tough environments and various tasks.😯 [![Video](htt… Seeing a brutal flex of industrial capability like this, and then these dismissive dunks, really drives home who is and who isn't a live …
AI High Signal

Ben Springwater argued that even AI adopters have not fully grasped what computer/browser use can do, pointing specifically to “5.6 Sol High in Codex” as an example. @mckbrando replied “correct,” endorsing the assessment.

It's my sense that most people, even AI adopter types, haven't really woken up to what computer/browser use can do - specifically 5.6 Sol… correct [https://x.com/benspringwater/status/2091961866943873388](https://x.com/benspringwater/status/2091961866943873388)
AI High Signal

A low-volume prediction market currently attributes the “0x-alpha” model to Zai_org, while noting that insiders may have incentives to leak information . Commentary cautions that participants may also try to misdirect others before betting, so this is unverified speculation rather than a confirmed model release .

0x-alpha model is from [@Zai_org](https://x.com/Zai_org) according to current prediction market (although the volume is low). Prediction … > Prediction markets are wild because insiders have an incentive to leak the info I think they have an even greater motivation to conv…
AI High Signal

An analysis of roughly 500,000 arXiv AI/ML papers reports a major shift toward Chinese open models in research: about 40% of papers now mention a Chinese open LLM versus 25–30% mentioning an American open model, reversing the 2024 split of roughly 10% versus 30%. The analysis cautions that papers lag model releases, so the figures reflect delayed adoption. Qwen appears in roughly one-third of papers that mention any LLM, while OpenAI’s closed models lead at about 37%; Llama peaked near 30% in April 2025 and has declined since, while DeepSeek saw a clear increase after R1’s January 2025 release. The share of AI papers mentioning any LLM rose from 10.43% in January 2023 to 55.49% in January 2026.

Over the weekend I had Codex parse 500K arXiv AI/ML papers since ChatGPT to understand which open models are used for research. In 2024, …
AI High Signal
  • A Team Tianjiao humanoid robot won gold at the 2nd World Humanoid Robot Games in Beijing with a 7.97-meter long jump—just 0.98 meters short of Mike Powell’s 8.95-meter men’s world record. The best robot mark in 2025 was reportedly only 1.25 meters, signaling a sharp improvement in humanoid athletic performance.
Humanoid robot from Team Tianjiao recorded a 7.97-meter long jump to win gold at the 2nd World Humanoid Robot Games in Beijing. Just 0.98…
AI High Signal
  • Jerry Liu argues that software and systems of record should become agent-native: agents should be able to use existing tools such as Slack, rather than build replacements or rely on the tool’s own agent. He expects agent consumption of software to become exponentially larger than human consumption.
  • Garry Tan predicts that systems of record will need to become “AI harnesses” or risk replacement by agents.
not sure i fully get this. - i want my agents to use slack - i don't want my agents to build a new slack - i don't want to use slack's ag… Prediction: systems of record will need to become AI harnesses or face replacement by agents
AI High Signal

@teortaxesTex argues that DeepSeek has fallen behind two other startups because it relies more heavily on internal RL, and characterizes the gap between “0731” and “ox-alpha” as “Opus-shaped.” The same account alleges that z.AI may have privileged access to a transfer station and Claude trajectories, while explicitly framing the claim as opinion.

Pretty interesting confirms some of my guesses, as well as obvious accusations to spell it out, DeepSeek has fallen behind the other two … To be maximally blunt: [https://z.AI](https://z.AI) probably has privileged access to a transfer station and can liberally use Claude tra… This embarrassing state of affairs will likely persist at least until Chinese data industry matches American stage of maturity, or really…
AI High Signal

A medical-AI practitioner says their team trains personalized-medicine models on rented GPUs, starts from Meta’s DINOv2 model trained on billions of Instagram images for cancer-detection tasks, and uses ChatGPT and Claude to accelerate medical-AI algorithm development. They argue that a blanket ban on generative AI could stall medical-AI progress.

Inaccurate take, people use the same techniques behind ChatGPT for medical AI as well. I train models to advance personalized medicine on…
AI High Signal
  • OpenAI announced that GPT-5.6 Sol API and credit pricing will fall by more than 20% for the next three months, linking the reduction to continued capability gains and improved efficiency.
  • @scaling01 speculates that Anthropic could cut its pricing within a month because its current pricing is viewed as excessive, and questions whether Model 2 improved Anthropic’s inference stack; this is commentary rather than a confirmed announcement.
As we continue to push the frontier of capabilities while improving efficiency, we're dropping API and credit pricing of GPT-5.6 Sol by o… you know the drill Lisan is pissed Anthropic pricing is as terrible as it is, so we are hopefully going to get a price cut within the nex…
AI High Signal
  • Hermes Agent parallelism: Teknium disputes a claimed one-agent-in-parallel limit, reporting that he runs about 25 sessions concurrently, with each using 15–25 subagents, and asserting that Hermes has no such limits.
Whoever claimed this is absolutely wrong. My Hermes runs 15+ subagents all the time whenever it makes sense to... We absolutely have no "…
AI High Signal
  • Enterprise-agent workflow thesis: @imjaredz argues that mobile could be the “ultimate form factor” for enterprise agents because workers are awake beyond desk hours; the post predicts an asynchronous model in which cloud agents run continuously and constantly ask humans for unblocking and steering.
Counterintuitively, mobile is the ultimate form factor for enterprise agents. It’s maximally capitalist. Your employees are awake for 16 …
AI High Signal

Figure says it is “accelerating again” and undergoing another “step change in capabilities”; it plans to announce a “critical update” tomorrow that it describes as necessary to solve general robotics.

Yes, Figure is accelerating again. We're going through another step change in capabilities Tomorrow we'll be sharing a critical update, o…
AI High Signal
  • Alibaba’s Qwen3.8-27B ranks #9 overall in Code Arena: WebDev with 1,595 points, making it the only model in its size class in the top 10; it trails the much larger Qwen3.8-Max by six places.
Exciting news: Qwen3.8-27B by [@Alibaba_Qwen](https://x.com/Alibaba_Qwen) just landed in Code Arena: WebDev at [#9](https://x.com/hashtag…
AI High Signal

Hark says it is building models intended to use computers better than humans, while its Handoff system operates across the internet—the dominant environment for computer use—rather than within a single application.

Hark is building models that can use computers better than humans Handoff operates across the internet, which is most of the computer use…