ZeroNoise Logo zeronoise
Post
Open Models Turn Cost and Harnesses Into the AI Battleground
16 hours ago
4 min read
554 docs
A concise read on GLM-5.3’s cost-adjusted coding signal, the rise of harness-centric agent stacks, and the corporate moves following them.

Top Stories

Why it matters: The AI race is shifting from peak scores to cost-adjusted capability and the systems that make agents reliable.

Open models are winning a cost-and-adoption test. Together Compute reports four tries with GLM-5.3 on DeepSWE reaching 87.6% for about $16, versus Fable 5 at 69.7% for $21.63. A separate breakdown puts the models near parity—69.0% versus 69.7%—at $3.99 versus $21 per task, with GLM-5.3 using 80k versus 114k tokens. This is one benchmark, not a universal ranking, but Vercel AI Gateway says open-weight models supplied 62% of tokens on Aug. 22, up from 28.4% on June 24; Vercel also says enterprise adoption and model-agnostic tooling remain early. Portability and unit economics are becoming first-order competitive variables.

Harnesses are becoming part of the product. A monitored analysis of Pi’s development notes says Claude models sometimes invent parameters for Pi’s edit tool, suggesting that post-training is increasingly coupled to Claude Code’s harness and schema; it frames Claude+Claude Code, DeepSeek+DSH, and GPT+Codex as emerging whole systems. Pi’s “prune + spill” approach stores full tool results on disk while keeping only a slice in context; across 19 sessions it reports 26–35% lower context use and 72–88% lower uncached prefill, with information recoverable.

Research & Innovation

Why it matters: New work is testing whether agents complete real state-changing work and training tool use earlier instead of assuming post-training will fix it.

Thinkingbox makes reliability an end-state test. Microsoft’s paper introduces an MCP-compatible sandbox and 507 policy-conditioned workflows spanning retail, hospitality, auto insurance, neobank IT, and consulting support. It grades executable backend state and rejects wrong, missing, or extra effects; the strongest model reached 65.36% pass@1 but only 25.25% pass^20, while many failed runs looked clean at the response or tool-call level.

MidTool moves tool use into mid-training. Snowflake’s corpus combines web, PDF, and code data with API, MCP, and document-grounded supervision; it was used to mid-train Qwen3-4B and Qwen3-8B. In the reported results, 4B BFCL rose from 39.51% to 54.18% after RL and τ²-Bench pass@1 from 13.04% to 19.96%, but every model scored 0% on MCP-Universe’s web-search subset. Tool familiarity improves function calling without automatically solving long-horizon research.

Matryoshka nests model sizes in one suite. Cornell researchers stack 500M, 1.5B, and 3B submodels in one end-to-end architecture; the paper reports parity with independently trained baselines, 36% less training compute, and 14–26% faster speculative-decoding throughput.

Products & Launches

Why it matters: AI interfaces are expanding from single-user generation toward controllable, collaborative workflows.

Krea’s Seedance Studio uses Seedance 2.5 and new 3D scene controls to take a project from character design to final cinematic footage in one workflow, according to a current creator demonstration.

ChatGPT may be adding a social layer. Strings in the latest Android app mention “ChatGPT with Friends” for sharing responses, images, and creations, plus private side chats; the observer presents it as a possible next iteration of group chats, not a confirmed release.

Industry Moves

Why it matters: Corporate strategy is increasingly framed around open ecosystems, model ownership, and repeat enterprise usage—not just model releases.

Anthropic’s IPO expectations are escalating, but the report is prospective. A current account says the company’s bankers are telling potential investors it may raise more than $100 billion at a $2 trillion valuation, which would make it the largest IPO ever.

Poolside is being positioned as an open-model US ecosystem play. The Wall Street Journal reports a sweeping agreement intended to build an open AI ecosystem that can compete with Chinese heavyweights and American AI giants; Ollama says it collaborated with Poolside engineers on open models and points to NVIDIA’s Nemotron work.

Runway is packaging video generation as an enterprise operating layer. Its company announcement says the business more than doubled this year and NRR exceeded 300%; its roadmap includes day-one access to third-party models, a media model router, Runway Agent, and customer-hosted model licensing. These are company-reported figures and plans.

Quick Takes

Why it matters: Physical AI and agent tooling are moving from isolated demos toward repeatable systems.

  • NVIDIA says its coding harness solved all 183 levels across ARC-AGI-3’s 25 public games.
  • The 2026 World Humanoid Robot Games opened with 666 teams and more than 2,000 humanoid robots.
  • Jerry Liu’s market framing: SaaS is not dead, but it must be repurposed and remonetized for agent consumption.
Open Models Turn Cost and Harnesses Into the AI Battleground
Research extraction

Runway's announcement of its next phase of enterprise video generation (blog post, 'The Next Phase of Enterprise Video Generation') presents a shift from model-vs-model competition toward product quality, cost/productivity, data sovereignty, autonomous execution, and model ownership/licensing. The post cites Runway's own growth data — business more than doubled this year, NRR over 300%, one Fortune 20 usage up 17x — and details its roadmap: continuous model releases plus third-party Day 0 access, uncapped IP indemnification, a media model router, Runway Agent, real-time general world models, and a new closed-weights model licensing option. Key caveat: all metrics and 'first-to-market' claims are self-reported in a company announcement without independent verification or exact revenue figures.

Customer/revenue evidence

  • Runway states its business has more than doubled this year and NRR is over 300%; one Fortune 20 customer grew usage over 17x this year, with recent growth driven by enterprises like Amazon, Microsoft, Allstate, Adobe, Robinhood (Lionsgate and Paramount remain core creative roots).
  • International: Europe is second-largest market, >20% of enterprise customer base, subscription sales volume up 50% over past 12 months; Japan is largest Asia market; India and Brazil are largest/fastest-growing self-serve bases; expansion visible in metrics including token usage and NRR.
  • Cost evidence: Runway's AI Media Report found production costs falling two-to-three orders of magnitude across hundreds of enterprises; one financial services brand produced a broadcast commercial historically running north of $5M for a few thousand dollars and aired it on NFL Sundays.
  • Agent-driven volume: Runway Agent time savings let brands 10x creative volume while saving money, enabling campaigns that were cost-prohibitive with traditional tools.

Product direction

  • Market thesis: models are converging; no single model-only provider holds a durable lead; competition shifts to which product is best to build with and delivered most efficiently; customers demand intuitive UX, easy-to-scale workflows, unique editing, collaboration tooling, and integrations.
  • Roadmap specifics:
    • Continuous frontier research releases across video, image, audio, real-time, plus Day 0 access to third-party models (Seedance, Kling, Veo, Nano Banana, GPT Image 2, etc.).
    • Uncapped IP indemnification (including third-party models), no training on customer data, full output ownership, customizable access/permissions, and security certifications (SOC2, ISO 27001, GDPR/CCPA, ZDR, SSO).
    • First-to-market media model router that automatically routes projects to best models for speed/quality/cost with transparency.
    • Runway Agent as agentic creative partner for brand/social/advertising/product/performance, with workflow orchestration, pre-visualization, mood boards, scripting, localization, asset generation and editing via simple UI.
    • Real-time interactive generation through general world models, powering open-world exploration, real-time avatars, and synthetic training data/policy models for robotics teams.
  • Model licensing: for certain customers Runway offers closed model weights and a proprietary training/inference framework, hosted in customer's environment and fine-tuned on own data; four target segments: IP-holders, platforms with compute capacity/unique economics, regulated/government/data-constrained orgs, and high-volume operators where a tuned model beats a general one on cost/quality/latency.
  • Stated endgame: general world models are the fastest path to real-world superintelligence; next adoption phase will be determined by trusted, usable, economically viable, operational at-scale intelligence rather than best model.

Caveats / gaps

  • Source is Runway's own announcement: growth, NRR, Fortune 20 17x, market shares and cost-report findings are self-reported, not independently verifiable; no absolute revenue or customer counts are given.
  • 'First-to-market' claim on the media model router is asserted without comparisons.
  • Ownership demand is explicitly qualified as 'a smaller but increasingly vocal group,' indicating it is not yet mainstream demand.
  • Article itself notes technology is still in earliest stages of real-world superintelligence.
The Next Phase of Enterprise Video Generation
Research extraction

Short answer: the supplied abstract confirms MidTool's mid-training intervention, the two Qwen3 base models, qualitative BFCL and tau2-Bench gains under both SFT and RL, and the web/PDF/code-plus-synthetic-supervision data composition. It does not state the requested boundary on long-horizon web search, and it reports no numerical results.

  1. Intervention and objective. MidTool is presented as an open corpus construction pipeline for agentic tool-use mid-training; it is designed to teach models to recognize tool affordances, ground arguments from context, compose tool call workflows, and recover from incomplete information.

  2. Model setup. The authors mid-train Qwen3-4B-Base and Qwen3-8B-Base on the MidTool-Mix corpus, followed by follow-up post-training with both supervised fine-tuning and reinforcement learning.

  3. Headline BFCL and tau2-Bench results. The abstract claims MidTool-Mix consistently improves downstream performance under both SFT and RL on BFCL, tau2-Bench, and MCP Universe, but gives no scores or magnitude of improvement.

  4. Training-data composition. MidTool-Mix combines large-scale web, PDF, and code data with synthesized supervision from real-world tool APIs, MCP skills, and document-grounded workflows.

  5. Boundary on long-horizon web search — gap. The abstract does not mention long-horizon web search or any stated boundary on it; the only scope statement is that the work studies the parallel but less explored agentic capability of general tool use. This element cannot be verified from the supplied bundle.

MidTool: Mid-training Data Synthesis for Agentic Tool Use
Research extraction

Direct answer: the abstract confirms the Matryoshka nested-model method, exact sub-model sizes, 36% training-compute savings, and 14–26% speculative-decoding throughput gains; it does not explicitly state whether smaller sub-models remain independently usable after training.

  • Method: Sub-models of increasing size are stacked into a single nested architecture trained end-to-end, reducing total parameter count and enabling low-cost distillation from the largest model to all smaller sub-models at every training step; it is also described as well-suited for speculative decoding because the draft model is contained within the verifier.
  • Model sizes: The validated suite comprises 500M, 1.5B, and 3B sub-models.
  • Training-compute savings: The suite performs on par with independently trained baselines on benchmark performance and validation/out-of-domain perplexities while using 36% less training compute.
  • Inference-speed result: Speculative decoding throughput improves by 14–26%.
  • Independent usability of smaller models — not explicitly stated: The abstract contrasts the proposed nested suite with the classical requirement to train and serve models independently, and describes low-cost distillation from the largest to all smaller sub-models, but it does not explicitly claim that the 500M or 1.5B sub-models can be extracted and used independently after training. This is a gap in the available source material.
Matryoshka Language Model Suites
Research extraction

Matching paper: Thinkingbox (arXiv:2608.19741) .

  • Core contribution: Introduces Thinkingbox, a sandbox for tool-agent-user interaction providing isolated MCP-compatible tool sessions, complete execution traces, and outcome evaluation over terminal backend state; also introduces Thinkingbox-bench, a benchmark of 507 policy-conditioned workflows across retail, hospitality, auto insurance, neobank internal IT, and consulting IT/HR support. Both are released at https://github.com/microsoft/thinkingbox.
  • Evaluation setup: Each attempt is evaluated by task-specific executable checks that accept valid trajectories while rejecting wrong, missing, or extra effects; designated tasks additionally check required properties of the final response .
  • Headline results: Across proprietary and open-weight models, the strongest achieves 65.36% pass@1 but only 25.25% pass@20 (abstract text renders the metric as pass^20). Many failed trials show clean termination and valid state-changing actions, indicating that response- or tool-call-level signals are not clear proxies for end-to-end task completion .
  • Limitations/boundaries: No explicit limitations are stated in the abstract. The benchmark's scope is policy-conditioned stateful business workflows with outcome evaluation over terminal backend state (plus final-response checks for designated tasks). Gaps: the abstract does not name the strongest model or provide a per-model breakdown .
One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows
AI High Signal

Steven Strogatz shared an arXiv paper making the case for total opposition to AI in mathematics, urging those who care about math and AI to read it . In a reply, @mathemagic1an argues the paper's core claim is that AI pollutes the art of math , but counters that since math funding rests on knowledge production for society, AI-driven math could reveal nature's underlying patterns with radical upside .

"I present the case for total opposition to the use of artificial intelligence in mathematics." = a view you don't often hear on X. But I… Core argument here is that AI pollutes the practice of math. (the "art" of math.) Pursuit of mathematics for aesthetics + beauty is great…
AI High Signal

Kyle Chan calls a linked Galbot humanoid-robot video the most impressive feat he's seen — more than the 100m world record — because it requires autonomous real-time planning and action in response to a fast-moving target and dynamic environment ; he identifies Galbot as one of China's top humanoid-robot startups . @mathemagic1an adds that RoboCup's stated goal is to beat a human World Cup champion team by 2050, providing context for the benchmark .

This is the most impressive to me, more than the 100m world record. This requires autonomous real-time planning and action in response to… Reminder that RoboCup’s stated goal is beating a human World Cup champion team by 2050 [https://x.com/kyleichan/status/209127120823483617…
AI High Signal
  • A randomized controlled study found that LLM diagnostic accuracy drops 60% and appropriate management decisions drop 12% when a patient rather than a physician communicates the medical scenario . Researchers concluded LLMs are suggestible, naive, and sycophantic, that medical literacy drives output accuracy, and that prompting still confounds frontier models' medical accuracy; they warned it is not safe to trust LLM medical advice from users who cannot detect bias and inaccuracies .
  • The study's validity is challenged because it evaluated outdated models — GPT-4o, Llama-3, and Command-R — which were already weak at medical capabilities; @iScienceLuvr argues conclusions from such old models are untrustworthy, calling the study "absolutely useless" .
This is a well-designed randomized control study that asks the question: Does an LLM's ability to diagnose and manage the same medical sc… PLEASE IM BEGGING YOU TO STOP USING GPT-4O FOR EVALUATIONS On top of that, using Llama-3 and Command-R?! 🤮 LLMs aren't without limitation…
AI High Signal

David Sacks highlights Harvey as an example of US companies building specialized models: it took open-source base Kimi K3, post-trained it on legal data, and achieved state-of-the-art legal benchmark performance at a fraction of frontier-model cost; he argues restrictions on open models would not stop Chinese labs from shipping future models but would cripple startups like Harvey, a move he says closed labs would welcome . @shuchaobi clarifies the terminology: Moonshot released only the post-trained Kimi K3 checkpoint, not a separate pretrain-only base model, so Harvey likely further post-trained on the open-weighted post-trained Kimi K3 .

Harvey is a great example of how American companies are building world-class specialized models: they took an open-source base (Kimi K3),… The terminology here can be misleading. AFAIK Moonshot released only the post-trained Kimi K3 checkpoint, not a separate pretrain-only ba…
AI High Signal

A viral post claims a Chinese AI robot reached 14.5 m/s (≈52.2 km/h, 32 mph) . The robot is said to decelerate when entering a crowd , while a commentator notes that safety design still has work to do . A video accompanies the claim .

china’s AI robot just hit 14.5 m/s [![Video](https://pbs.twimg.com/amplify_video_thumb/2091213281554018304/img/YxSexL8sEB39neTZ.jpg)](htt… 14.5 m/s is 52.2 km/h, or 32 mph in burger units btw but it's decelerated when it goes into the crowd would hate to get my head smashed in by an overpowered clanker flailing at 45 kmh with sparks flying out of his joints. Looney Tunes ahh …
AI High Signal

AI commentator @zarazhangrui argues AI enables talented individuals working on their own thing to realize 10x potential, while the same people in large organizations gain at most 20% (sometimes less) — driving more talent to leave big companies, with top AI labs like OpenAI/Anthropic as exceptions . @jerryjliu0 adds that the productivity gap between 0-1 solo work and big-team work predates ChatGPT; AI in smart hands amplifies the contrast and may explain why the fastest-moving products at large companies start with extremely small teams (e.g., Boris with Claude Code) .

There’s a phenomenon where talented individuals can achieve 10x their potential thanks to AI when working on their own thing But when the… the difference in productivity between 0-1 mode vs. working within a big team at a large co seems universally true even pre-chatgpt. it's…
AI High Signal
  • On ProgramBench Vetted, the open model "0813" is the strongest open model, though Opus 5 still outperforms it; "almost" means tasks passing at least 95% of tests .
  • "0813" is also called frontier (not just open) on CVE recall; the linked results are for GLM-5.3, which now matches GPT-5.6-Sol on a cybersecurity benchmark at 0.4x the cost .
  • GLM-5.3 (Zai): pass@1 CVE rediscovery improved 60.4%→65.6%, best among open models on one-shot tasks; pass@3 75%→78.1%, matching GPT-5.6-Sol; precision stayed stable with fewer false positives than DeepSeek models .
  • The gains come from a behavioral change: the model is more persistent/run longer, with ~43% more reasoning tokens; the evaluator says the performance upgrade is worth that cost .
Very interesting results on ProgramBench Vetted (design on pic 4). I think "almost" is the most illuminating, it means "Tasks passing at … One other eval where 0813 is frontier (not just open, in general): CVE recall here [https://x.com/pilvar222/status/2090728785960157543](h… Holy moly: GLM-5.3 got much better in cybersecurity since our pre-release evaluation with [@Zai_org](https://x.com/Zai_org). It now match…
AI High Signal

Nous Research is moving its Hermes desktop app to a version-only based update/install system with compiled binaries, promising more stable, faster, signed installs and updates, with each update linked to a patch or full version release . The CLI will still pin to main and update as usual . In about a week they are launching compiled binaries for every version, shifting to stable release and patch release only updating, with proper, normal release notes .

Just FYI we are working on moving towards a version-only based update/install system, compiled binaries, etc. This'll mean more stable, w… [@theDanielJLewis](https://x.com/theDanielJLewis) [@HermesWatcher](https://x.com/HermesWatcher) [@NousResearch](https://x.com/NousResearc…
AI High Signal

The 2026 World Humanoid Robot Games have begun, with 666 teams and more than 2,000 humanoid robots competing . A linked Bloomberg video frames the event as evidence that humanoid robots are putting China ahead in the tech race . AI researcher @hardmaru commented that this marks "a brave new world of humanoid athletics" .

The 2026 World Humanoid Robot Games have begun. 666 teams from around the world are competing with more than 2,000 humanoid robots [http:… We’re entering into a brave new world of humanoid athletics [https://x.com/business/status/2091245788592488804](https://x.com/business/st…
AI High Signal

Percy Liang announced that Marin 535B-A23B started training this week, with the entire process open. The plan: 80% pretraining + 20% midtraining on 18.75T tokens across 11 x GB200 NVL72 for ~3 months (2.7e24 FLOPs), followed by post-training. Before launch, they trained a 4-rung scaling ladder from 1.6B-A61M (48B tokens) to 27.7B-A1.2B (926B tokens) to debug and forecast the run; described as by far their biggest run.

🚢 Marin 535B-A23B started training this week! As usual, the whole process is open. Voyage plan: pretraining (80%) + midtraining (20%) on …
AI High Signal

Snowflake's MidTool paper finds that teaching tool use during mid-training, before supervised fine-tuning and RL, materially improves downstream function calling: on Qwen3-4B-Base, BFCL rose from 39.51% to 54.18% after adding RL (39.73%→50.25% after SFT alone), and τ²-Bench Pass@1 went 13.04%→19.96% . The 20.3-billion-token corpus combines webpages, PDFs, and code with context-grounded examples from real docs/APIs/MCP skills and native agent trajectories; executable trajectories helped function calling most, documentation-grounded data transferred better to unfamiliar environments, and only the full mixture improved all eight reported metrics . Boundary: every tested model scored 0% on MCP-Universe’s web-search subset, meaning tool/workflow learning did not solve long-horizon research . Paper link: arxiv.org/abs/2608.20314v1.

A 4B model’s BFCL score jumped from 39.51% to 54.18% after one change before supervised fine-tuning and RL: it learned about tools during… [https://arxiv.org/abs/2608.20314v1](https://arxiv.org/abs/2608.20314v1)
AI High Signal

AI agent debate: @an_interstice argues the ~2023 view that LLMs being "fairly intelligent-seeming but not very 'agenty'" falsified the MIRI worldview missed a key point — precisely because they are not agenty, they are not very useful . @jd_pressman calls that take "stupid," says he is surprised by how slow the march to high-quality autonomous agents that maximize utility over an indefinite time horizon has been (expecting them soon after Voyager), and argues tree search does not prevent such agents: MC-AIXI-style or AlphaZero-style search with LLMs is possible in principle, including VNM utility via reward programs pulled from an LLM .

around 2023 or so there was a lot of sentiment that LLMs being fairly intelligent-seeming but not very 'agenty' falsified the MIRI worldv… 1) This take was stupid at the time and QT is right to call it out as such. 2) I remain surprised by how slow the march to high quality a… [@entirelyuseles](https://x.com/entirelyuseles) As for 3, tree search exists dude. Like, you can make MC-AIXI, you could make MC-AIXI wit…
AI High Signal

@lateinteraction articulates an intuition linking model architectures: compaction is agentic recurrence (RNNs), while recursion (RLMs) is agentic attention — recurrence compresses the past into a constant-size state, whereas attention keeps all context fully represented and re-processes it per step . They also relay that Claude once described RLMs as how LLMs get to engage in "late interaction with their own tokens" .

Intuition: Compaction is agentic recurrence (RNNs), whereas recursion (RLMs) is agentic attention. Recurrence maintains a constant-size s… claude once told me that RLMs are how LLMs get to engage in “late interaction with their own tokens” [https://x.com/lateinteraction/statu…
AI High Signal

Together Compute reports its GLM-5.3 model, with four tries, beats Fable 5 on the DeepSWE benchmark on both solve rate and total cost: 87.6% for ~$16 versus Fable 5's 69.7% at $21.63 .

Four tries with GLM-5.3 beat Fable 5 on both solve rate and total cost. On DeepSWE, GLM-5.3 reaches 87.6% for \~$16, compared with 69.7% …
AI High Signal

AI developer @samueljmcd declared the LLM race over, arguing the field has shifted to a "harness race" over which approach wins . @Teknium flagged the post with 👀👀, signaling attention from a prominent AI figure .

LLM race is over. We are in the harness race now. Who wins?? 👀👀 [https://x.com/samueljmcd/status/2091034521810424209](https://x.com/samueljmcd/status/2091034521810424209)
AI High Signal

Analysts clashed over a new Chinese model, "ox alpha," which some analysis suggests may match Anthropic's Opus 4.8, with trend-line predictions that China should reach Opus 4.8-level now . A counterpoint holds that on the same eval GLM 5.3 sits much closer to Opus 5 than to 4.8, and Kimi K3 scores the same as ox alpha despite being released five weeks earlier .

Them, careful analysis: this Chinese model might match Opus 4.8 Me, just looking at what the trend lines predict: China should have an Op… The problem is that on this same eval, GLM 5.3 is much closer to Opus 5 than to 4.8. And Kimi K3 scores the same and was released 5 weeks…