ZeroNoise Logo zeronoise
Post
Open-Weight Models Reach Majority Share on Vercel as Agents Reprice the AI Stack
16 hours ago
6 min read
2718 docs
Open-weight models reached 62% of token share on Vercel AI Gateway while agents consumed nearly five times human token volume, shifting investment attention toward model-agnostic harnesses, specialized infrastructure, proprietary data access, and redesigned go-to-market.

1. Funding & Deals

Pre-financing, in-place access to proprietary data is the clearest deal thesis in the current slice. A founder who has spent a year speaking with AI labs and data-holding institutions says labs have “burned through” open-internet data and now want clinical, chemistry and drug-discovery, and regional-language data, while institutions almost never sell or hand over copies because legal, privacy, and IP concerns stop the conversation before pricing. The proposed layer licenses access while training runs where the data sits; the founder says they have mapped roughly 700 potential institutions and 200+ AI labs, are starting with healthcare and drug discovery, and remain bootstrapped with no raise. This is a pipeline thesis rather than a financing event: diligence whether researchers will accept in-place iteration, who can authorize it, whether labs will bypass the intermediary, and whether synthetic data closes the gap—the founder’s own open questions. Andrew Chen’s shorthand—acquisitions moving from users in 2012 to engineers in 2021 to training data in 2026—captures the strategic direction.

2. Emerging Teams

OdoReach has converted a WhatsApp-policy pain point into first paid demand. Its founder says the product uses the official Meta WhatsApp API rather than extensions, charges ₹699 per month with no markup on Meta’s fees, and is positioned as avoiding the account-ban problem faced by extension-based marketers. Fourteen businesses are reported to be using it, with ₹9,044 collected in 30 days; the customers came from talking in the same WhatsApp groups rather than pitching. The product is still buggy and early, so the signal is problem validation and founder-led distribution—not yet retention or a durable moat.

Suhail’s build thread shows team formation under live execution pressure. The latest update says “a bunch of people” are starting the following week, while one “super annoying bug” is currently killing the team. That is a useful hiring and execution signal, but the update discloses no product traction.

3. AI & Tech Breakthroughs

Agent harnesses are moving into specialist engineering work. Clem Delangue reports that NVIDIA built a coding harness to optimize CUDA GPU kernels and achieved a 100% score on ARC-AGI-3’s 25 public games, solving all 183 levels. He argues that agents will lower the barrier to running, optimizing, and post-training models and kernels. The benchmark claim is a reported result, but the strategic signal is that scarce kernel expertise is being packaged as an agent workflow.

The compiler and kernel layers are attacking deployment bottlenecks. A current post reports Mojo 1.0 open-sourced under Apache 2.0, with an MLIR pipeline intended to target CPUs, Nvidia GPUs, and mobile NPUs from one codebase rather than forcing Python prototypes to be rewritten in C++ or Rust. Separately, an independent developer reports that the Apache-licensed fast_trimul library for AlphaFold3-family models matches OpenFold-3 output within about 0.0006%, runs 4.5–6.8× faster on short sequences, uses roughly 2.2–2.4× less peak VRAM, and avoids recompilation for new sequence lengths. Those figures are self-reported, but they show why narrow software optimizations can create usable capacity when memory is the constraint.

More agents are not automatically more intelligence. Exponential View’s summary of an Anthropic multi-agent experiment says that when common evidence pointed to the wrong answer, most model families chose correctly only 17–36% of the time after discussion, while a single agent given the full evidence got it right nearly every time; Mythos 5 reached about 85%. The product implication is to treat diversity, evidence allocation, and dissent mechanisms as design problems rather than assuming that adding agents improves reliability.

Agent products still need workload routing and deterministic boundaries. LlamaIndex’s ParseBench post says specialized OCR tools are generally much cheaper than coding agents on short documents, while coding-agent harnesses become more competitive on long documents because they can search snippets and use prompt caching. A separate verification project illustrates the same boundary: its deterministic verifier passed 66/66 canonical cases, but the live end-to-end pipeline passed only 19/66, prompting a split between verifier correctness, production-contract integrity, and model generation.

4. Market Signals

Agent demand is accelerating while model usage shifts toward open weights. a16z says agents burn nearly five times as many tokens as human users, up 14× since February. On Vercel AI Gateway, open-weight models accounted for 62% of token share on Aug. 22, versus 28.4% on June 24; closed models fell from 71.6% to 38%. The post expects further movement as enterprise harnesses, CLIs, IDEs, and SDKs become model-agnostic. The combination supports investment in routing, context, tools, and observability rather than assuming value remains concentrated in one model provider. LlamaIndex CEO Jerry Liu makes the adjacent commercial point: SaaS is not dead, but must be repurposed and remonetized for agent consumption.

A stealth-model episode shows why provenance and pricing remain part of the moat. A Reddit discussion quoting an article says Ox Alpha appeared on OpenRouter from an anonymous third-party provider as a free coding and sustained-agent-work model. A developer claims near-frontier coding performance and possible GLM-family lineage, but those are community reports; the same discussion flags that the model may be free only temporarily, is not downloadable, and has unknown pricing. Treat it as a trial candidate and competitive watch, not an underwriting-grade benchmark.

Model price/performance is becoming a routing problem. Bindu Reddy’s operator chart puts DeepSeek Flash at roughly $0.05 per task, describes Fable as top-scoring but premium-priced, and calls models below the quality-cost frontier a “kill zone.” She separately characterizes Anthropic’s Opus 5 and Sonnet 5 as costing more with few quality gains than their predecessors. These are not independent evaluations, but they are a useful warning against equating the newest frontier release with the best economic choice.

AI-era PLG still turns into sales, but later and with a different org mix. SaaStr puts the threshold for adding a real sales team around $100M–$250M ARR in the AI era, versus roughly $30M–$50M for 2015–2022 PLG companies. Emergence Capital’s survey found 36% of venture-backed B2B software companies cut SDR/BDR headcount while only 19% increased it; sales engineers and professional services expanded more often. Vercel’s COO said an agent reduced a 10-person lead-qualification function to about 1.25 people while SDR quotas rose 30%. Underwrite technical selling, implementation, and customer success capacity even when prospecting is automated.

Data-center deployment now requires political permission as well as power. Exponential View argues that AI labs’ decade of messaging—promising enormous gains while warning that the technology could take jobs or become dangerous—has “exploded in their face” at the county level. It treats local opposition as inseparable from that messaging and from communities’ perception that data centers tangibly serve an out-group. This is an essayist’s framing rather than a forecast, but it is a real diligence variable for infrastructure-heavy companies.

5. Worth Your Time

  • Read — ParseBench paper / Appendix D. LlamaIndex links the paper and ExtractBench behind its short-document versus long-document cost/accuracy comparison.
  • Read — Mojo 1.0 architecture breakdown. The linked discussion goes deeper on the MLIR-based, heterogeneous deployment thesis.
  • Inspect — fast_trimul. Review the open kernel implementation behind the developer’s AlphaFold performance and VRAM claims.
  • Read — Why one AI is better than four. The essay connects multi-agent hidden-profile failures with the economics of pricing a useful unit of work.
  • Read — Everyone Ends Up With a Sales Team. The current SaaStr analysis provides the ARR threshold and sales-function split behind the GTM signal.
Open-Weight Models Reach Majority Share on Vercel as Agents Reprice the AI Stack
Suhail

Founding team & funding: Suhail, ex-Mixpanel CEO/founder (per his profile), is launching a new AI venture: he secured a seed round and acquired a domain/name , starting "with 2 8xB200s" after a period working on image models .

Technical & compute: He validated a basic RLVR post-training stack , got a key research piece working and is scaling it , and is using an autonomous AI scientist for optimizations . He acquired 64 B300s and locked down "much greater quantities of compute" while learning about the datacenter frontier , though he also hit GPU losses and networking delays .

Team & hiring: He made a first hire, is seeking #2 for post-training (RLVR/OPSD) or low-level model optimization , grew to a team of 3 , and has more people starting next week . He reports Silicon Valley people are generously helping .

5/ Funding secured. Seed round done. 6/ domain / name acquired 1/ it all started w 2 8xB200s excited to be back in the game again 2/ spent a lot of time reviewing the absolute fundamentals again; missed a lot in the world while working on image models there’s so much… 8/ basic RLVR post-training stack validated ![](https://pbs.twimg.com/media/HL6LiPhboAIl8wK.jpg) 12/ got a very key piece of research working and need to scale it up; lost all my GPUs today though so now I am GPU poor more coming but … 3/ Time to let my autonomous ai scientist rip on some new optimizations ![](https://pbs.twimg.com/media/HKZJymUa0AAp-8r.png) 10/ 64 B300s acquired - if you search hard enough, you'll find what you need 14/ much greater quantities of compute locked down; ready to fly; learned a lot about the frontier of the datacenter industry this week 9/ made the first hire ❤️ Looking for [#2](https://x.com/hashtag/2): post training (RLVR/OPSD/etc) or low level model optimization 13/ first day going from team of 1 to team of 3 ❤️ 16/ a bunch of people starting next week This one super annoying bug is killing us 11/ banging my head against the proverbial hill climbing wall but met a bunch of people who set me in the right direction; SV is very muc…
Latent.Space
  • Around Christmas 2025, agents began working reliably because model capability and the agent harness (environment, tools, context, guardrails) improved together; the author's thesis is that models keep absorbing harness capabilities into their weights, leaving a "harness for human attention" .
  • Anthropic's Claude Code (terminal-based coding agent led by Boris Cherny, launched Feb 2025) was the first built to seize the moment when model capability crossed harness reliability — handing the loop back to the model via bash/file access and permission rules rather than per-change human approval; it grew to roughly $1B ARR within six months .
  • Harness choice can matter as much as the model: Harness-Bench ran the same model on the same 106 tasks in different harnesses and got 52.4–76.2 (a 23.8-point spread with zero model change); OpenAI tripled GPT-5.6 Sol's ARC-AGI-3 score, 13.3%→38.3%, using only harness changes (retained reasoning and compaction) .
  • Training is moving inside the harness: OpenAI's codex-1 was trained with RL on real-world coding tasks in varied environments, and GPT-5.1-Codex-Max is the first model natively trained to operate across multiple context windows via compaction; Anthropic then deleted ~80% of Claude Code's system prompt after the model absorbed it .
  • Next absorption candidates: multi-agent orchestration, tool selection, memory, and self-improving harnesses; what cannot be absorbed is the human-facing layer — permissions, identity, trust, and legibility — so the harness becomes the agent's interface to scarce human attention .
  • Prediction: within a year, every agentic AI company will ship a "human attention policy surface" (like AGENTS.md) governing when the agent may interrupt, keep working, and decide alone vs. ask approval — and it will become a learnable component where every correction is training data .
The Evolution of the Agent Harness
Exponential View

An Exponential View essay reads the local backlash against AI data centers as the result of AI labs' own messaging: for nearly a decade they promised huge gains requiring fast, massive capital for software and 21st-century infrastructure, while warning the technology could possibly take your job, 'maybe kill you,' and should be stewarded by only a few — the pitch that 'AI is too important for America not to let us get on with it' has now 'exploded in their face' at the county level . The author calls data-center siting a 'Gordian knot' inseparable from that hapless messaging and from the fact that the concrete blocks, whatever their abstract economic or healthcare benefits, 'tangibly serve an out-group' local communities dislike; the rest of the essay examines attitudes toward data centers and whether they actually benefit local communities .

🏦 The problem with petards
Paul Graham

Paul Graham says founders often don't know how well they're doing — both those making mistakes and those doing great — because running a startup is unlike other work; by the third startup you'd have a comparison, but with the first you don't . His key evaluation heuristic: growth rate is the "high bit" (in base 10) for judging a startup; if growth is good — unless it comes from giving away money — everything else is probably good enough .

It's very common for founders not to know how well they're doing. In both directions. It's obviously common for founders to be making mis… I think the reason is that running a startup is so unlike other kinds of work. By your third startup, you'd know if you were doing great.… Want to know my secret for telling how well a startup is doing? What's their growth rate? That's the high bit, and in base 10. If their g…
andrew chen

Andrew Chen observes that company acquisition rationale has shifted from gaining users (2012) to hiring engineers (2021) to acquiring training data (2026), signaling that proprietary training data is becoming a key strategic asset for AI-focused acquirers .

2012: acquire company for users 2021: acquire company for engineers 2026: acquire company for training data
Sriram Krishnan

Sriram Krishnan says demand for cyber defense solutions at nation-state/large enterprises is hard to overstate, and recent events from Mythos/Fable/OAI-HF over the past few months have strengthened the view that a totally different set of solutions is needed .

hard to overstate the demand for cyber defense solutions at various nation state/large enterprises. from Mythos/Fable/OAI-HF, these past …
Exponential View
  • Anthropic's multi-agent experiment: when common evidence points to the wrong option and only a few agents hold the correct facts, most model families chose correctly in only 17–36% of runs; a single agent given the full evidence base got it right nearly every time; only Mythos 5 escaped at ~85% .
  • LLM systems lack diversity (low-variance) — e.g., 30 agents given the same coding task result in 18 naming their git branch identically — and lack human-like institutions that protect lone dissenters; Thinking Machines advocates an ecosystem of AIs raised with different values and purposes to "keep the weirdness alive" .
  • Token price elasticity is modest: a 10% price cut lifts token use only 12–18%, not enough to trigger a strong Jevons paradox; the bottleneck may be pricing "a useful unit of work," not token cost .
  • Top-tier firms are pulling away in AI adoption: since October 2023, the top 1% of firms raised AI spend per employee by $6,542, while the median rose only $9.63 .
🔮 Why one AI is better than four #598
a16z

AI agents now burn nearly 5x the tokens human users do, up 14x since February, and humans are no longer the majority user of AI — a marker of accelerating agentic AI adoption . Source points to a16z's Charts of the Week for the underlying data .

Humans are the minority user of AI Agents burn nearly 5x the tokens people do, up 14x since February Charts of the Week: [https://www.a16…
Bindu Reddy

Bindu Reddy (@bindureddy), CEO of Abacus AI, shared a cost-vs-quality analysis of AI models, claiming the Pareto frontier is currently 'brutal': DeepSeek Flash offers 'insane quality at ~$0.05/task' , Gemini is 'the mid-range value king' , Sol is 'pushing 81 at an excellent cost' , and Fable has 'top score, premium price' . She warns that models below the frontier line are in the 'kill zone' — costing more or performing worse .

🚨 AI Models - Quality vs. Cost The Pareto frontier for quality vs. cost is BRUTAL right now: 🔹 DeepSeek Flash — insane quality at \~$0.05…
clem 🤗

Open-weight models set a record 62% share of tokens on Vercel AI Gateway on Aug 22, up from 28.4% on Jun 24 (~2 months earlier), while closed models fell to 38% from 71.6%; the post expects this shift to continue as enterprise adoption is still early and tooling becomes model-agnostic . HuggingFace CEO Clement Delangue amplified this, arguing the overwhelming majority of AI workloads will ultimately be based on open models .

Today's a record day for open weight share of tokens on Vercel AI Gateway. Aug 22 (today) 🟦 Open: 62% 🟨 Closed: 38% Jun 24 (\~2 months ag… Ultimately the overwhelming majority of AI workloads will be based on open models! [https://x.com/rauchg/status/2091234105702887879](http…
Deep Learning
  • Independent researcher unveiled CWAA (Complex Wave Associative Memory), a sequence-mixing architecture that replaces quadratic attention with a damped complex oscillator recurrence, giving O(T) linear memory scaling .
  • At ~10.4M parameters, trained 5,000 steps on WikiText-103, CWAA V6 reached validation PPL 139.09 vs 149.19 for a vanilla Transformer baseline . This result is not compute-matched: CWAA trains at 0.38s/step vs 1.08s/step for the Transformer on the same T4, and the author plans wall-clock-matched and multi-seed runs to confirm .
  • CWAA-only scaling shows 15.7k tok/s and 7.27GB VRAM at 4,096-token context; standardized inference/VRAM comparison against Transformer is still being re-run under identical conditions .
  • Architecture details: RetNet-style multi-scale decay initialization with half-lives spread geometrically from ~8 to ~512 tokens, static rather than data-dependent decay, and complex oscillatory phase dynamics; the author has not benchmarked against Mamba3 and notes the state dynamics/parameterization differ from Mamba's complex-valued update .
  • Community caution: oscillatory sequence models have historically drifted or destabilized over longer runs; the author reports V6 stayed stable for 5,000 steps but longer-run stability remains untested .
  • Author disclosure: AI tools did much of the implementation, CWAA is slower than the optimized Transformer despite better PPL, and the author calls for ablations and long-context tests before making bigger claims . Code: https://github.com/Ridhvik-2024/CWAA-V5.
A transformer built on complex wave dynamics — beats vanilla Transformer at 10M You need to compare training throughout and make a compute apples to apples comparison. Use the same wall clock time for both. Also do 2 … Comparing by step count isn't a true compute match, especially since the V6 rewrite runs at \~0.38s/step while the Transformer baseline t… synthetic tasks are exactly what’s needed to stress-test the recurrence limits. This architecture uses a RetNet-style multi-scale initial… I haven't actually compared it against Mamba3 yet. CWAA uses a damped complex oscillator for the recurrent state, with the decay/frequenc… I’ve played with oscillating ML and transformer setups for quite a while. My experience was that they could be difficult to keep stable o… Thanks, this is exactly the kind of feedback I'm looking for. V6 stayed stable through 5k steps, but I don't want to say that into “it's … Fair. AI did a lot of the coding, so I won't pretend this is 100% hand-written. I designed the architecture and experiments and I'm using…
Bindu Reddy

Abacus AI CEO Bindu Reddy argues Anthropic's latest frontier models, Opus 5 and Sonnet 5, are regressions that spin, cost more, and show few quality gains, recommending prior Opus 4.8 and Sonnet 4.6 .

Anthropic's latest generation models Opus 5 and Sonnet 5 feel like legit regressions They spin a lot, cost more and don't show many quali…
Natural Language Processing

An r/LanguageTechnology post asks whether scaling synthetic post-training data yields diminishing returns as model power grows (most synthetic examples being shallow variations of already-known knowledge) and suggests small sets of challenging real-world tasks with ground-truth answers may be the higher-value signal . It names Parsewave as a company working on post-training data generation for engineering-type tasks, along with evaluation and traces, framed as a possible emerging trend in data collection rather than a company analysis . A commenter asked whether the post was an ad for the company, a credibility flag on the mention .

Parsewave and the Problem of “More Data” in Post-Training Is this an ad for your company?
Bindu Reddy

Bindu Reddy (@bindureddy), CEO of Abacus AI, says Anthropic's new Opus 5 and Sonnet 5 models are "legit regressions": the latest generation spends far more tokens with almost no quality gains, so she recommends sticking with Opus 4.8 and Sonnet 4.6 .

Anthropic's new models Opus 5 and Sonnet 5 are legit regressions The latest generation spins and spends a lot more tokens with almost no …
Bindu Reddy

Bindu Reddy (CEO of Abacus AI) flags that Chinese open-source models are slow, get stuck and spin often, and are frequently severely rate limited, adding that post-training tweaks and bug fixes appear to be what's needed .

Chinese open source models would be great if they weren’t so slow They get stuck, spin a lot and often times are just severely rate limit…
Bindu Reddy

DeepSeek's new "Flash" vision variant is reportedly very good for attachments, better than "Luna" and 4x cheaper than Luna's prices, but DeepSeek has not open-sourced it yet .

DeepSeek's Flash new vision variant is very good for attachments Better than Luna and 4x cheaper than Luna's prices The only catch - Deep…
Amjad Masad

Replit CEO Amjad Masad said the team shipped 7 features in one week, including Free Mode, Conversations, Routines, steer conversations, importing Skills from GitHub, control models across Workspaces, and black-box pen tests .

A week has 7 days That means 7 ships [https://x.com/raoufcode/status/2091311652494803424](https://x.com/raoufcode/status/2091311652494803… The team has been busy lately. This is what shipped on Replit this week: - Free Mode - Conversations - Routines - Steer conversations - I…
Jerry Liu

Jerry Liu (CEO of LlamaIndex) posted that there's a real opportunity for AI labs to bake even higher switching costs into coding agents like Claude Code, Codex, and Grok Bot , citing his own inertia to switch apps because of existing skills, routines, system instructions, and project setup . He maintains an external wiki/artifacts per project that any app could point to, but this loses the nuances of conversation history ; enabled memory features let apps index and remember context for subsequent sessions, reducing the need to retype context . He calls efficiently feeding the right context the biggest pain point and suggests a good, self-improving context graph as the answer .

i think there's a real opportunity for labs to bake in even higher switching costs between claude code/codex/grok bot etc. whenever a new…
Suhail

Mixpanel founder Suhail asked on X whether anyone has gotten DeepSeek v4 flash 0731 working with "miles" for training, then replied that he would do it himself .

Has anyone gotten deepseek v4 flash 0731 to work with miles for training? ok i shall do it myself
Deep Learning

A developer released fast_trimul, an open-source (Apache-2.0) drop-in library for fused Triangle Multiplicative Updates across AlphaFold3-family protein models, built on Python CuTe DSL and designed as vendor-agnostic, modular, and production-friendly with a PyTorch fallback . Against OpenFold-3 it matches output within ~0.0006%, removes kernel-launch overhead via graphed execution (~22ms vs ~53ms at small N), runs 4.5–6.8× faster on short sequences, uses ~2.2–2.4× less peak GPU VRAM (fitting ~1.4× longer sequences), and avoids recompilation for new sequence lengths unlike torch.compile . It targets the memory-heavy Triangle Multiplicative Update in protein-structure models, the model family whose creators won a 2024 Nobel Prize . GitHub: https://github.com/tiagomonteiro0715/fast_trimul

I wrote a GPU kernel that speeds up AlphaFold style protein models