ZeroNoise Logo zeronoise
Post
Open-Weight Models Reach Majority Share on Vercel as Agents Reprice the AI Stack
16 hours ago
6 min read
2718 docs
Open-weight models reached 62% of token share on Vercel AI Gateway while agents consumed nearly five times human token volume, shifting investment attention toward model-agnostic harnesses, specialized infrastructure, proprietary data access, and redesigned go-to-market.

1. Funding & Deals

Pre-financing, in-place access to proprietary data is the clearest deal thesis in the current slice. A founder who has spent a year speaking with AI labs and data-holding institutions says labs have “burned through” open-internet data and now want clinical, chemistry and drug-discovery, and regional-language data, while institutions almost never sell or hand over copies because legal, privacy, and IP concerns stop the conversation before pricing. The proposed layer licenses access while training runs where the data sits; the founder says they have mapped roughly 700 potential institutions and 200+ AI labs, are starting with healthcare and drug discovery, and remain bootstrapped with no raise. This is a pipeline thesis rather than a financing event: diligence whether researchers will accept in-place iteration, who can authorize it, whether labs will bypass the intermediary, and whether synthetic data closes the gap—the founder’s own open questions. Andrew Chen’s shorthand—acquisitions moving from users in 2012 to engineers in 2021 to training data in 2026—captures the strategic direction.

2. Emerging Teams

OdoReach has converted a WhatsApp-policy pain point into first paid demand. Its founder says the product uses the official Meta WhatsApp API rather than extensions, charges ₹699 per month with no markup on Meta’s fees, and is positioned as avoiding the account-ban problem faced by extension-based marketers. Fourteen businesses are reported to be using it, with ₹9,044 collected in 30 days; the customers came from talking in the same WhatsApp groups rather than pitching. The product is still buggy and early, so the signal is problem validation and founder-led distribution—not yet retention or a durable moat.

Suhail’s build thread shows team formation under live execution pressure. The latest update says “a bunch of people” are starting the following week, while one “super annoying bug” is currently killing the team. That is a useful hiring and execution signal, but the update discloses no product traction.

3. AI & Tech Breakthroughs

Agent harnesses are moving into specialist engineering work. Clem Delangue reports that NVIDIA built a coding harness to optimize CUDA GPU kernels and achieved a 100% score on ARC-AGI-3’s 25 public games, solving all 183 levels. He argues that agents will lower the barrier to running, optimizing, and post-training models and kernels. The benchmark claim is a reported result, but the strategic signal is that scarce kernel expertise is being packaged as an agent workflow.

The compiler and kernel layers are attacking deployment bottlenecks. A current post reports Mojo 1.0 open-sourced under Apache 2.0, with an MLIR pipeline intended to target CPUs, Nvidia GPUs, and mobile NPUs from one codebase rather than forcing Python prototypes to be rewritten in C++ or Rust. Separately, an independent developer reports that the Apache-licensed fast_trimul library for AlphaFold3-family models matches OpenFold-3 output within about 0.0006%, runs 4.5–6.8× faster on short sequences, uses roughly 2.2–2.4× less peak VRAM, and avoids recompilation for new sequence lengths. Those figures are self-reported, but they show why narrow software optimizations can create usable capacity when memory is the constraint.

More agents are not automatically more intelligence. Exponential View’s summary of an Anthropic multi-agent experiment says that when common evidence pointed to the wrong answer, most model families chose correctly only 17–36% of the time after discussion, while a single agent given the full evidence got it right nearly every time; Mythos 5 reached about 85%. The product implication is to treat diversity, evidence allocation, and dissent mechanisms as design problems rather than assuming that adding agents improves reliability.

Agent products still need workload routing and deterministic boundaries. LlamaIndex’s ParseBench post says specialized OCR tools are generally much cheaper than coding agents on short documents, while coding-agent harnesses become more competitive on long documents because they can search snippets and use prompt caching. A separate verification project illustrates the same boundary: its deterministic verifier passed 66/66 canonical cases, but the live end-to-end pipeline passed only 19/66, prompting a split between verifier correctness, production-contract integrity, and model generation.

4. Market Signals

Agent demand is accelerating while model usage shifts toward open weights. a16z says agents burn nearly five times as many tokens as human users, up 14× since February. On Vercel AI Gateway, open-weight models accounted for 62% of token share on Aug. 22, versus 28.4% on June 24; closed models fell from 71.6% to 38%. The post expects further movement as enterprise harnesses, CLIs, IDEs, and SDKs become model-agnostic. The combination supports investment in routing, context, tools, and observability rather than assuming value remains concentrated in one model provider. LlamaIndex CEO Jerry Liu makes the adjacent commercial point: SaaS is not dead, but must be repurposed and remonetized for agent consumption.

A stealth-model episode shows why provenance and pricing remain part of the moat. A Reddit discussion quoting an article says Ox Alpha appeared on OpenRouter from an anonymous third-party provider as a free coding and sustained-agent-work model. A developer claims near-frontier coding performance and possible GLM-family lineage, but those are community reports; the same discussion flags that the model may be free only temporarily, is not downloadable, and has unknown pricing. Treat it as a trial candidate and competitive watch, not an underwriting-grade benchmark.

Model price/performance is becoming a routing problem. Bindu Reddy’s operator chart puts DeepSeek Flash at roughly $0.05 per task, describes Fable as top-scoring but premium-priced, and calls models below the quality-cost frontier a “kill zone.” She separately characterizes Anthropic’s Opus 5 and Sonnet 5 as costing more with few quality gains than their predecessors. These are not independent evaluations, but they are a useful warning against equating the newest frontier release with the best economic choice.

AI-era PLG still turns into sales, but later and with a different org mix. SaaStr puts the threshold for adding a real sales team around $100M–$250M ARR in the AI era, versus roughly $30M–$50M for 2015–2022 PLG companies. Emergence Capital’s survey found 36% of venture-backed B2B software companies cut SDR/BDR headcount while only 19% increased it; sales engineers and professional services expanded more often. Vercel’s COO said an agent reduced a 10-person lead-qualification function to about 1.25 people while SDR quotas rose 30%. Underwrite technical selling, implementation, and customer success capacity even when prospecting is automated.

Data-center deployment now requires political permission as well as power. Exponential View argues that AI labs’ decade of messaging—promising enormous gains while warning that the technology could take jobs or become dangerous—has “exploded in their face” at the county level. It treats local opposition as inseparable from that messaging and from communities’ perception that data centers tangibly serve an out-group. This is an essayist’s framing rather than a forecast, but it is a real diligence variable for infrastructure-heavy companies.

5. Worth Your Time

  • Read — ParseBench paper / Appendix D. LlamaIndex links the paper and ExtractBench behind its short-document versus long-document cost/accuracy comparison.
  • Read — Mojo 1.0 architecture breakdown. The linked discussion goes deeper on the MLIR-based, heterogeneous deployment thesis.
  • Inspect — fast_trimul. Review the open kernel implementation behind the developer’s AlphaFold performance and VRAM claims.
  • Read — Why one AI is better than four. The essay connects multi-agent hidden-profile failures with the economics of pricing a useful unit of work.
  • Read — Everyone Ends Up With a Sales Team. The current SaaStr analysis provides the ARR threshold and sales-function split behind the GTM signal.
Open-Weight Models Reach Majority Share on Vercel as Agents Reprice the AI Stack
Summary
Coverage start
1 day ago
Coverage end
16 hours ago
Frequency
Daily
Published
14 hours ago
Reading time
6 min
Research time
11 hrs 59 min
Documents scanned
2718
Documents used
22
Citations
28
Sources monitored
119 / 120
Insights
124
View
Skipped contexts
209
View
Source details
Source Docs Insights Status
Hunter Walk 0 0
SaaStr 2 1
andrewchen 0 0
VC Adventure 0 0
Elad Blog | Substack 0 0
AVC 0 0
Above the Crowd 0 0
Entrepreneur Ride Along 107 8
r/SideProject - A community for sharing side projects 420 23
Future(s) Studies 946 17
Artificial Intelligence (AI) 262 14
Software As a Service Companies — The Future Of Tech Businesses 704 27
Investing In AI 0 0
Big Technology 0 0
The Gradient 0 0
Import AI 0 0
Sam Altman 0 0
The community for ventures designed to scale rapidly | Read our rules before posting ❤️ 107 3
Co-Founder: Find Your Co-Founder Here 0 0
Entrepreneur 57 1
Naval 1 0
Machine Learning 16 2
Deep Learning 22 4
Natural Language Processing 2 1
Venture capital news and articles, for the VC industry 0 0
Newcomer 0 0
Jerry Liu 6 4
Harrison Chase 0 0
Cristóbal Valenzuela 4 2
Amjad Masad 3 1
Arthur Mensch 0 0
clem 🤗 3 2
Aidan Gomez 0 0
Kanjun 🐙 0 0
Suhail 18 2
Guillaume Lample @ NeurIPS 2024 0 0
Clouded Judgement 0 0
Bindu Reddy 5 5
Parag Agrawal 0 0
Harry Stebbings 0 0
Keith Rabois 0 0
Fred Wilson 0 0
Brad Feld 1 0
Exponential View 2 2
The Pragmatic Engineer 0 0
Latent.Space 1 1
Mark Suster 0 0
Benedict Evans 0 0
Allie K. Miller 0 0
Elizabeth Yin 💛 0 0
Roelof Botha 0 0
Andrew Reed 0 0
Luciana Lixandru 0 0
The Pragmatic Engineer 0 0
Elad Gil 0 0
Nathan Benaich 0 0
sarah guo 4 0
@jason 12 0
Vinod Khosla 0 0
Daniel Gross 0 0
Ann Miura-Ko 🦖 0 0
Mike Volpi 0 0
Aravind Srinivas 0 0
Ajay Agarwal 0 0
Leo Polovets 0 0
David Sacks 0 0
Lenny's Newsletter 0 0
Interconnects 0 0
Not Boring by Packy McCormick 0 0
Marc Andreessen 🇺🇸 0 0
Chris Dixon 0 0
Sriram Krishnan 1 1
a16z 1 1
benahorowitz.eth 0 0
martin_casado 0 0
andrew chen 1 1
Scott Kupor 1 0
David Ulevitch 🇺🇸 0 0
Dalton Caldwell 0 0
Y Combinator 0 0
Jessica Livingston 0 0
Paul Graham 9 1
Invest Like The Best 0 0
Garry Tan 0 0
Michael Seibel 0 0
Sam Altman 0 0
TechCrunch 0 0
Plug and Play Tech Center 0 0
No Priors: AI, Machine Learning, Tech, & Startups 0 0
Lex Fridman 0 0
Lightspeed Venture Partners 0 0
500 Global 0 0
Google for Startups 0 0
ThisWeekinStartups 0 0
Two Minute Papers 0 0
My First Million 0 0
Lenny's Podcast 0 0
All-In Podcast 0 0
Garry Tan 0 0
Y Combinator 0 0
Acquired 0 0
Foundation Capital 0 0
20VC with Harry Stebbings 0 0
Sequoia Capital 0 0
Greylock 0 0
Stanford eCorner 0 0
a16z 0 0
Jeremy Howard 0 0
Aravind Srinivas 0 0
Cassie Kozyrkov 0 0
Andrej Karpathy 0 0
Alexandr Wang 0 0
Naval Ravikant 0 0
Clément Delangue 0 0
Elad Gil 0 0
Fei-Fei Li 0 0
Andrew Ng 0 0
Demis Hassabis 0 0
Sam Altman 0 0
Yann LeCun 0 0