ZeroNoise Logo zeronoise
Post
Nvidia–Poolside’s Reported Bet on the Open-Weight Operating Layer
17 hours ago
6 min read
1800 docs
A reported Nvidia–Poolside transaction leads a period in which early teams are productizing agent memory, coordination, and security while inference costs, vertical adoption, and incumbent distribution become the sharper investment questions.

1. Funding & Deals

Nvidia is reportedly combining capital, talent, and open-weight model development in one Poolside transaction. A post citing the WSJ says Nvidia plans to spend $6 billion on a powerful open-weight model, license Poolside’s technology, bring more than 100 Poolside employees into Nemotron, and invest another $1 billion in Poolside at a $12 billion pre-money valuation. The stated target is to challenge DeepSeek and Kimi while competing with OpenAI and Anthropic.

This is strategic corporate financing rather than a clean seed or Series A comparable. The diligence question is whether the combination of Poolside technology, transferred talent, and Nvidia’s model-building program produces a durable open-weight advantage; confirm the reported terms before underwriting.

Stripe’s reported OpenRouter purchase makes routing an exit thesis—but not a settled one. SaaStr’s 20VC recap says Stripe paid around $7 billion for OpenRouter four months after its $1.3 billion round, describing it as a leading LLM-routing layer. The same analysis says enterprises may prefer a few models and in-house routing, and gives only a one-in-three chance of a standalone routing business versus absorption into Stripe infrastructure. For early-stage investors, routing needs defensible control of model flow or a distribution advantage; a model menu alone is unlikely to be enough.

2. Emerging Teams

A repeat founder is testing persistent roles as the company operating system. A developer and founder with more than 10 years of experience who previously ran a SaaS business to roughly $5 million ARR says a new company is seeing “real traction.” He gives persistent agents CEO and CMO roles, feeds them company context, lets them disagree, makes the final decision himself, and then has agents handle execution, research, and testing. He says the value is memory that carries decisions and observed results forward, and forces agents to distinguish BUILT, VERIFIED, NOT VERIFIED, BROKEN, and UNKNOWN. Treat this as an operating-model signal rather than an underwriting datapoint.

Intimassy shows how basic distribution fixes can unlock early monetization. Its 11-year software-engineer founder says Reddit feedback led to a proper domain, tripling organic search visits, and an iOS launch that produced three paying users from 30 organic downloads on day one. The post headline reports 1,500 daily users and about $100 per day. The signal is discoverability and platform coverage—not another model feature; retention, paid acquisition, and durability remain unshown.

MUON is an early bet on coordination as an agent primitive. A research engineer at a YC startup built an open-source desktop app, MCP integration, and CLI that connect Claude Code, Cursor, Codex, and OpenCode into a shared memory and coordination graph. The author says it is still early, has broken features, and is being dogfooded on itself; its Polyform Noncommercial license permits personal or day-job use but excludes a company-wide shared brain without an enterprise arrangement. The category signal is stronger than the current traction evidence.

3. AI & Tech Breakthroughs

ShardFlow attacks inter-region inference latency with speculative decoding. The framework splits HuggingFace transformers across multiple GPU machines; in a two-T4 setup across GCP regions with roughly 86 ms round-trip latency, K=8 drafting commits 4.07 tokens per round trip instead of one. On Qwen2.5-7B, the builder reports 4.92 TPS for the non-speculative baseline versus 28.10 peak and 20.31 average TPS with a neural drafter and CUDA Graphs; graph capture reduced draft latency from 112 ms to 25 ms. These are builder-reported results, but reproducibility would make distributed inference over public WAN a meaningful deployment option.

Agentic document work is converging on retrieval first, vision second. Jerry Liu describes a two-pass pattern: a cheap open-source parser scans tens to thousands of files, then a just-in-time VLM screenshots and dissects only the relevant pages. That avoids running expensive VLM OCR over entire file dumps, while exposing current weaknesses in grounding, parser versatility, and cost.

Ship Safe packages the agent attack surface into a local developer tool. The MIT-licensed scanner checks for prompt injection and agent hijacking, dangerous MCP configurations and permissions, secrets, CI/CD and supply-chain risks, auth/API vulnerabilities, and RAG or memory poisoning. It runs locally with npx ship-safe; its agent workflow proposes a fix, shows the diff, asks for approval, applies it, and verifies the result. The project is still seeking security-tooling feedback, so coverage and false-positive economics are unproven.

4. Market Signals

Codex adoption is spreading beyond tech into functions with domain-specific workflows. a16z reports that its fastest-growing Codex adopter categories since February are legal at 108x, sales and recruiting at 41x each, marketing at 26x, and healthcare at 24x. This is a vendor-reported chart, so use it as a directional vertical-discovery signal rather than market-share data.

Inference price competition is an adoption tailwind and an application-margin trap. Suhail says Chinese model subsidization at the inference layer is helping new agentic coding products compete with Anthropic and OpenAI, and calls the trend healthy. A separate SaaS discussion argues that every AI action adds variable cost: raising prices causes churn, usage caps frustrate customers, and routing to cheaper models only partly closes the gap—quietly turning conventional SaaS into usage-based pricing. Underwrite token cost, pricing architecture, and gross margin together.

Agents are putting systems of record and link economics under pressure. Garry Tan predicts that systems of record will need to become AI harnesses or face replacement by agents. Separately, a current post says UK publishers asked the CMA to keep ChatGPT and Perplexity off Google’s default-search choice screen because chat answers generate no click-through or referral traffic; the CMA has yet to decide whether chatbots count as search engines. The practical risk is that incumbents retain data while losing the interaction and distribution layer.

Governance demand is real, but the headline statistic is not clean enough to underwrite. A post claims that 78% of organizations had not taken meaningful AI-compliance steps despite deploying agents on sensitive data. A commenter says that figure conflates at least three different measurements, while agreeing that PII leakage and prompt injection receive less attention than visible accuracy failures. The practical control discussed is explicit human confirmation before sensitive execution rather than trusting a model’s silent “processed successfully” claim.

5. Worth Your Time

  • Read — Jerry Liu’s two-pass RAG thread. The clearest current architecture note on cheap broad parsing followed by just-in-time VLM inspection, including the grounding and tool-quality gaps.

  • Inspect — ShardFlow. The repository and implementation notes behind the WAN-inference benchmark, including the Rust relay, KV-cache handling, and model slicing.

  • Read — The Speed of Thought: What If Intelligence Has a Universal Limit?. A contrarian capital-allocation frame arguing that frontier gains are roughly logarithmic in compute and that the next value may accrue more to applied AI, tooling, integration, and cost reduction than to recursive self-improvement.

  • Read — 20VC x SaaStr’s AI deal analysis. Useful valuation framing: the article contrasts systems-of-record software at roughly 5.3x revenue and slower or non-SOR software near 2.7x with OpenRouter near 70x trailing revenue.

Nvidia–Poolside’s Reported Bet on the Open-Weight Operating Layer
David Ulevitch 🇺🇸

a16z GP David Ulevitch amplified a story about Flock surveillance cameras: an Illinois state senator who opposed Flock on privacy grounds reversed her stance after a bullet pierced her son's bedroom wall; police used a single license plate captured by Flock to identify and arrest the suspect within hours, and the police chief said the case could not have been solved that fast without the cameras . Ulevitch commented "Many such cases!!!" , signaling a pattern of lawmaker opposition yielding to demonstrated success — a positive market signal for public-safety surveillance technology.

An Illinois state senator fought Flock on privacy grounds. Then a bullet came through her son's bedroom wall and nearly killed him. Polic… Many such cases!!! [https://x.com/austinjustice/status/2091568688738501020](https://x.com/austinjustice/status/2091568688738501020)
Paul Graham

Paul Graham observes that LinkedIn users have been editing their work histories to add "AI" and remove "DEI" , a possible signal of shifting professional branding that investors may want to weigh when evaluating founder resumes.

LinkedIn users have been editing their work histories to add "AI" and remove "DEI". ![](https://pbs.twimg.com/media/HQbR74Xa8AA76pe.jpg)
Paul Graham

Paul Graham, asked what he'd do if he were 17, said he'd learn to build LLMs from scratch and train models as powerful as possible with whatever hardware he could access . He explicitly said he would not try to start a startup immediately, instead building a knowledge foundation first, arguing that far better startup ideas grow out of deeply understanding LLMs than from acting on what he knew at 17 . The thread signals his view that the strongest future founders will come from hands-on, deep LLM expertise rather than early startup attempts.

Someone asked what I'd do if I were 17. I'd learn how to build LLMs from scratch, and then train ones as powerful as I could with whateve… Notice that what I would not do is try to start a startup. Instead I'd build the foundation of knowledge to base a startup on later. Way …
martin_casado

Martin Casado, a16z GP, says he was given access to a new model that he expects to be one of the most significant releases of the year, calling it potentially "the most?" significant drop this year; he shares no details about the model, its builder, or timing beyond "this year" .

Mind blown from a new model I just got access to. I think this will be one of the most (the most?) significant drops this year. Excited. …
Garry Tan

YC CEO Garry Tan predicts systems of record will need to become "AI harnesses" or face replacement by agents .

Prediction: systems of record will need to become AI harnesses or face replacement by agents
Nathan Benaich

Nvidia is reportedly spending $6B to build one of the world's most powerful open-weight AI models, and per WSJ will license Poolside's technology and bring 100+ of its employees into the Nemotron project . Nvidia is also investing another $1B in Poolside at a $12B pre-money valuation . The goal is to challenge Chinese open-weight leaders such as DeepSeek and Kimi while competing directly with US frontier labs including OpenAI and Anthropic .

Nvidia is reportedly spending $6 billion to build one of the world’s most powerful open-weight AI models. According to the WSJ, Nvidia wi…
Nathan Benaich

AI investor Nathan Benaich pushed back on fears that AI has made markets too expensive, arguing the global growth outlook would be dramatically worse without AI and calling AI "the only major new growth engine in sight" and "so early still" . He also shared an FT piece on the US widening its AI-driven investment gap with Europe .

everyone loves writing about whether ai has made markets too expensive but no one gives airtime to the reality that the global growth out… [https://www.ft.com/content/77b94c4a-4b4b-4983-9138-7db6926150f4?shareType=nongift](https://www.ft.com/content/77b94c4a-4b4b-4983-9138-7d…
Suhail

Suhail (former Mixpanel CEO) asked whether anyone has gotten DeepSeek v4 flash 0731 working with 'miles' for training , said he would do it himself , and followed with 'Victory.' in a follow-up post .

Has anyone gotten deepseek v4 flash 0731 to work with miles for training? ok i shall do it myself Victory. ![](https://pbs.twimg.com/media/HQcTgc0agAA1Ve4.jpg)
Vinod Khosla

Vinod Khosla endorsed wajo_ai's agent demo, saying "Really the way agents need to work!" in response to a demo by @shivanipod . The demo shows wajo_ai agents with "a phone, email, phone number, credit card and their own identity built to be secure for their bigs" , with product link http://wajo.ai. This signals an emerging theme of autonomous AI agents possessing independent communication and payment infrastructure.

Really the way agents need to work! [https://x.com/shivanipod/status/2091394612862890181](https://x.com/shivanipod/status/209139461286289… Hear my [@wajo_ai](https://x.com/wajo_ai) agent save me oxygen. [http://wajo.ai](http://wajo.ai) Agents with a phone, email, phone number…
The community for ventures designed to scale rapidly | Read our rules before posting ❤️

A commenter citing Carta data says the majority of seed rounds still use SAFEs, with priced rounds making up about one-third .

Stage benchmarks are inflating: Series A reportedly used to require ~$1M ARR but 'seemingly' now needs $2M-3M .

A commenter flags that rapid AI-driven ARR growth is 'creating all kinds of chaos in VC' and expects new stage buzzwords to follow .

Investor-sentiment signals on pre-seed: an angel says pre-seed is now used to get into high-pedigree teams early on the best terms 'while offering bad terms' and advises founders to bootstrap instead; another commenter says VCs use pre-seed to 'warehouse bets that are too early for even their early stage fund' .

Counter-view: pre-seed can benefit founders on dilution and is most appropriate for hardtech/deeptech/biotech where IP is developed and derisked inside the company, and pre-seed backing strengthens leverage in university IP and grant negotiations .

Yeah can go either way. Looking at Carta data it's majority of Seed round are still SAFEs, but priced makes up about 1/3. Guess it really… Seed was the original. We then got investment inflation so they had to add a prefix hence pre-seed. Now that landscape is also changing a… Ehh, when I started (before the dotcom boom) the first round was your Series A. Seed rounds were post dotcom bust. Also, the rapid ARR gr… The original pre-seed was friends & family. Pre-seed is used now to invest in high potential teams (with pedigrees) and get the best term… Not unpopular at all. Pre-seed is just 'we haven't figured out product-market fit yet but we're calling it a round so it sounds legit.' V… Couple of counter arguments. In some proper settings pre-seed staging can be good for founders in terms of dilution. Most pre-seed should…
Suhail

Tech founder Suhail (@Suhail) argues that AI that improves AI — inference systems, training models, etc. — is the only sensible benchmark moving forward; all other work is downstream and possibly disrupted .

AI that makes AI better (inference systems, training models, etc) is the only sensible benchmark to me moving forward. All other work is …
Bindu Reddy

Abacus AI CEO Bindu Reddy expects a very large AI breakthrough within the next few weeks , citing AI acceleration and that "multiple stealth labs are making very meaningful progress" . She also signals that a small continually learning model may be imminent .

Excepting a very huge AI breakthrough in the next few weeks AI is accelerating at an insane pass and multiple stealth labs are making ver…
a16z

a16z reports the fastest-growing OpenAI Codex adopters since February are outside tech: Legal up 108x, Sales 41x, Recruiting 41x, Marketing 26x, Healthcare 24x . The data is presented in a16z's Charts of the Week edition .

AI power users are showing up outside tech. The fastest-growing Codex adopters since February: Legal: 108x Sales: 41x Recruiting: 41x Mar…
Bindu Reddy

Bindureddy (CEO of Abacus AI) reports that Anthropic will launch three new models in the coming days: Opus 5.1, Sonnet 5.1, and Fable 5.1 . He claims Opus 5.1 and Sonnet 5.1 will reverse regressions and strictly outperform Opus 4.8 and Sonnet 4.6 , and Fable 5.1 will top the leaderboards .

Multiple new Anthropic models dropping in the next few days - Opus 5.1 - Sonnet 5.1 - Fable 5.1 Opus and Sonnet will reverse the regressi…
Bindu Reddy

@bindureddy predicts OpenAI will cut GPT-5.6 SOL prices by 80% , which will skyrocket usage and adoption and wipe out competition from Anthropic and open-source AI .

PREDICTION- OpenAI will slash GPT 5.6 SOL’s prices by 80% This will sky rocket usage and adoption It will also wipe out competition from …
Harry Stebbings

Harry Stebbings (@HarryStebbings) named his top 5 founders ever met (out of 1,000 founder interviews): Alan Chang (Fuse Energy), Jack Zhang (Airwallex), the ClickHouse CEO, Li Qiao (Fireworks), and Max Junestrand (Legora) . He also noted he has invested in many of the founders he has interviewed .

I have interviewed 1,000 founders and had the fortune to invest in many of them. The top 5 founders that I have ever met (not in order): …
Suhail

Suhail (@Suhail, account described as past CEO/Founder @Mixpanel and starting something new) reports an AI-agent debugging workflow that saved him pain: when encountering a bug, search GitHub PRs and issues in parallel with another agent to find a reference solution . He observes that this makes GitHub likely the primary place agents are already communicating with one another . Signal: agent-driven development is increasingly anchored to public code platforms, relevant to agent infrastructure and developer-tooling investment theses.

Has saved me a lot of pain recently debugging issues: “Moving forward, when you encounter a bug or issue, search GitHub PRs, issues with …
Machine Learning
  • An r/MachineLearning user built and shared a minimal, educational implementation of SynthID-Text-style statistical watermarking for LLMs, motivated by Anthropic's announced plan to add watermarks to model responses; the watermark is a subtle statistical pattern introduced during token selection, not a visible message .
  • Skeptics argue reliable text watermarking is not possible: output can be rewritten by another non-watermarking LLM or translated to another language and back, destroying the statistical signature; the implementer, after reviewing Anthropic's planned algorithm, doubts its effectiveness .
  • Counter-evidence: a recent benchmark showed watermarking accuracy remains at 70% even after backtranslation using Chinese as an intermediate language .
  • Supporters argue watermarking is still useful for the ~99% of users who will not try to strip it and enables applications such as online AI-content filters .
Implementing Watermarking for Language Models [P] Reliable watermarking for linguistically based inferencing is not possible, imo. You can feed the output to multiple other llms that don'… Yeah I do agree with this sentiment. The main motivation for me was to try and understand what anthropic are planning on doing. And after… A recent benchmark showed that even after backtranslation using Chinese as intermediate language, watermarking accuracy remains at 70%. W… Your argument is that it won't work because it can be "removed" by taking extra steps. Ok, sure. You and others are missing the point tha…
Machine Learning

ShardFlow, a distributed LLM inference framework built over the past few months (repo: rautaditya2606/Shardflow), splits any HuggingFace transformer across N GPU machines and uses neural speculative decoding to overcome WAN latency . In benchmarks on two T4 nodes in separate GCP regions (~86ms RTT via an AWS EC2 TCP relay), Qwen2.5-7B achieved 4.92 TPS non-speculative baseline, 14.3 TPS peak with a neural drafter, and 28.10 TPS peak / 20.31 TPS avg with CUDA Graphs on the drafter; Qwen2.5-14B with NF4 4-bit quant hit 14.43 TPS avg on the same two nodes . The key insight is that speculative decoding turns WAN latency into a per-round rather than per-token cost: with K=8 drafting, ~4.07 tokens are committed per round trip vs 1 . Capturing the full 0.5B draft forward pass as a CUDA Graph and replaying it with one driver call cut draft latency from 112ms to 25ms, versus ~1500 kernel launches per round that left the GPU idle 65% of the time . Additional stack elements include a zero-copy Rust TCP relay, StaticCache with in-place KV rewind for graph compatibility, and meta-device model slicing to avoid loading 15GB into CPU RAM . Repo: https://github.com/rautaditya2606/Shardflow.

28 TPS on Qwen2.5-7B across two separate cloud regions over public WAN using speculative decoding + CUDA Graphs [P]
Investing In AI

Investing in AI argues intelligence has a hard physical ceiling and human brains may already be near it — evolution landed on a ~20-watt architecture whose synaptic operations sit within an order of magnitude of thermodynamic limits — so unbounded superintelligence is an unlikely premise .

Empirical support cited: frontier model capability gains are "roughly logarithmic in compute — each increment of performance costs exponentially more," the signature of an asymptote, not an explosion, mirroring chess engines' diminishing post-1997 gains .

Investment implication: capital allocated on the assumption of recursive self-improvement is mispriced — the author projects another decade of AI improvements, a visible bend in the 2030s, and remaining gains from tooling, integration, and cost reduction rather than raw cognitive horsepower; this favors applied AI (where NYC is expected to beat San Francisco as the future AI hub) over west-coast AGI bets .

Cautionary flag: many hard problems (climate, aging, interstellar travel) are not intelligence-limited — climate is a coordination problem, chaotic systems have Lyapunov horizons, Wolfram's computational irreducibility applies, and experiments run at the speed of the physical world — so intelligence is a lever, not a wish .

The Speed of Thought: What If Intelligence Has a Universal Limit?