ZeroNoise Logo zeronoise
Post
AI Infrastructure’s Next Test: Efficient Inference, Grounded Agents
8 hours ago
6 min read
2377 docs
A concise radar on the split between inference capacity and demand, the rise of grounded document agents, and early AI-native operating models.

1. Funding & Deals

The period’s clearest deal signal is a reported Stripe acquisition of OpenRouter for more than $7B—but it is a strategic datapoint, not an early-stage financing comp. A current-period venture post relaying Bloomberg says the deal followed a $1.3B round roughly 82 days earlier; it describes OpenRouter as founded in 2023, with 8M users, more than $100M in annualized inference volume, and backing from Sequoia and a16z. Its product is a routing layer: one API in front of 400-plus models, rather than a model lab. The post’s thesis is that aggregation, routing, price discovery, and settlement may capture durable AI value as models commoditize.

Do not use the reported 5x-in-82-days multiple for seed underwriting. The useful question is whether this represents genuine repricing of the routing layer or a strategic premium from Stripe; the source leaves that unresolved.

2. Emerging Teams

An AI-native Canadian law firm combines unusually strong founder–problem fit with an early, self-reported traction signal. Its founder has 25 years of contracting experience, including a decade as general counsel at a large BC private company and a prior CEO role in international aviation. The firm uses flat fees and a 48-hour target, has AI perform the first pass from its own playbooks, versions every draft with an audit trail, keeps client data out of model training, and retains founder review and signature. Demand was described as constant, with the first month tracking toward low-to-mid five figures.

This is a service-delivery model rather than legal-copilot SaaS: the founder says the platform was built for the firm’s own use, not to sell to other law firms, with startups and SMBs as the target market. The diligence question is whether the flat-fee, rapid-turnaround workflow remains reliable as volume rises without weakening legal review.

Factory Brain is a small but clear infrastructure thesis around “learn once, reuse later.” The private-preview project stores business terminology, schema relationships, KPI definitions, validated questions and SQL, corrections, and follow-up context so that an LLM is reserved for new or genuinely complex questions. It is also being designed around RLS/CLS, governance, and keeping business data in the customer environment. The founder is explicitly asking for architecture criticism rather than presenting traction; the current product signal is the attempt to make semantic memory and query reuse part of the analytics stack, not another chat-with-data wrapper.

3. AI & Tech Breakthroughs

Document agents are being evaluated against completeness and evidence, not just plausible answers. LlamaIndex’s new ExtractBench is an open benchmark covering 370 enterprise documents, 4,869 pages, eight business domains, 67 document types, and 14 systems; its ground truth combines cross-model agreement with human adjudication, synthetic long lists, and manually checked forms. The vendor reports that Agentic Plus leads at 95.6% value F1, with the best grounding scores at 8.1¢ per page.

The more important result for diligence is the failure profile: LlamaIndex reports that commercial VLMs fall below 35% recall on documents longer than 50 pages, while coding agents return no evidence by default; it says Agentic Plus holds 94.4% on the longest documents. Those are vendor-reported results, but the dataset, harness, and paper are public, so long-document recall and grounding can be independently tested rather than accepted from a demo.

Long-context efficiency still has an exact-retrieval problem in genomics. A developer’s 1M-token DNA experiment reports roughly 25% needle-in-a-haystack recall—chance level for a four-token DNA vocabulary—and similarly poor 25–27% results for HyenaDNA, versus 50–60% recall for a much shorter 16K context. The accompanying discussion points toward selective state-space models or hybrid architectures that retain occasional uncompressed attention. This is a self-reported research thread, not a validated benchmark, but it is a useful warning against treating million-token context as equivalent to million-token memory.

Text provenance also remains brittle. A developer reports that, across nearly 300 watermark tests, inserting invisible Unicode variation selectors into about 30% of characters reduced a watermark score from 45 to below 1 in every one of ten trials; the same post says code is often barely watermarked because its token distribution is low-entropy. Treat this as an attack report requiring replication, not as a general defeat of all watermarking schemes.

4. Market Signals

The inference-glut question is becoming a two-market asset-quality problem. An Investing in AI analysis expects AI data-center power to remain tight through 2027 and loosen by mid-2028 without an aggregate glut, while warning that the average hides a split between new facilities able to host 130–600 kW liquid-cooled racks and older capacity. It estimates 120 GW of scheduled capacity through August 2028, but only 50–60% realization because of transformers, turbines, and interconnection queues; the resulting AI fleet would rise from 35.6 GW to 78.5 GW into a market described as 94% full.

Demand is not automatically falling with inference prices: the essay says token prices fell from about $20 to $0.07 per million while Google’s reported monthly volume rose 330-fold, with measured elasticity of −1.03. Its model also estimates that raising reasoning queries from 12% to 45% lifts average energy per query 2.6x. The positioning implication is to underwrite workload mix and routing, not simply installed megawatts.

The analysis projects roughly 4 GW of older provisioned capacity becoming effectively stranded, with price pressure beginning around Q2 2027—well before aggregate vacancy shows weakness. Its recommended monitors are hyperscaler capex split between training and serving, H100-class spot GPU-hour pricing, preleasing on capacity under construction, and interconnection or turbine-order cancellations. The main caveats are that the realization haircut is a judgment call and token growth leans heavily on unaudited Google figures.

AI adoption is asymmetric by workflow. Andrew Chen’s framing is that workplace AI succeeds where it compresses patterned drudgery—forms, process steps, boilerplate, and updates—while consumer products need novelty, parasocial connection, and authenticity. He extends the same problem to sales and marketing: uniform AI messaging is easy to generate but fails in adversarial settings where the message must be fresh and differentiated.

The “agent test” is becoming a product and valuation filter. Jason Lemkin says SaaStr’s agents built an ad-creative operation without Canva or Notion ever entering the workflow, despite SaaStr having paid for and liked both products. He argues that no-code products built around replacing a missing specialist are exposed to native AI substitution, while Gartner data suggests fewer than 10% of enterprises have successfully deployed an agentic application. His practical test is to give an agent the job a product does without instructing it to use that product, then mark valuations from current growth rather than legacy financing marks.

5. Worth Your Time

  • Watch — Lenny’s Podcast: OpenAI’s Head of Product Design, Ian Silber. Silber describes a future ChatGPT as a proactive, voice-rich universal input that decides whether to answer or act, hides model and mode choices from most users, and supports durable repeatable workflows rather than one-off chats.
  • Read — ExtractBench. Use it as a concrete evaluation starting point for long-document completeness, grounding, perception failures, and cost—not just clean-PDF accuracy.

  • Read — Is There An Inference Glut Coming?. The useful parts are the distinction between modern and stranded capacity and the monitorables that could reveal softness before aggregate utilization does.

AI Infrastructure’s Next Test: Efficient Inference, Grounded Agents
Research extraction

Direct answer

Source is the vendor's announcement blog alone (linked HF/GitHub/arXiv are not part of the bundle). The post reports Agentic Plus at 95.6% value F1 / 8.1¢ per page, first overall and best grounding at both levels, on an open 370-document/4,869-page benchmark with ground truth built three ways . Several methodology specifics cannot be verified from the text: the three reported segments sum to 159 of 370 documents, "accuracy" and "F1" are used interchangeably, and cost/grounding metric definitions are absent .

Findings

  1. Benchmark scope and openness. Claimed to be open and reproducible; 370 enterprise documents (4,869 pages), 8 business domains, 67 document types, each type with its own schema; claimed to be the only benchmark jointly evaluating long-record completeness, real scans and handwriting, word- and page-level grounding, and measured cost; 14 systems evaluated; Agentic Plus is the new LlamaExtract tier that debuts at the top of the board . Dataset, code, and paper links are listed .
  2. Methodology: evaluation axes. Documents are tagged along five independent axes — task challenge, perception, table structure, document length, business domain — and every score is reported per axis . A challenge taxonomy is presented in tables (long-list completeness, needle-in-haystack, dense documents, perception, table structure, grounding and output trust) with a covered/partial/absent legend . Caveat: in the provided markdown the table cells are empty, so which challenges are marked covered/partial/absent is not legible from the text .
  3. Methodology: ground truth. Built three ways: real documents where systems from different model families run the same schema, agreed values become candidate truth and disagreements are resolved by a person; synthetic long lists built data-first so every value and box is known before rendering and completeness can be scored exactly at thousands of rows; and 169 regulatory/tax forms where a reviewer checks every field by hand and places bounding boxes for 84% of fields . Every document's ground truth comes from one of these three documented pipelines and every system gets identical inputs; no extractor's output is trusted on its own, ours included .
  4. Reported overall results. Mean value accuracy across all 370 documents is reported; Agentic Plus posts 95.6% value F1 at 8.1¢ per page with the best grounding scores at both levels . Full overall board: LlamaExtract Agentic Plus 95.6%/8.1¢, Agentic 89.5%/3.1¢, Cost Effective 86.8%/1.0¢; Codex (GPT-5.5) 93.6%/27.8¢, Claude Code (Opus 4.8) 87.1%/16.2¢; Reducto Deep Extract 90.4%/34.4¢, Extend (Max Context) 86.3%/10.0¢, Datalab 64.5%/3.5¢; Qwen3.6 35B 87.3%, Gemini 3.5 Flash 79.8%/1.0¢, Lift (9B OSS) 77.3%, GPT-5.4 Nano 74.9%/0.21¢, Gemma4 26B 66.2%, NuExtract3 47.9% .
  5. Reported segment results. Dense multi-domain business docs (72 docs, 22 types, ≤10 pages): Agentic Plus 96.6% at 8.3¢/page, Codex 95.7%, Reducto 94.2% . Multi-section reports (70 docs, 26 types, 11–50 pages): Agentic Plus 93.3% at 7.7¢/page, Codex 91.2% . Repeated records at extreme scale — up to 26,725 rows (17 docs, 12 types): Agentic Plus 94.4% at 7.5¢/page, Reducto 92.0%, Claude Code 88.1% .
  6. Failure-case context. Short documents make everyone look good: eight of fourteen systems score above 90% on documents under ten pages, and only price separates them . On documents longer than fifty pages, every commercial VLM falls below 35% recall while retaining high precision; Agentic Plus is said to barely move, holding 94.4% on the longest documents . The two frontier coding agents have opposite perception blind spots (one fails rotated pages, the other fails scans and handwriting); Agentic Plus is claimed to be the only system above 93% on all three perception challenges . Grounding is called an open challenge: VLMs and coding agents return no evidence by default and score zero at both grounding levels; among systems that do return boxes, Agentic Plus leads at both levels . Cost framing: at a million pages a month, each cent per page is another $10,000, so a couple points of accuracy at four times the cost is a bad trade . Price does not predict accuracy across the measured systems; the accuracy leader charges less than a quarter as much as the closest peer, and Cost Effective comes within a point of Claude Code at a sixteenth of the rate .
  7. Reproducibility. Dataset, eval harness, and paper are open and schemas are frozen; clone and run commands are provided .

Conflicts / gaps / uncertainty

  • The three reported result subsets (72 + 70 + 17 = 159) do not account for the claimed 370-document corpus; the post does not state whether the segments are exhaustive or describe the remaining documents .
  • The multi-section subset is labeled 11–50 pages, so the claim about "documents longer than fifty pages" has no explicit segment mapping; the extreme-scale subset reports only "up to 26,725 rows" with no page counts, and the reported segment metric is not explicitly recall .
  • Metric terminology is inconsistent: "value F1" and "mean value accuracy" are both used, and neither F1 nor the two grounding levels (word vs page) is defined in the bundle .
  • Cost methodology is not described; the post only says 9 of 14 evaluated systems carry a measured price .
  • The coverage-matrix marks are not visible in the provided text, so the "only benchmark that jointly evaluates..." claim cannot be checked against the table matrix in this bundle .
ExtractBench: The Most Comprehensive Extraction Benchmark
Lenny's Podcast

Ian Silber, OpenAI's head of product design (3 years in role; previously Artifact, the Instagram founders' news app, and 8 years at Instagram), shared OpenAI product strategy and design-market signals on Lenny's Podcast .

  • OpenAI's ChatGPT/Codex vision is a "super app": one universal input that adapts to the context of your life, becomes proactive (ChatGPT Work already suggests actions from calendar/Slack), gains far richer voice interaction, and shifts from answering questions to doing tasks for you. The goal is durable, repeatable workflows and a simpler default experience where users don't think about models or modes . UI is also moving beyond pure chat — manipulable "writing blocks," tappable follow-up questions, and specialized image-generation tools .
  • OpenAI's design strategy rests on "capability overhang" across ChatGPT's ~1 billion monthly active users: keep a simple mainstream experience while experimenting with cutting-edge features (desktop app, Codex, ChatGPT Work) on power users first, then distill them into the main product over time .
  • Product-paradigm signal: the "blank box" chat interface is AI's central design problem — designing one interface that can shapeshift across the full use spectrum, from recipe ideas to automating farms, is "uncharted territory" .
  • Market signal: engineers are seeing 10–100x productivity gains from AI coding agents, while the design process remains messy and largely un-compressed; Lenny's tech-workforce sentiment survey finds designers and user researchers the most unhappy, overwhelmed, anxious, and least optimistic role in tech .
  • Startup team composition is shifting: founders now discuss ratios like 2 designers to 1 rock-solid engineer (vs. a historical ~1 designer per 15 engineers), and Silber advises new founders to hire the best generalists who move fluidly across design/PM/engineering as those roles blur .
  • AI as designer: "I think it already is an incredible product designer" — accessible to everyone, though not yet best-in-class at visual or interaction craft. Durable human value lies in understanding users, inventing new interaction paradigms (iPhone multitouch, Snapchat as examples with no training data), and holding a point of view . As anyone can build anything with AI, the "human element" of design becomes a primary differentiator and great companies are actively seeking designers .
  • OpenAI design hiring: no AI background required — they screen for curiosity, prototyping skill, and increasingly systems thinking (composable building blocks, per Notion's early model); Silber echoes "the model we have today is the worst the model will ever be," with new models shipping weekly/daily across labs .
  • Cautionary flag on AI hype: Silber endorses Elena Verna's "Stop the AI Confidence Theater" — most public AI success stories aren't actually working yet, and the honest position is that tools are changing too fast for anyone to have it figured out .
OpenAI’s Head of Design: This is the best time in history to be a designer | Ian Silber
David Sacks

Anthropic's Dario Amodei says his policy proposals are designed to slow frontier AI companies while helping smaller competitors: California's SB53 exempts companies below revenue/training-cost thresholds (he cites $500M for SB53), and testing he advocated at CAISI/White House is stricter for frontier models, differentially advantaging challengers including open-weights . He supports pre-deployment testing for frontier models and testing open-weights as they approach the frontier, plus Demis Hassabis's FINRA-like entity idea, and notes AI is structurally centralizing due to scaling laws and compute/chip concentration .

David Sacks counters that federal pre-approval ('DMV for AI') would create long queues, handicap the US vs China, and undermine Anthropic's own pricing power, which depends on staying ahead of open models; he warns the US could become an island of costly closed models while the rest of the world uses broader choice, and says Anthropic is on track to become one of the most valuable companies in history, able to shape rules while competitors wait . He argues such gatekeeping would reinforce centralization by putting capability decisions in a federal bureaucracy hand-in-glove with frontier labs , and alleges Anthropic hired senior Biden AI-policy officials and built aligned organizations to push its regulatory frameworks . He also cites Anthropic's AI-fear campaigns (50% entry-level job-loss claim; 60 Minutes 'blackmail' study) as narratives shaping regulation . Framing: Amodei believes frontier AI is too powerful to distribute; Sacks believes it is too powerful to centralize .

1/2 Thanks Gavin for an especially thoughtful exchange. I don't usually spend much time on social media but I wanted to engage here becau… Some thoughts on Dario’s post: 1. Dario does not actually address Gavin Baker’s account of what he said – something he could easily deny …
Paul Graham

Paul Graham, YC co-founder, advises founders to focus on building stuff and understanding their users rather than on getting into YC, as that is what gets them in .

If you want to get into YC, don't focus on getting into YC. Focus on building stuff and understanding your users. That's what gets you in…
Paul Graham

Paul Graham (@paulg) argues that rising economic inequality is driven less by tax policy changes and more by technology making it easier to start companies and letting them grow faster once started, linking to his essay at paulgraham.com/richnow.html . This frames tech-driven company formation and scaling as a structural driver of wealth concentration, a relevant theme for early-stage tech investing.

There's a pervasive myth that "tax policy changes" are the reason economic inequality is growing. Actually it's because technology makes …
a16z
  • Travis Kalanick is back with Atoms, an industrial AI company, in a fireside chat with Ben Horowitz and Erik Torenberg; he says he's been working the whole time, and the venture comes after "eight years, thousands of employees, multiple industries" and "the biggest check Ben Horowitz has ever written."
  • Atoms' framing: manufacturing, real estate, and logistics as "the CPU, storage, and network of the physical world"; the chat covers "bits, atoms, and the three computing primitives," whether a delivered meal can cost less than the grocery store, and Kalanick's view that "ride-sharing was the gold medal. Transport is full of silver" — signaling expansion beyond ride-hailing into physical-world infrastructure.
  • Kalanick's operating advice: "if it's getting easy, it's about to get really hard"; entrepreneurs on "easy street" are about to "get your ass whooped," so keep pushing while winning.
Travis Kalanick: "A lot of folks think I'm back. I've been working my ass off the whole time. I just haven't been talking about it." Eigh… Travis Kalanick says when it gets easy, it's time to push hard: "I always used to say, if it's getting easy, it's about to get really har…
@jason
  • Anthropic CEO Dario Amodei argues AI is structurally power-concentrating via scaling laws; open-weights models are not a sufficient fix because they shift concentration to compute/chip owners .
  • Anthropic explicitly supports regulations that exempt smaller AI companies (e.g., SB53's $500M threshold) and differential testing for frontier vs. off-frontier models, which "hurts the business interests of the frontier labs and helps challengers, including open-weights" .
  • Amodei endorses the reported federal approach of pre-deployment testing for frontier models (and open-weights near frontier) and Demis Hassabis's FINRA-like regulator idea, noting the industry has pivoted from opposing state rules to accepting a federal path .
1/2 Thanks Gavin for an especially thoughtful exchange. I don't usually spend much time on social media but I wanted to engage here becau…
andrew chen

Andrew Chen (@andrewchen) highlights an emerging wave of American open-weight AI models, thanking Meta, NVIDIA, Google, Thinking Machines, OpenAI, and Microsoft for leading the way . He names Muse, Nemotron, Gemma, Inkling, GPT-OSS, and Phi as "just the start" of this movement , signaling early-stage momentum and investor attention around open-weight AI from both major labs and the newer entrant Thinking Machines.

Thank you Meta, NVIDIA, Google, Thinking Machines, OpenAI, Microsoft, for leading the way in American open weight AI models 🇺🇸 Muse, Nemot…
andrew chen

@andrewchen (a16z) argues AI has exploded at work but is slow to disrupt consumer apps because work drudgery follows patterns AI can compress, while consumer attention demands novelty, parasocial validation, and authenticity — AI-generated "slop" fails there, e.g., an AI-distorted background ruins a video's authenticity in social, entertainment, dating . He sees similar problems in adversarial business activities like sales/marketing, where uniform AI cold emails are unwanted; fresh messaging requires going "outside the model" and is an "adversarial creativity" human-in-the-loop challenge, possibly solvable by models that observe real-time timelines (e.g., the past 24 hours of X) .

Quick observation on why AI has exploded at work, but has been slow to disrupt consumer apps: At work, AI follows patterns and gets rid o…
Aravind Srinivas

Gergely Orosz, a prominent engineer and former Perplexity advocate, publicly criticized Perplexity as 'disappointment after disappointment' over the last 6-12 months . Perplexity CEO Aravind Srinivas acknowledged the criticism, stating the reminder email did not go out to Orosz, he was refunded, and the company is 'upgrading our support across the board' .

Perplexity, the last 6-12 months, is disappointment after disappointment. I used to be a huge advocate for the service thanks to how good… You’re right [@GergelyOrosz](https://x.com/GergelyOrosz), we got this wrong. Reminder email didn’t go out to this user. He has been refun…
martin_casado

Martin Casado (a16z) endorsed a post by @davidsenra sharing Travis Kalanick's (@travisk) view on VC behavior: only 10% of VCs are able to "do no harm" and only 1% are actually helpful . Kalanick framed operators as grandmasters of chess and VCs as chess enthusiasts who check in quarterly with opinions, and said most VCs are like lions that instinctively attack a limping antelope—"Can't even fucking help it" .

“ A super high bar for a VC is: do no harm.” [@travisk](https://x.com/travisk) says only 10% of VCs are able to do no harm, and only 1% a… Haha so good. Also true. [https://x.com/davidsenra/status/2089081748600815844](https://x.com/davidsenra/status/2089081748600815844)
Harry Stebbings
  • Harry Stebbings ranks growth leader Matt Swulinski "easily top 3" (with Alex Schultz and Brian Hale); Swulinski scaled Wispr Flow to over $100M ARR via a UGC machine and scaled Superhuman from founder-led onboarding to a growth machine at $50M ARR.
  • For AI/SaaS, distribution is becoming a critical moat in a crowded AI market, so SaaS should apply the e-commerce playbook — every dollar tied to conversion, deploy UGC creators, constantly test creative, and diversify channels.
  • Paid acquisition is the fastest way to validate a PLG funnel: organic-only validation is too slow, while paid lets teams test positioning, messaging, and conversion within a single week.
  • Scaling to $10M ARR requires only three core channels — video intent on Meta, search intent on Google, lifecycle retention via email/SMS — rather than running ten channels poorly.
  • Scaling paid ads on Meta requires 400-500 new creative assets per month (via UGC revenue-share, agencies, internal teams) because creative increasingly acts as the targeting algorithm; Swulinski repeats the 400-500/month minimum or "you're going to get outcompeted," citing an example of 3-4 videos/week across five agencies.
  • Best growth leaders test spend incrementality by measuring spend elasticity against ARR growth and running strict holdout tests to avoid buying conversions that would happen organically.
  • Within three years, lean human teams could operate like boards of directors: 20% of time on strategy while autonomous AI agents handle 80% of operational execution.
  • Marketing teams should be systems thinkers building self-improving AI workflows for 10x personal leverage; repetitive manual marketing roles are becoming replaceable.
I have interviewed 100 of the best growth leaders in the world. [@MattSwulinski](https://x.com/MattSwulinski) is easily top 3. (alongside… "You probably need at least 400 to 500 new creatives a month. Otherwise, you're going to get outcompeted. Once they do that, they create …
Deep Learning
  • Developer "aloshdenny" released an open-source attack claiming to remove Claude's text watermark without rephrasing: repo https://github.com/aloshdenny/claude-awm and interactive demo https://aloshdenny.com/claude-awm/. The author built a custom Claude/OpenAI/Gemini text-watermark generator plus detector — described as relying on Tournament Sampling built upon standard Gumbel-max sampling — and ran attacks on gpt-oss-20b and Qwen outputs .
  • The only attack that beat detection 10/10 times was inserting invisible Unicode variation selectors into ~30% of characters, dropping the watermark score from 45 to under 1; it survives normalization because these are real, meaningful codepoints normalizers can't strip .
  • Common countermeasures failed across nearly 300 test runs: em-dash/hyphen swaps, markdown stripping, and AmE→BrE spelling didn't move detection; only deleting 40% of every word crossed the threshold, and it wrecks the text .
  • Code output is barely watermarked: watermark strength tracks how uncertain the model is about the next token, so low-entropy code often comes out effectively unwatermarked with zero attack .
I figured out a loophole to remove Claude watermark WITHOUT rephrasing
a16z

a16z spotlights neoclouds: companies that pivoted from crypto mining now own some of tech's hottest assets (power rights, data centers, GPUs) . CoreWeave, 25 quarters in, generates more quarterly revenue than Azure, AWS, or Google Cloud each did at 30 quarters , signaling the scale of this infrastructure trend; a16z links to its Charts of the Week for further data .

Right place, right time: many neoclouds spent years mining crypto, then AI showed up and turned their power rights, data centers, and GPU…
@jason

In a This Week in AI episode, Eragon's Josh Sirota argues that logging into a SaaS platform will soon feel like buying a CD instead of streaming — users want the output, not the interface . The panel debates whether software developers will disappear or transition into software managers, and whether agents still can't fake some human judgment . The episode raises whether Anthropic now watermarks every word Claude writes, asking if this solves the AI slop/Dead Internet problem or is a nuisance, and reacts to Zuckerberg's vision for AI abundance . Other segments cover Jason's "No Clankers" proposal, the "distillation double standard," "when you're behind, go open…," "everyone is vibecoding," and Texas freezing new data centers .

Software is dead?! Eragon’s Josh Sirota says that logging into a SaaS platform will soon feel like buying a CD instead of streaming your …
andrew chen

a16z's @andrewchen credits Meta, NVIDIA, Google, Thinking Machines, OpenAI, and Microsoft for leading American open-weight AI models, naming Muse, Nemotron, Gemma, Inkling, GPT-OSS, and Phi as the start of this trend . He also observes that many US AI critics come from countries with very few AI startups and models, framing the open-weight push as a US-led competitive advantage .

Thank you Meta, NVIDIA, Google, Thinking Machines, OpenAI, Microsoft, for leading the way in American open weight AI models 🇺🇸 Muse, Nemot… loving the X feature where you can click on peoples’ profiles and see what country they are from lotta US AI haters from countries with v…
Harry Stebbings
  • Harry Stebbings ranks Matt Swulinski as one of the top 3 growth leaders he has interviewed; Swulinski scaled Wispr Flow to over $100M ARR with a UGC machine and scaled Superhuman to $50M ARR .

  • Growth lessons from the episode: modern SaaS should use the e-commerce playbook (every dollar tied to conversion), because distribution is a critical moat in a crowded AI market ; paid acquisition is the fastest way to validate a PLG funnel, often within a week ; only three channels are needed to reach $10M ARR — video intent on Meta, search intent on Google, and lifecycle retention via email/SMS ; scaling paid ads requires 400–500 new creative assets per month on Meta to avoid audience fatigue ; top growth leaders run strict holdout tests and measure spend elasticity against ARR growth to confirm incremental revenue .

  • Prediction: within three years, lean human teams will operate like boards of directors — 20% strategy, 80% autonomous AI agents — and marketers who aren't systems thinkers become replaceable; teams should build self-improving AI workflows for 10x leverage .

  • For answer engine optimization in ChatGPT, Swulinski's top advice for founders is to focus on YouTube, Reddit, and the social narrative, since long-form YouTube reviews rank long-tail and are a "really high citation on ChatGPT" .

I have interviewed 100 of the best growth leaders in the world. [@MattSwulinski](https://x.com/MattSwulinski) is easily top 3. (alongside… How to crush answer engine optimization and rank [#1](https://x.com/hashtag/1) in ChatGPT "The most important thing is YouTube, Reddit, a…
Harrison Chase

@patrickc argues that agentic coding harnesses should not be primarily terminal-based: the terminal is great for quick, precise commands but has extremely low information density and minimal UI affordances, and TUIs are at most for occasional use; he draws an analogy to dynamic-language REPLs taking a long time to break out via Jupyter notebooks . In reply, LangChain's Harrison Chase (@hwchase17) explains how deepagents is architected: the agent loop runs separately from a "backend" that must expose filesystem-like read/write/edit operations (database, object storage, or real filesystem), with sandbox backends also exposing an "execute" command — separating "the brains from the hands" per Anthropic's managed-agents post — and it is built on LangGraph for deployment via MCP, A2A, and other standard endpoints . The same backend separation powers a local TUI coding experience (dcode) and a cloud setup via LangSmith deployments with sandboxes on Modal, Daytona, or E2B, where web and Slack frontends share one backend (open-swe) ; non-coding agents can use a "fake" backend for file/context interactions without a full sandbox, and a managed deepagents option is available .

I love agentic coding harnesses, but they shouldn't be primarily terminal-based. The terminal is great for quick and precise commands, bu… totally agree! here's how we architected deepagents to enable this deepagents runs connected to a "backend". this backend needs to expose…
The community for ventures designed to scale rapidly | Read our rules before posting ❤️
  • An NBER working paper, "Beyond Demo Day: Sorting and Value Added in Startup Accelerators," studying ~750,000 US startups across 329 accelerators finds YC historically generated extraordinary value but its estimated value-add fell dramatically by 2022, and that roughly 60–80% of accelerators appear worse than building without one . The post author shared the paper PDF (https://www.nber.org/system/files/working_papers/w35063/w35063.pdf), noting it is freely available on the NBER website .
  • Three compounding problems are hypothesized for YC's stumble: oversized batches diluting scarce partner/investor attention and bespoke introductions; a shift to younger founders with less industry experience just as domain judgment (insurance, defense, healthcare, institutional finance, manufacturing) became more valuable when anyone can ship quickly; and AI as a paradigm shift where YC's pattern recognition from Stripe, Airbnb, Dropbox, Coinbase and SaaS generations transfers poorly. YC's real moat — Bookface and the accumulated network — may amplify the decline through a feedback loop: weaker selection → fewer defining winners → weaker network → less value for future founders . By 2022 YC's incremental contribution looked "surprisingly small," with prestige "largely past glory" .
  • Commenters split on whether frontier LLMs commodity startups: one side says most Launch HN startups' products can be replicated by a frontier LLM in a few days with an experienced engineer guiding it, and open-source clones beat stealth startups to market . The counter: no founder can replicate scaled SaaS like Atlassian (pioneer of network-effect flywheel growth) or Figma (real-time collaborative design canvases) with Claude Code — marketing, positioning, distribution, and deep engineering require industrial experience . Replies assert a frontier LLM can build an Atlassian or Figma clone , and note the same replication argument was made against Dropbox, Airbnb, and Calendly .
  • A commenter argues the YC model is "broken": YC burned hundreds of millions on wrappers, concluded wrappers are worthless, and is pivoting to funding young founders in domain-expertise fields (defense, biotech, law). But domain experts can now use frontier models to do technical work themselves, so the technical edge of the young Stanford-wunderkind archetype is gone; short-term, domain experts get a "brief moment," while long-term OpenAI and Anthropic "own it all" .
  • One view holds that with product-building cheap, distribution is "the only thing that matters," making accelerator networks stronger than ever for both B2B and B2C; YC's problem is diluting its batches by funding too many teams .
  • A counter-view: accelerators don't help much — they mostly filter. Founder personality, team, and marketing account for ~80% of success, and the failure rate is still ~90%, exactly what it was in 1996; second to start almost always beats first, because ideas are worthless and execution decides .
  • In enterprise healthcare software, distribution isn't the real issue: founders fail because they don't understand clinical workflows, their product fails on edge cases or can't integrate with incumbent EHRs, and idealized sales get torn apart pre- and post-sale. CEO access is easy for VCs to supply; connecting with skilled end users and admins who know the systems is the hard part. YC historically lacks a strong healthcare track record .
YC may have already peaked, and the data is starting to show their stumble (I will not promote) Here's the paper: [https://www.nber.org/system/files/working\_papers/w35063/w35063.pdf](https://www.nber.org/system/files/working_papers/… A lot of the Launch HN threads I see are for tech startups whose product can be replicated by a frontier LLM in a few days guided by an e… Unless you have spent decades in tech and across many parts of a scaled company, I can assure you that no founder will ever be able to re… If you don't think a frontier LLM can build an Atlassian or Figma clone then, sincerely, shut the fuck up Thank you! You could say the same (replication) thing about Dropbox or Airbnb or Calendly, yet here we are YC is dead. They’ve finally realized (after burning hundreds of millions across multiple years) that wrappers are worthless and are now p… i would argue otherwise. building a product today is cheap. distribution is the only thing that matters. and the network that such accele… We've known for a LONG time that accelerators don't actually help much. They more so filter. What's been studied is that the personalitie… That's a typical complaint of first time founders, especially when things don't work out for them. Ideas aren't worth anything. None of t… In enterprise healthcare software, distribution isnt actually the issue, the problem is that far too many builders dont understand the re…
@jason

@Jason posed a speculative question asking why Stripe would buy OpenRouter, with an attached video .

Why would stripe buy openrouter? Why would they do that? [![Video](https://pbs.twimg.com/tweet_video_thumb/HP4WhB7WYAEXugp.jpg)](https://…