ZeroNoise Logo zeronoise
Post
AI’s New Battleground: Proof, Price, and the Runtime
17 hours ago
6 min read
1998 docs
Current signals move value away from generic model access and toward independent evaluation, cheaper and faster execution, workflow-specific distribution, and infrastructure that can run reliably in the background.

1. Funding & Deals

Vals is turning model evaluation into a financed infrastructure layer. Vals announced a $40M Series A at a $400M valuation led by a16z, with existing investors 8VC, Pear VC, and Bloomberg Beta, plus HRT Ventures and Next Ladder. The round accompanied the general availability of Vals Smith, new risk benchmarks including the RSI Index with CoreWeave and ReverseEngBench with Columbia, Tufts, UC Berkeley, and UCLA, and broader coverage in Vals Index 2.0. Vals says revenue is up 8x versus all of 2025, customers doubled, and the team tripled in six months; a16z names Rayan Krishnan and Langston Nashold as the founders it is backing.

The thesis is unusually clear: domain experts turn real workflows into benchmarks and automated graders, replacing leaderboard theater with evidence that a model can do useful work. The diligence question is whether Vals becomes a recurring decision layer for model buyers rather than a catalog of benchmark products.

Uplift is an early hard-tech bet on manufacturing economics in housing. The company launched with a $7M seed led by a16z. Its co-founders say they previously built at SpaceX and Tesla and are applying a Starlink/Cybercab-style design-for-manufacturing playbook to homes: deploy the first units over the next three months, then scale toward production rates comparable to the largest car companies. The announcement is a strong pedigree and thesis signal; the stated deployment and production milestones are the execution proof to underwrite.

2. Emerging Teams

HomeTurf is an early mortgage-marketplace experiment with real, if still modest, supply-side traction. The founders say the product lets borrowers anonymously post a mortgage need while NMLS-verified lenders bid against one another, with contact information withheld until a winner is chosen. It is live in five states with about 30 lenders actively bidding; borrowers are free, lenders are temporarily free, and the planned monetization is a flat lender subscription rather than per-lead fees or paid placement. One co-founder spent years at Wells Fargo and another has long mortgage-startup experience.

The acquisition loop is currently paid ads, Facebook groups, word of mouth, the co-founder’s industry network, and cold calls; the founder describes uptake as early but gaining momentum. The main diligence risk is regulatory and operational: the team says state-by-state compliance requires legal work and creates substantial monetization hurdles.

Suhail’s unnamed inference startup is a watchlist signal, not yet an investable story. The public update is hiring model-optimization and inference-performance engineers, says compute is already locked down, and calls for an in-person San Francisco team. It gives no product or traction detail, but the infrastructure-first hiring and secured compute suggest an attempt to move quickly from zero to a production-backed company.

3. AI & Tech Breakthroughs

The model race is moving into execution economics. A Gemini 3.7 Flash launch post says the workhorse model targets coding and agentic workflows at half the introductory price of 3.6 Flash, with reported gains in software engineering from 37.0% to 65.3% and enterprise automation from 13.4% to 30.4% after moving from 3.5 to 3.7 in three months. OpenAI is previewing GPT-5.6 Sol’s Ultrafast mode at up to 14x the speed for select API customers, while Perplexity is moving Sonar into an Agent API with grounded search, multi-step research, code execution, built-in tools, and multi-model access; it reports more than doubling the best Sonar score on BrowseComp and WideSearch.

The technical implication is not simply “better models”: latency, token cost, tool orchestration, and access to multiple models are becoming the product surface. That favors infrastructure and applications that can route work intelligently, rather than products whose only differentiation is access to one model.

World-model evaluation is exposing where attractive metrics stop being useful. The open-source worldproof project compares predicted rollouts with ground truth and physical invariants rather than scoring task success. On real SO-101 robot-arm footage, a copy-last-frame baseline reached 0.983 SSIM and 53.9 dB PSNR over a six-step horizon, with flat error across the horizon—meaning the setup could not distinguish a good model from a do-nothing predictor. On DROID footage, the only discriminative region was roughly steps 8–24; both the short and long ends tied.

For robotics and physical-AI diligence, benchmark design is itself an investable layer: ask whether the horizon, frame rate, task speed, and metric can actually rank systems. The project is Apache-2.0 and runs its evaluation path on a laptop without a GPU, which lowers the cost of independently checking model claims.

AI-for-science may need records of decisions, not just records of results. Nathan Benaich argues that papers show the “happy path” while omitting failed experiments—the history models need to learn scientific taste. A useful record must distinguish an underpowered assay, degraded reagent, and wrong hypothesis; it must also address the counterfactual problem that only the chosen experiment produces an outcome. He proposes that funders reserve part of a grant for testing branches that would otherwise be discarded, so labs do not repeatedly rediscover the same dead ends.

4. Market Signals

Model access is becoming a switchable input rather than a durable relationship. Suhail says he stopped using Anthropic models and has no loyalty to whichever model he uses next; his criteria are “fast, cheap, reliable.” Martin Casado separately says AI code generation feels saturated and that his remaining work is architectural and semantic—scale, performance, and trade-offs. Paul Graham says startup interest in tuning open-weight models has returned after a period when it was considered a waste of time. These are directional operator and investor observations, not a market-wide measurement, but they favor companies that own distribution, data, evaluation, and cost control rather than thin model wrappers.

Enterprise AI GTM should follow buyer exposure and proof portability, not prestige by default. a16z’s Lighthouse/Landgrab framework puts regulated, high-consequence markets where proof travels in “lighthouse” territory: a few credible customers make the category safe to buy. Established-budget markets with recoverable mistakes and fragmented buyers are “landgrab” territory, where math and coverage matter more than a marquee logo. The current repost makes the operating implication explicit: founders can sell in Ohio, Chicago, or St. Louis instead of competing for the same San Francisco logos.

For diligence, ask which motion the company has earned and whether it can execute it. POCs should have fixed milestones and a defined conversion point; otherwise an enterprise “trial” can become an open-ended science project that consumes the founding team.

Dual-valuation rounds are becoming a cap-table diligence issue. Newcomer reports that Starcloud’s $170M Series A was split into a first Benchmark-led tranche priced at a $250M valuation and a second tranche that closed at more than four times that price. The article says roughly 25% of recent deals seen by MVP Ventures used dual valuations, a view echoed by six other early-stage investors who described the mechanism as moving from rare to pervasive. Investors should request tranche-by-tranche pricing and model how the structure affects employee reference prices and future dilution.

Background agents are becoming an operational product category. LangChain’s Harrison Chase says agents running in the background will move work beyond direct prompting, with cron schedules as an early implementation; a companion post says scheduled self-prompting agents can be built in a few lines. Sequoia’s current framing groups the system into an open agent layer, a compounding evaluation loop, and a governed runtime. The commercial opportunity is shifting toward state, scheduling, observability, and control—not just another chat surface.

5. Worth Your Time

  • Read AI needs science’s search history. It is the clearest current argument for collecting failed and unchosen scientific paths rather than training only on polished results.
AI’s New Battleground: Proof, Price, and the Runtime
TechCrunch
  • Manon Mehta, co-founder and managing partner of Unshackled Ventures, founded the firm 11 years ago with Nitin after an investment-banking career, a stint at a company acquired by Intel, and an attempt to start a company with an H1B co-founder . Unshackled built an R&D technology lab structure that can sponsor visas and deploy capital via payroll on a typical YC post-money SAFE, pre-assigning IP to the founder's Delaware C-corp with no taxable event; it required ~300 hours of legal work across immigration, labor, securities, tax, and employment law . The firm reports 300 immigration filings, 13 visa types, and 100% success across four administrations, including an O1 pathway for a DACA founder .

  • Fund history: 2015 debut fund of $4.5M against a $3.5M target, only $2M invested, 79 LPs; second fund of $20M with 99 LPs . Some first-backed founders now have companies backed by Blackstone and Fortress and going public .

  • Thesis: immigrant founders face four structural barriers – work authorization, no friends/family capital, shallow networks, and VC pattern recognition . Mehta argues pattern-matching creates a "familiarity tax" – higher valuations, larger checks, fewer bets, concentrated returns – while outlier wealth is created in structural gaps . Unshackled underwrites founders at day zero, before product or traction, pricing in unknown risk .

  • Founder evaluation: after every pitch the team scores IQ, AQ (adversity), EQ, and SQ (social), then compares scores to separate bias from signal; archetypes are "technical visionaries" (PhD/postdocs, higher IQ/AQ) and "systems disruptors" (regulatory/complex-systems operators, higher EQ/SQ) .

  • H1B policy: current administration's H1B modernization allows holders to self-petition and be their own employer – a reform Unshackled pushed in a white paper nine years ago .

  • Market view: Mehta expects the AI era to drive a return to building, manufacturing, and scaling in the US, and points overlooked founders to CDFIs, grants, and university/local competitions; he cites Eric Ries's "Incorruptible" framing of for-profit as maximizing human flourishing (Novo Nordisk, Costco) .

The High Risk, High Reward Model of Backing Immigrant Founders l Build Mode
TechCrunch

Unshackled Ventures' thesis and structure. Manon Mehta, co-founder and managing partner of Unshackled Ventures, on TechCrunch's Build Mode podcast, describes a fund that backs immigrant founders at day zero, before most investors will take the risk . Mehta — ex-investment banker, previously in marketing at an Intel-acquired company — started the firm ~11 years ago with co-founder Nathan after failing to raise for a co-founder on an H1B visa, hearing “don't the best immigrants find a way” . The four problems it solves for immigrant founders: legal work authorization, no friends-and-family money, shallow networks, and VC pattern recognition . Its legal structure is an R&D technology lab that can deploy capital via payroll to sponsor visas; investments are typical YC post-money SAFEs, with all IP pre-assigned to the founder's Delaware C corp so downstream VCs face no tax event or liability — a design requiring ~300 attorney hours across immigration, labor, securities, tax, and employment law .

Record and policy shift. Unshackled reports 300 immigration filings across 13 visa types, 100% success under four White House administrations, including a DACA founder moved to an O1 visa after a 4-month effort . Mehta credits the current administration's H1B reform — for the first time, H1B holders can self-petition and be their own employer — and argues immigrants (15.4% of the US population) can produce even more .

Fund history and proof points. Debut fund was $4.5M in 2015 (target $3.5M; only $2M invested while building the structure) with 79 LPs; the second fund was $20M with 99 LPs . A decade on, companies Unshackled backed first are now being backed by Blackstone and Fortress and going public “in the last quarter” .

Founder evaluation framework. Unshackled scores every pitch on IQ, AQ (adversity/resilience), EQ (self-awareness, growth mindset) and SQ (social quotient), proceeding only if the team wants to learn more; it underwrites founders before product or customers and prices in the unknown rather than debating market size . It invests in two archetypes: “technical visionaries” (PhD/postdoc-heavy, higher IQ/AQ) and “systems disruptors” (drawn to regulatory and complex systems, higher EQ/SQ) .

Market critique and AI outlook. Mehta argues VC pattern matching is a “familiarity tax” yielding higher valuations, fewer bets, and a concentrated system, while wealth is created by teams bet against — so structural gaps are where outliers play and deserve non-concessionary dollars . He calls much of the industry “pretty lazy” and says the job is to maximize the value of taking risk, not minimize it . He predicts the AI age pushes society back toward creation and manufacturing over the next 5-7 years, a tailwind for builders and local communities .

The VC Backing the American Dream with Manan Mehta, Unshackled Ventures
No Priors: AI, Machine Learning, Tech, & Startups

Founder path: Erik Allebest and a friend bought the Chess.com domain at a bankruptcy auction in 2005 for $56,000 while he was at Stanford GSB; nearly every VC told him it was uninvestable, so he bootstrapped with a CTO co-founder, launched in 2007, began charging for membership within ~18 months, and quickly became profitable.

Scale: Chess.com now has ~10M DAU, 40–50M MAU, 250M+ registered members, ~$200M+ expected revenue this year, and a 650-person fully remote team.

Growth durability: Before Covid and The Queen's Gambit, Chess.com had ~1M DAU; the 2020 Covid/Queen's Gambit/news-cycle wave and a 2023 wave (short-form content, Mittens bot, cheating scandal) each reset the baseline higher, and by 2024 Allebest concluded growth was compounding rather than a fad.

Capital story: After years without VC, General Atlantic became Chess.com's first PE investor in ~2020/21 via a secondary purchase ; CVC later bought into another secondary transaction (GA reinvesting), with no capital going into the company — Allebest cited CVC's deep experience in gaming, sports and community businesses as a fit.

Competitive moat: Chess is not patentable and there are tens of thousands of chess apps; Allebest credits Chess.com's market leadership to relentless UX, community and content, while open-source Lichess coexists as the 'Linux of chess.'

Product expansion: Chess.com launched Gambit (gambit.com) poker, applying the 'chess playbook' — a skill rating ('poker for ratings') rather than chip counts, bots or money — with plans to roll out more classic games; he expects players to care about poker rating as much as money.

AI/tech: Chess.com uses AI across support automation, data/analytics (LLM-based), an internal 'GNS' auth/knowledge layer, martech and agentic development to shorten idea-to-ship time; it is building an AI chess coach and personalized stat comparison, and runs ML/statistical anti-cheating models trained on how humans vs computers play.

Thesis signal: Humans still value human competition even after AI became superhuman at chess; neural-net engines made the game more exciting, and AI coaching/puzzles now help humans enjoy and learn it. Allebest is optimistic AI will accelerate human learning but flags the need for guardrails and worries about economic concentration.

How Chess.com Became the World’s Biggest Chess Community with CEO Erik Allebest
Garry Tan
  • Brian Lovin, who helped design Notion Agent at Notion, reviewed Grok Bot: the app shifts from thread-centric to agent-centric interaction, where users manage named/personified specialist agents that coordinate ('hire a team of specialists'). Lovin sees the model as fine but notes you can't queue messages, so multitasking inside a long-running agent task is awkward, and each agent's single long-running thread forces users to duplicate agents to keep context bubbles around sensitive data; he's unsure whether the market wants a Jarvis-style omni-agent or multiple coordinating agents — possibly both for different audiences .
  • Grok Bot hides thinking steps/tool calls, unlike almost every other AI agent: Notion Agent customers demanded those steps back as a loading indicator and to spot the model going off track, though Lovin suspects a year of model progress makes hiding the noise acceptable now (he guesses agents' mid-task progress updates run on Grok 4.6); the app also has no model picker, with opaque routing that requires a 'model take the wheel' mindset that feels uncomfortable short-term .
  • Multi-agent coordination is immature — his agents go in loops talking to each other before stopping — yet delegating agents to work simultaneously across the cloud computer, his local machine, and Cursor cloud coding sessions 'works really well', and consolidating skills/MCPs/connectors behind a single 'Plugins' concept matches frontier consensus; agents also adapted on request to a Notion-database memory layer, creating ~20-30 pages .
  • Market caution: this agent-tool shape is becoming common across many companies and open source, so it's unclear how users will justify one $200/mo subscription vs another — near-term differentiation may come down to vibes/tribal affiliation/token subsidization until a capability-level differentiator appears — and Lovin doubts most people have enough to automate in daily life, asking for retention graphs .
  • Garry Tan replied publicly: 'I agree, you need topics per-bot please!'
First impressions of Grok Bot: - With Grok Bot you don't juggle threads, you juggle named/personified agents. This seems like a \~fine me… I agree, you need topics per-bot please! [https://x.com/brian_lovin/status/2087917528001794393](https://x.com/brian_lovin/status/20879175…
20VC with Harry Stebbings
  • Canva, still private, cut its 2026 growth outlook from ~30% to ~20% (a one-third slowdown) with ~$3B GAAP revenue last year; CEO Melanie Perkins blamed the cost of subsidizing users with frontier-model AI, and Canva is now building its own in-house image model . Panelists debate whether this marks the start of structural decline ("30 on the way to 20 on the way to 10") as ChatGPT/Claude absorb prosumer creative work; agents no longer suggest tools like Canva/Notion . Figma also dropped ~20% and warned agentic usage will impair gross margins . One panelist pegs Canva's value at ~$12B vs its ~$50B 2021 valuation, with existential-risk perception the key variable .

  • AI talent competition has created a three-tier comp structure: even early-stage companies now need "God tier" packages (seven figures cash, 10x equity) for a handful of superstars; the hiring bar is now "did you get an offer at Anthropic or OpenAI?" . Google's exodus — Jeff Dean leaving after 27 years, Demis Hassabis stepping back — coincided with a ~$200B stock drop, underscoring the value of top AI researchers . Venture appetite for moonshot AI-for-science startups is at an all-time high .

  • Regulatory and infrastructure risk: Rep. Ro Khanna plans a "data center bill of rights" letting local communities reject AI data centers; NIMBYism is spreading, though panelists expect state/county competition to resolve it and call power availability the real bottleneck . Elon Musk's Terrafab ($16.8B first installment, Intel in the consortium) and Intel's first equity raise since 1979 signal AI-capex-driven vertical integration pressure .

  • Anthropic is reportedly close to an IPO (within 60 days); the panel argues it should go public now while ahead of ChatGPT. OpenAI completed a $7B secondary; a $1M Anthropic stake from 2020 is now worth $51M, distorting startup hiring expectations across the ecosystem . DeepSeek is raising $8B at a reported $74B valuation, and ByteDance banned distillation of US models — signs of intensifying AI competition .

  • Founder-led thesis: Jason argues any portfolio company not run by a founder is "going to be a zero" in this age; Revolut's leaked compensation ratchet (up to ~40% at $500B) and dilution concerns illustrate the premium on founder control. Pitchbook data cited shows returns are being compressed by unprecedented dilution and high entry prices, with seed investors now modeling ~4x effective entry after 75% dilution .

  • Enterprise SaaS stress: Atlassian's beat (biggest jump since 2015) was real, but the panel flags monetization of free Loom seats as a sign of stress; HubSpot ($10B) faces low-end AI-native competitors growing at unprecedented speed, threatening SMB software incumbents .

Canva Slahes Growth | Talent Exodus at Google | Revolut's $50B CEO Package | Musk's $55B Terrafab
  • a16z podcast with Joe Schmidt and Andy (former Samsara/Meraki sales leader) introduces a Lighthouse vs Land Grab framework for enterprise AI startups: Lighthouse suits regulated, high-exposure markets with few logos where proof must travel; Land Grab suits established-budget markets where startups show ROI math .
  • Portfolio/example companies: Harvey (AI for law) as lighthouse, winning top law firms and making the category safe to buy ; Stu (accounts receivable automation, founded by Tark and Ben) as land grab, selling mid-market with a math-based pitch ; Pylon (a16z-backed AI customer support) as land grab starting at modest ACVs ; Decagon (customer support) excelling at evangelizing benchmarks ; Further AI (insurance, Joe on board) as lighthouse via governance-first approach ; Applied Intuition (portfolio) as persistent lighthouse in autonomous software .
  • Thesis: AI is creating a new "big software" moment where enterprises reimagine core systems (CRM, HR) with agents rather than do skeuomorphic replacements, favoring platform sales over PLG wedges .
  • Practical caveats: AI POCs risk becoming endless "science projects"; define fixed end dates and success criteria . Founders should spend ~1% time on strategy and 99% talking to customers . Early-stage companies should hire sales ops early and aim for 100% quota attainment .
The GTM Advice Behind Billion-Dollar AI Companies
Harry Stebbings

From Harry Stebbings' podcast conversation with @jasonlk (Jason Lemkin) and @rodriscoll :

  • 2021-era valuations are resetting hard: Airtable, valued at $11B, sold for $2B; Canva, valued at $44B, is slashing growth targets and being cannibalized by OpenAI . As software growth slows (e.g., 30% to 20%) and AI inference/serving costs rise, private-market valuations are resetting toward lower public-market revenue multiples .
  • A* talent is unaffordable for most startups against OpenAI/Anthropic compensation packages — the AI talent exodus from Google is so severe that even Demis Hassabis and Jeff Dean can't hold people . Lemkin recommends a 'god-tier' comp layer for the one-to-five superstars (vs 'regular human beings' and 'the AI guys'), with outsized equity and cash — at $200M revenue, four god-tier employees won't break the model .
  • Dilution is the worst it has ever been: record SBC and inflated entry prices are compressing venture returns, and 'superhero pay packets' (Nik at Revolut, Elon Musk) are normalizing . Investors should model up to 75% dilution — effectively 4x their initial entry price — for realistic outcomes .
  • Founder presence is existential in the AI era: any investment where the founder leaves should be written to zero — non-founder executives lack the agility, vision, and obsession for radical pivots, turning late-stage portfolio stars into 'functional zeros' .
  • Hesitancy on AI in 2023 will be paid for in 2026-27 via competitive displacement and revenue erosion ; only undeniable top-line growth can defend software multiples against AI-displacement narratives ; meaningful liquidity arrives only at the end of the journey, after platform shifts are navigated .
There are three ugly truths right now that we are not admitting. 1. The valuations we have from 2021 are representative of true value tod… "We're now seeing three bands of compensation: regular human beings, the AI guys, and the one to five superstars. We have to provide them…
Sequoia Capital

Harrison Chase, co-founder/CEO of LangChain, first noticed in 2022 (GPT-3 era) as an early thinker on building harnesses to turn LLMs into collaborators/agents; LangChain frames an agent as three parts — harness orchestrating model and context — and recommends owning all three, including model switching to avoid lock-in .

All agents share a core loop of an LLM calling tools; harnesses customize this with middleware/hooks, and for out-of-distribution tasks with bespoke cognitive architectures; LangChain recommends starting with off-the-shelf harnesses (Claude Code, Codex, DeepAgents) and customizing as the task moves out-of-distribution .

DeepAgents is LangChain's model-agnostic general-purpose harness with file systems, skills, and subagents; its "model profiles" swap per-model implementations for in-distribution sub-tasks like file editing .

Harbor, from the makers of Terminal Bench 2, is becoming the industry-standard open-source eval runner: tasks define a Dockerfile-based environment, verifier tests (code/unit tests/LLM or agent as judge), and prompt; benchmarks such as Frontier Bench compare harnesses, models, and reasoning efforts .

LangSmith is LangChain's evals/observability platform; LangSmith Engine (launched in past few months) is a background agent that curates traces, creates issue boards, and proposes fixes to prompts, context, or harness code; LangChain dogfoods it and built an "issue bench" to benchmark it .

Benchmarks drive LangChain's roadmap: running Codex/Claude Code/DeepAgents on their engine benchmark showed Codex winning via aggressive script-writing, prompting a "codexification" sprint on Engine .

LangChain attributes most agent failures to poor context rather than model quality, making observability into the context window critical; they automate feedback with cheap fine-tuned SLMs as online judges, and with Harvey (legal AI) significantly cut LLM-as-judge cost .

When to Build Your Own Agent Harness | Harrison Chase, LangChain
a16z

a16z's AI sales guidance contrasts two go-to-market playbooks for enterprise AI: "lighthouse" (win marquee logos so proof travels) vs "landgrab" (win budgets that already exist), citing Harvey and Applied Intuition as lighthouse examples and Stuut, Decagon, and Pylon as landgrab examples, with a section on why most companies need both playbooks . The discussion draws on a16z GP Andy McCall's operator record — building Meraki's sales org and, as CRO, taking Samsara from single-digit millions to >$1B ARR — with chapters on Samsara's ELD landgrab and Meraki's free-access-point playbook . McCall's tactical advice: spend 1% of time on sales strategy and 99% on execution, and hire sales operations earlier than most founders do — "It can be literally one person" handling territory alignment and commissions, since unresolved sales-org mechanics "become speed bumps otherwise" .

Lighthouse or Landgrab: Choosing Your AI Sales Strategy In this conversation, a16z's Joe Schmidt, Andy McCall, and Elena Burger sit down … a16z GP Andy McCall on advice for starting a sales career, and the hire founders wait too long to make: "I made this mistake early in my …
a16z
  • a16z's Joe Schmidt published "Lighthouse or Landgrab?", arguing AI founders sell either proof or math via two enterprise GTM playbooks: Lighthouse (category creation; win marquee logos whose adoption proves the category) or Landgrab (buyer already knows the problem; win on math, speed, and volume). The deciding map has two questions — how exposed is the buyer who signs (career risk from a mistake) and does social proof travel (do buyers watch each other): high exposure + portable proof = lighthouse; recoverable mistakes + fragmented buyers = landgrab .
  • Lighthouse cases: Harvey (AI for legal) — Allen & Overy signed late 2022 and Paul Weiss early 2023, unlocking the market; Harvey now has hundreds of millions in ARR and an $11B valuation. Hebbia (AI intelligence for financial services) — won the world's largest PE firms, hedge funds, and consultancies first, then expanded to 40%+ of the largest asset managers by AUM, including KKR and BlackRock .
  • Landgrab cases: Stuut (accounts-receivable automation: collections, payments, cash application, dispute resolution) raised a $29.5M Series A led by Andreessen Horowitz; it went wide early in the lower middle market (manufacturers, distributors, logistics across Michigan, Ohio, Texas) instead of chasing Fortune 100 logos, deploying in under a week vs 6–18 months for traditional rollouts; customers reported 40% more cash flow, 70% fewer manual tasks, and 37% shorter collection periods . Decagon (AI customer support) — founders ran roughly 100 customer conversations in a month before building, then scaled 0 to 8-figure ARR in 18 months, signed 100+ new enterprise customers in 2025, and tripled valuation to $4.5B in under six months .
  • Competitive dynamic: in landgrab markets incumbents keep adding AI each quarter (e.g., Zendesk's copilot vs Decagon replacing the support function), so speed is the whole game — per a16z's Alex Rampell, "you need to get distribution before the incumbent gets innovation" .
  • Sequencing: the best companies deliberately move lighthouse-to-landgrab — Affirm broke through with Casper, then every mattress company, then adjacent big-ticket categories like exercise equipment; the signal that the transition is earned is buyers approaching with allocated budgets and asking for a demo rather than asking who went first .
  • Trap flags for diligence: lighthouse traps include becoming hostage to a marquee logo, prestige without payback (non-recurring, low-ACV, non-replicable), pilot purgatory (six-month POCs that never convert), and building a product for one lighthouse customer; landgrab traps include "dying of indigestion" (poisonous low-ACV, hard-to-onboard customers without qualification discipline), grabbing land you can't hold (scaling coverage before product readiness creates churn), and mistaking canvassing a metro for reaching the real 50,000-company market .
  • In the accompanying post, Schmidt tells founders they don't all need the SF/NYC marquee-logo sales motion: "Or go out and sell in Ohio, go out and sell in Chicago, go out and sell in St. Louis. Find people who need your solution" .
Lighthouse or Landgrab? Joe Schmidt says not every enterprise AI company needs a billboard in San Francisco: "You just realize that you have the same two competi…
Sam Altman

OpenAI is previewing Ultrafast mode for GPT-5.6 Sol, offering up to 14x speed, launching first in the OpenAI API to a select group of customers with expanded access as capacity grows . OpenAI CEO Sam Altman amplified the announcement on his X account .

Previewing Ultrafast mode: GPT-5.6 Sol at up to 14x the speed. Launching first in the OpenAI API to a select group of customers with expa… /ultrafast [https://x.com/OpenAI/status/2087947721936359705](https://x.com/OpenAI/status/2087947721936359705)
a16z
  • Vals announced a $40M Series A at a $400M valuation, led by a16z, with existing investors 8vc, Pear VC, Bloomberg Beta and new investors HRT Ventures, Next Ladder .
  • Vals Smith is now GA — custom coding benchmarks from any GitHub repo with 120 free credits; also shipped Vals Index 2.0 with wider economy coverage .
  • Launched risk benchmarks: RSI Index with CoreWeave, ReverseEngBench (cyber) with Columbia, Tufts, UC Berkeley, UCLA; initial mental-health work, with environmental impact, military, biosecurity to come .
  • Traction: revenue up 8x vs all of 2025; customers doubled and team tripled in 6 months; results cited in model cards from OpenAI, Anthropic, Google, Meta, xAI .
  • a16z backs Vals as the independent evaluation layer for the AI economy, grading models on real workflows via domain-expert-built benchmarks and automated grading; founders Rayan Krishnan and Langston Nashold .
Today we're announcing our $40M Series A at a $400M valuation, led by [@a16z](https://x.com/a16z) , with participation from existing inve… We're thrilled to invest in Vals. A frontier model can look brilliant on a leaderboard and still struggle with the messy work that actual…
Paul Graham
  • Paul Graham observes that tuning open-weight models is back among startups, after a year ago the consensus was that it was a waste because model companies' next versions would surpass them soon .
  • He expects more back-and-forth on this, and if AI tops out, open weights win — leading either to a thousand model companies with proprietary training data as moats, or four after a new JP Morgan combines them .
A year ago lots of startups were tuning open weight models. Then the consensus became that this was a waste of time, because the model co… I wouldn't be surprised if there are several more back and forths. I have no idea who ends on top. If AI tops out somehow, open weights w…
Garry Tan

Heart Aerospace flew the world’s largest electric aircraft on August 12, per @gustaf . Garry Tan points to this as proof that "YC is the YC for hard tech," signaling YC’s push into hard-tech startups like electric aviation .

On August 12, [@heartaerospace](https://x.com/heartaerospace) flew the world’s largest electric aircraft. [https://www.youtube.com/watch?… YC is the YC for hard tech [https://x.com/gustaf/status/2087927880017965357](https://x.com/gustaf/status/2087927880017965357)
Paul Graham

Paul Graham observes that making an AI model faster is a strategically smart move when customers pay per token, since speed directly reduces customer cost — a relevant lens for evaluating AI startup differentiation and pricing economics .

Making your model faster is a very smart move when your customers are paying by the token.
a16z
  • a16z's Joe Schmidt, Andy McCall, and Elena Burger discuss two enterprise AI sales playbooks: lighthouse (win marquee logos so proof travels) vs landgrab (win budgets that already exist), naming Harvey and Applied Intuition as lighthouse examples and Stuut, Decagon, and Pylon as landgrab examples .
  • Andy McCall's operator track record: built Meraki's sales org and, as CRO, took Samsara from single-digit millions to >$1B ARR .
  • Takeaways for founders: spend 1% of time on sales strategy and 99% on execution, hire the sales org based on your strategy, and most companies ultimately need both playbooks .
Lighthouse or Landgrab: Choosing Your AI Sales Strategy In this conversation, a16z's Joe Schmidt, Andy McCall, and Elena Burger sit down …
martin_casado

@martin_casado says he feels AI code generation is saturated: as a VC dev, nearly all his time is now spent on architectural/semantic issues (scale, performance trade-offs), and he hasn't seen meaningful improvement since Opus 4.5 .

Maybe it's because I'm a VC dev. But I feel AI code generation is saturated. Nearly all the time I spend now are architectural or dealing…
Aravind Srinivas

SpaceXAI's Grok 4.6 is now available in Perplexity and Perplexity Computer . Perplexity CEO Aravind Srinivas says it was benchmarked as an orchestrator on Perplexity's Wide-And-Deep-Research benchmark using the Perplexity Computer harness, sits on the Pareto frontier of performance vs cost, and is available to all Pro and Max users on Perplexity . On the WANDR benchmark, Grok 4.6 matches Fable 5 results at over 60% lower cost .

Grok 4.6 is now available in Perplexity and Perplexity Computer. On WANDR, it sits on the Pareto frontier of performance and efficiency, … Congrats to [@SpaceXAI](https://x.com/SpaceXAI) on one more amazing model: Grok 4.6. We benchmarked it as an orchestrator on our Wide-And…
David Ulevitch 🇺🇸

Uplift launched, backed by a $7M seed led by a16z . The company builds homes the way cars are built, on a production line . Co-founders (cnitschelm and trevoroleary) previously built at SpaceX and Tesla, applying the Starlink/Cybercab playbook—design for manufacturing, ship volume, drive cost down per version—to housing . They plan to deploy first homes within 3 months and scale toward production rates rivaling the largest car companies, with volume to house thousands then millions per year, aiming to 'move rent' . They are hiring engineers and taking orders . a16z GP David Ulevitch (per account description) endorsed: 'a very special team doing the work that matters most to the most number of people' .

We are launching Uplift, backed by a $7M seed led by [@a16z](https://x.com/a16z). We build homes the way cars are built: on a production … This is a very special team doing the work that matters most to the most number of people. 👇 [https://x.com/cnitschelm/status/20879019918…
Y Combinator

Heart Aerospace, the Y Combinator W19 startup co-founded by Anders Forslund, completed the first flight of its X1 aircraft — the largest battery-electric aircraft ever flown, with a 106-foot wingspan and a takeoff weight over 25,000 pounds . The 27-minute flight ran on roughly $5 of electricity; the company's next milestone is the ES-30, a 30-seat hybrid-electric airliner targeting service entry in 2031 .

Congrats to [@AndersForslund1](https://x.com/AndersForslund1) and [@heartaerospace](https://x.com/heartaerospace) (W19) on the first flig…