ZeroNoise Logo zeronoise
Post
AI’s Verification Layer Is Becoming the Investable Bottleneck
1 day ago
6 min read
2300 docs
Mundo’s multimodal-data financing, autonomous sales traction, and new research on synthetic environments and hardware-aware agents point to the same shift: evaluation, provenance, and control are becoming as important as model capability.

1. Funding & Deals

Mundo AI raised a $20M Series A led by GreatPoint Ventures, with participation from Y Combinator, Next Frontier Ventures, and E12 Ventures; a previously unannounced $4M seed brings total capital to $24M.

Mundo is positioning the company as a data-and-evaluation layer for “perceptual intelligence”: AI that understands speech, video, gestures, environments, and human interaction. It says its datasets and evaluations are already used by leading AI labs for speech-to-speech interaction, fine-grained video understanding, and emerging modalities, and argues that progress depends on a continuous feedback loop between better data and better evaluations. The investment question is therefore less “which model wins?” than whether Mundo can make multimodal data and evaluation demand repeatable and defensible.

2. Emerging Teams

OSOA is an early paid test of an autonomous sales closer. The product finds and verifies prospects, writes outreach in the owner’s voice from the owner’s accounts, negotiates within preset limits, and closes without human involvement after setup; the founder says four businesses paid in the first week at $59–$179 per month, from 20 founding spots. The failure report is more informative than the demo: the system needed hard guardrails to avoid closing unwanted deals, while prospect verification produced false positives and nearly sent outreach to dead leads before a rebuild. The founder is raising a $500K SAFE, and the product began as an internal sales system for the founder’s first company. For diligence, policy enforcement and prospect verification matter more than message quality.

MarketOwl is a distribution signal, but not yet product-market-fit evidence. A first-time European founder who has been building solo since 2023 says an earlier version generated five to six demo calls a week but retained high churn; after shifting to Reddit and iterating toward an AI marketer, the founder reports testing with 12 companies at an average 15% positive-reply rate and getting zero-follower Threads accounts to posts with more than 1,000 comments. The proposed moat is a playbook layer combining social posts, DMs, ads, and SEO, built from three years of campaign experience. That makes durable customer impact—and avoiding channel or account bans—the key validation test, not raw automation volume.

A newly out-of-stealth earthquake-prediction startup is a useful diligence caution. Its team combines a serial technology founder who runs an SMB-focused IT company with a doctor of seismology; the founders say their methodology, developed using Greek seismological data, predicts California earthquakes up to 48 hours ahead. They later define the target as within 125 km, one magnitude, and 24 hours, reporting 80% accuracy and 83% recall. The claim is not yet underwriting-grade: a commenter who said they read the paper characterized it as off-the-shelf ML trained on historical Greek data, while another questioned the accuracy definition and requested false-positive and false-negative rates.

3. AI & Tech Breakthroughs

SPADE turns synthetic-environment design into a learnable post-training loop. Its Environment Designer writes executable, long-horizon training environments while a Reasoning Agent solves them; on Qwen3-30B, the reported game-suite average was 58.3, or 8.1 points above base and 5.3 above the strongest fixed-environment baseline, while tool-use environment design improved every tested backbone. Code and checkpoints are available. The investable implication is a potential reduction in the cost of creating diverse post-training tasks, although the framework’s gains remain bounded by the capabilities of the model generating the environments.

Hawkeye packages hardware-specific kernel optimization into a test-driven coding-agent workflow. Researchers from Harvard, Stanford, Together AI, and Caltech describe an open-source framework whose unit tests pair a human-authored solution kernel with profiling metrics and a usage guide. They report matching or exceeding torch.compile on tested workloads across NVIDIA Ampere, Hopper, Blackwell, and AMD MI350, and an 18.9× geometric-mean speedup over expert-authored Triton kernels for emerging attention variants. If reproducible, this moves scarce accelerator expertise into reusable agent infrastructure rather than leaving every new chip or kernel pattern to specialist manual optimization.

Protege makes healthcare evaluation an outcome-measurement problem, not simply a model-quality problem. The article argues that medicine lacks a shared absolute ground truth, that clinical datasets are guarded, and that benchmarks often measure test design or physician preference rather than clinical merit; objectively evaluating a decision requires following the patient forward over time. The context gap is concrete: the median patient record contains about 8,500 tokens, while five of six public healthcare-AI benchmarks provide less than median context and several provide fewer than 200 tokens per case. Protege says it uses real-world partner data to build tests tied to health gains for clinicians and patients.

The broader calibration remains uneven. Import AI’s summary of a METR analysis reports major acceleration in vulnerability discovery, minor and hard-to-measure acceleration in mathematics, and no measurable acceleration across seven AI algorithmic-progress areas. That makes domain-specific evaluation a prerequisite for distinguishing a real capability phase change from a compelling demo.

4. Market Signals

Agent demand is rising while open-weight usage and model costs are repricing the stack. Exponential View reports that open-weight tokens’ share doubled in the last year and is approaching a 1:1 ratio with closed-weight tokens, even as the number of closed-weight tokens grew sevenfold; it also says agents now use 14× their February token volume while human token usage grew 2.8×. A separate essay argues that AI-token costs could fall by more than 1,000× in under a decade, with cost declining 2–5× per year as utility improves, shifting value toward applications that become viable at lower prices.

The infrastructure corollary is a power and supply-chain contest. The essay characterizes Nvidia’s CUDA and accumulated tooling as owned, packaging capacity as rented, and power generation as absent; it says Nvidia’s four largest customers have contracted roughly 9.8 GW of nuclear power while also shipping their own silicon.

Systems of record are being pushed toward agent-native interfaces, with evaluation emerging as the routing moat. Jerry Liu says agents should use existing software such as Slack rather than rebuild it or rely on the vendor’s agent, because software and systems of record need to become agent-native. Garry Tan’s current formulation preserves deterministic APIs, ACLs, SQL, and data structures but says vendors must add the AI harness and full customer solution or risk being subsumed. In the adjacent model-routing layer, Brendan Foody argues that 80% of building a good router is building a good evaluation, with routing logic the easy part.

Launch attention remains weak evidence of software durability. An audit of 2,291 Product Hunt launches reports that 24.4% were hard-dead and 28.9% no longer existed as independent products; AI products died at essentially the same rate as non-AI products, 24.5% versus 24.4%. Most deaths occurred in the first year or two, and the author warns that website availability is only a proxy and may understate the true failure rate.

5. Worth Your Time

  • Read — Import AI 470. The most useful single synthesis in this period: uneven real-world acceleration alongside SPADE’s synthetic environments and Hawkeye’s hardware-aware kernel agents.

  • Read — The Oracle Problem. A strong case for longitudinal, outcome-based health-AI evaluation and for testing against full patient-record context rather than compressed vignettes.

  • Inspect — AQuA’s model-development loop. Its bounded configuration diffs and sealed evaluator are a useful pattern for making agent-led model search auditable, with explicit caveats around unequal compute, undisclosed features, and label construction.

  • Read — AI fact-checker citation audit. The author found 12 dead or nonexistent URLs among 215 citations and traces the problem to letting the prose model author provenance; the proposed remedy is retrieval-owned citations plus mechanical URL and text-support checks.

AI’s Verification Layer Is Becoming the Investable Bottleneck
Research extraction

Direct answer: Mundo’s $20 million Series A was led by GreatPoint Ventures, with participation from Y Combinator, Next Frontier Ventures, and E12 Ventures; including a previously unannounced $4 million seed round, the company says it has raised $24 million.

  • Founding team/background: The supplied announcement does not name Mundo’s founders or provide their backgrounds, so this part of the objective remains unresolved in the source bundle.
  • Product thesis: Mundo defines “perceptual intelligence” as AI’s ability to understand rich real-world sensory experiences—including speech, video, gestures, environments, and human interaction—and use that understanding to interact naturally.
  • Specific product wedge: The company says its datasets and evaluations are used by leading AI labs to build multimodal systems, including natural speech-to-speech interaction, fine-grained video understanding, and emerging modalities without established learning methods.
  • Core operating thesis: Mundo argues that progress will require a continuous feedback loop between better evaluations and better data, with datasets and benchmarks evolving alongside models rather than relying on better models alone.
  • Problem targeted: It characterizes current AI as capable at reasoning over structured information but weaker on unstructured real-world signals such as overlapping speech, background noise, incomplete camera views, interruptions, hesitation, implicit communication, and meaning absent from transcripts.
Mundo raises $24M to build the data layer for perceptual intelligence | Mundo AI
a16z
  • Protege’s thesis: Protege CEO Bobby Samuels is pitching the “Oracle Problem” as a central healthcare-AI bottleneck. Clinical decisions lack a shared, absolute ground truth, clinical datasets are closely guarded, and benchmark design can become a marketing artifact; determining whether a model is truly right requires following patients forward over time.
  • Technical wedge: Protege says it sampled millions of records from nearly a trillion tokens of EMR text across billions of notes through its data-partner network. Its goal is to build real-world tests that measure whether AI produces health gains for clinicians and patients, rather than only higher scores on abstract cases. The evaluation gap is substantial: the median patient record is about 8,500 tokens (mean about 39,000), while five of six public healthcare-AI benchmarks provide less than median context and several provide fewer than 200 tokens per case.
  • Market signal: The article cites OpenAI reporting more than 300 million weekly users asking ChatGPT health-related questions, while Protege’s data shows patient references to AI rising sharply after ChatGPT’s launch and AI soon being used to write nearly one-third of SOAP notes. This creates a growing need for independent, outcome-oriented verification infrastructure as healthcare AI moves into live care settings.
AI use by patients and clinicians inflected shortly after the launch of ChatGPT Protege CEO Bobby Samuels the Oracle Problem in health AI… The Oracle Problem: an invisible bottleneck to AI and medicine
Paul Graham

The post links to an essay titled “How Universities Should Prepare Founders,” addressing university support for startup-founder formation.

How Universities Should Prepare Founders: [https://paulgraham.com/prepare.html](https://paulgraham.com/prepare.html)
Garry Tan

Garry Tan is bullish on data-center expansion, stating that “Datacenters create jobs and prosperity” and linking to a Garry’s List post whose URL frames data-center opposition as “killing $1 trillion in AI infrastructure.” A quoted @BayAreaNewLibs post argues that a national data-center moratorium would “kneecap” the industry and says data centers have helped put San Francisco on the recovery track, signaling policy risk for AI-infrastructure investment.

Datacenters create jobs and prosperity, actually [https://garryslist.org/posts/data-center-nimbys-are-killing-1-trillion-in-ai-infrastruc… Supporting a national data center moratorium, kneecapping the industry that's single-handedly put SF on the recovery track, seems like a …
Garry Tan
  • OpenAI is signaling a platform-first strategy: Sam Altman said the company should be “more of a platform company than a product company,” merging ChatGPT and Codex into a unified interface for personal or company-wide AI, with an API for developers to build on top.
  • The platform is intended to span the cost-performance curve: OpenAI aims to offer high-end AI for scientific discovery alongside inexpensive models for high-volume work, while avoiding direct competition in every product category and targeting adoption by 100 million new businesses and 8 billion people. Garry Tan characterizes this positioning as OpenAI becoming “the much more open lab.”
Sam Altman: "I think we should be more of a platform company than a product company.” “We just merged ChatGPT and Codex together. So we u… Current meta: OpenAI is actually the much more open lab [https://x.com/dnapway/status/2091588622046871786](https://x.com/dnapway/status/2…
Garry Tan
  • AI-native systems of record: Systems of record may need to become AI harnesses or risk replacement by agents. The underlying deterministic APIs, ACLs, SQL, and data structures are expected to remain, but software vendors will need to build the AI harness and full customer solution or be subsumed by it.
Prediction: systems of record will need to become AI harnesses or face replacement by agents There will still be API’s and acl’s and sql and underlying data structures that are deterministic It’s just that software companies have …
Garry Tan

Garry Tan endorsed Conductor Cloud, saying it made him “so much more productive” and eliminated his need to keep his MacBook Pro open.

Conductor Cloud has made me so much more productive and I don't have to keep my Macbook Pro cracked open anymore
a16z
  • Healthcare AI’s bottleneck is shifting from model capability to verification. Clinical decisions lack a shared, objective ground truth; proprietary datasets are difficult to inspect; and static benchmarks can measure physician preference, prompt design, or test construction rather than clinical merit. Reliable assessment requires tracking patient outcomes over time and measuring real-world health impact, not just benchmark scores.
  • Protege, led by CEO Bobby Samuels, is building evaluation infrastructure for health AI. The company says it uses real-world data from partner networks to create tests for clinicians and patients, drawing on a random sample of millions of records from nearly one trillion EMR tokens across billions of notes; it cautions that the resulting figures are descriptive rather than causal.
  • Healthcare AI adoption is creating immediate demand for trustworthy evaluation. The article cites OpenAI reporting more than 300 million weekly users asking ChatGPT health-related questions, while Protege’s data shows patient requests for clarification about model outputs rising to roughly 1,686 per million notes partway through 2026; it also projects that AI will soon write nearly one-third of SOAP notes.
The Oracle Problem: an invisible bottleneck to AI and medicine AI use by both patients and clinicians inflected shortly after the launch of ChatGPT Protege CEO Bobby Samuels on the Oracle Problem in h…
martin_casado

The post cautioned against drawing firm conclusions from FT market-share data for Fable, saying the cause was unclear and could relate to cost, speed, aggressive moves by OAI, ZDR, open source, negative press or sentiment, refusal to serve some domains, or biased data.

Lots of folks jumping to conclusions on the FT market share data for Fable. The reason is not at all obvious to me. Is it due to cost? Sp…
Y Combinator
  • Mundo raised a $20M Series A and is building datasets, evaluations, and research with frontier AI labs and companies to improve AI perception and understanding of the real world across audio, video, and emerging modalities.
Congrats to [@hellomundoai](https://x.com/hellomundoai) on their $20M Series A! Mundo partners with frontier AI labs and companies to bui…
Y Combinator

Y Combinator says more than 20 of its startups have recently published AI research at top conferences including NeurIPS, ICLR, and ICML, signaling that major labs do not have a monopoly on frontier AI research.

Big labs don’t have a monopoly on AI research. More than 20 YC startups have published research recently at top conferences like NeurIPS,…
a16z
  • Ben Horowitz said he rejected the Databricks founders’ $200,000 request and wrote a $10 million check instead. The founding team comprised six PhD students and professor Ion Stoica, and had built open-source Spark against the already well-funded Hadoop ecosystem.
  • His investment thesis was that academic spinouts can underestimate their opportunity: technically strong teams facing a time-sensitive competitive window should commit to building a full-scale company rather than optimize for a modest outcome. The comments were made on Lenny’s Podcast in 2025.
.@bhorowitz on why he refused the Databricks founders' $200K ask and wrote them a $10M check instead: "So there were six PhD students, an…
martin_casado

Martin Casado says the apparent consensus is that Fable’s position in FT market-share data may be explained by ZDR, consistent with his own experience; he cautions that this remains unconfirmed, but argues that if true it would be a strong signal to AI labs about the impact of data-retention policies.

Consensus on this seems to be ZDR which tracks to my own experience as well. If that really turns out to be the case, it's an incredibly …
a16z
  • Sam Altman argues that high-risk bets are worthwhile when success would create exceptional value, citing the 2015 pursuit of AGI despite widespread skepticism that it was possible.
  • His selection thesis favors researchers and founders with non-consensus views, fresh approaches, high energy, and unconventional thinking over incremental variations; strong conviction can be valuable even when the bet may be wrong.
.@sama on how to pick high-risk bets, from AGI in 2015 to the researchers he hires today: "When we started, people thought it was totally…
@jason
  • Jason floated the possibility that Hugging Face is for sale and suggested AWS or SpaceXAI as potential buyers; this is unverified acquisition speculation, not a confirmed transaction.
  • He characterized Nvidia as making a $6B bet on open source and pursuing a “full stack” strategy.
we are live talking about the insane world robot games... it's the end of days folks! Also, Hugging Face is for Sale?!?!?!?!?! AWS or Spa…
@jason

Jason frames AI adoption as a two-stage workforce shift: companies may first eliminate unproductive roles and dismiss employees who refuse to use AI, while AI-enabled employees could uncover a dozen new business opportunities that require hiring ten more people who know how to use AI. He advises founders to fire non-users and hire or promote AI adopters.

Stop one Implementing AI: you eliminate positions that are unproductive and fire the ten people who refuse to use AI Step two implementin…
andrew chen

Andrew Chen describes a desired agentic-coding workflow in which eight parallel coding agents operate on the same repository and automatically handle overlapping work, merging, testing, Git operations, and conflicts, leaving the human to adjudicate only substantive issues. He specifically points to worktrees as the mechanism for agents to resolve overlap themselves before human review.

lazymaxxing agentic coding: When i have multiple agents working on the same code repo, if there's conflicts/overlaps, they just figure it… All the people talking to me about worktrees. Yes, exactly. I want the agents to just figure that out they're overlapping and to do it th…
@jason
  • Jason Calacanis floated that Hugging Face may be for sale and suggested AWS or SpaceXAI as potential buyers; the post provides no confirmation, valuation, or transaction details.
  • He characterized Nvidia's activity as a $6B bet on open source and a “full stack” strategy, without specifying whether this refers to an investment, acquisition, or another commitment.
we are live talking about the insane world robot games... it's the end of days folks! Also, Hugging Face is for Sale?!?!?!?!?! AWS or Spa…
Aravind Srinivas

Perplexity AI appointed Andrew Gordon Wilson as research lead, reporting to Denis. He will lead new work on continual learning, synthetic data, long-horizon reinforcement-learning environments, and architectures; the team is also hiring for this research agenda.

Excited to welcome Andrew Gordon Wilson to our research team. He will be reporting to Denis and lead new research efforts on continual le…
a16z
  • Frédéric Renken and Steijn Pelle spent a year doing manual back-office work at doctors’ offices before founding LassieAI to automate those workflows, indicating firsthand healthcare-administration experience behind the product.
Frédéric Renken and Steijn Pelle during the year they spent doing manual back office work at doctors' offices, before starting [@LassieAI…