ZeroNoise Logo zeronoise
Post
Agents Are Becoming Both the Customer and the Competitor
8 hours ago
6 min read
2723 docs
The strongest signals are strategic capital moving into physical AI, agents beginning to replace narrow B2B products, and a widening split between cheap specialized intelligence and expensive frontier-model economics.

1. Funding & Deals

Uber’s investment in Zipline is a distribution-led physical-AI bet, not a clean early-stage comp. Zipline says Uber is becoming an investor while it scales to more than 1 million autonomous Uber Eats deliveries per day. Jason Calacanis says the partnership gives Zipline access to Uber Eats’ existing partner network, customer support, and customer-acquisition infrastructure; he also discloses that he put millions into a recent late-stage Zipline round, so the current evidence supports a strategic-capital signal rather than a Seed/Series A valuation benchmark.

The diligence question is whether the urban last-mile model travels beyond Zipline’s rural medical-delivery track record: public discussion explicitly distinguishes those environments, while another response flags noise and privacy as adoption risks.

2. Emerging Teams

TryNearbyCom has the clearest current early-stage traction signal. The YC S26 startup says it is live with more than 120 paying restaurants across Southern California; after 10 months of restaurant visits with creators and conversations with hundreds of owners, it reports more than 130% growth since the batch started and over 90% retention since November. The signal is a potentially repeatable local-creator distribution model; before underwriting it, verify the retention cohort, revenue quality, and restaurant-level payback.

Forge is a founder-signal rather than a traction story. A 17-year-old developer says he has built a C++ deep-learning framework from scratch since January, loaded real GPT-2 weights into a Forge implementation, and matched Hugging Face’s output token-for-token. The project is still CPU-only, lacks a KV cache, and is working toward CUDA and performance fixes, so the investment question is whether the unusual systems depth can become a team and product rather than whether this is already a company.

Agent Facets is an early agent-supply-chain thesis. Its builder argues that public agent “skills” should be managed as dependencies, with version ranges or pins, immutable artifacts, integrity checks, and reproducible installs across Claude Code, Codex, OpenCode, and other clients. The motivation is credible infrastructure pain: the post cites reports of malicious skills and a Snyk figure that 36% of scanned skills contained prompt injection, while explicitly saying that figure may not be fully accurate. The current signal is the category definition—portable, reviewable agent capabilities—not disclosed revenue.

3. AI & Tech Breakthroughs

A JAMA study puts autonomous clinical AI on the agenda, but not yet in production. In a public summary of the paper, Khosla and coauthors report 159 simulated OSCE cases in which physicians rated Google’s AMIE better than physicians at eliciting complaints (97% vs. 50%), systems review (88% vs. 35%), medical history (85% vs. 50%), family history (50% vs. 21%), and medication history (68% vs. 45%). The authors argue that AI-alone may eventually outperform physician-only or physician-AI hybrids in some cognitive workflows, while listing workflow, liability, regulation, reimbursement, and medical education as unresolved barriers and pointing to possible deployment in some workflows by 2030. Because the evaluation is simulated and the claims here come through the authors’ summary, the near-term investment signal is in evaluation, governance, and clinical deployment infrastructure—not proof that autonomous care is ready.

Faraday is a more concrete step toward scientific agents, while open-ended discovery remains a separate gate. The paper’s abstract describes Replica, a scalable paper-replication task space with an auto-generated rubric judge, and Faraday, a 27B agent that uses coding agents as tools and surpasses Claude Opus 4.8 and GPT-5.5 on held-out replication tasks. Import AI reports that Replica covers 100 ML and AI-for-science papers converted into 310 tasks, with Faraday exceeding the comparison systems on 73% of in-distribution ML tasks and 60% of held-out AI-for-science tasks according to its rubric judge. By contrast, DiG-bench’s 70 hidden-rule games remain difficult for frontier models: only Opus 5 and Fable 5 with Claude Code beat any Tier 7 task, while individual humans reached 100% on the tests. Replication may therefore be an earlier commercial wedge than genuinely open-ended scientific discovery.

CR-NN is an unverified low-memory attention bet. Its author claims matrix-free attention with O(N log N) complexity and a 16.2× speedup over Flash at 50K tokens, plus 0.015 GB at 1.36M tokens versus 12.3 GB for a KV cache; the project includes negative results and is explicitly seeking collaborators to validate the idea at scale. Treat the numbers as a replication target, not an established efficiency result.

4. Market Signals

AI demand is expanding, but usage is highly concentrated and frontier-model spend may be nearing a ceiling. Exponential View reports July AI revenues at three times the prior year’s level and an annualized run-rate above $210 billion. Its tracking also says the top 10% of OpenAI enterprise users consume 8.3× as many tokens as the typical firm, while Fable 5 usage is flat at only 6% of business tokens and 11% of spend—an attributed signal that the best model is not automatically the economic default.

Agentic replacement is now visible in churn, while some incumbents are using agents to expand. SaaStr says it canceled Notion after seven years because its internally built 10K agent absorbed Notion’s remaining job, running the Monday staff meeting from revenue, campaign, collections, and pipeline data. It argues that quiet, low-touch accounts are especially exposed because a narrow workflow is easier for an agent to absorb, and says Marketo lost a 10-year relationship after its API stopped working for SaaStr’s agents. The counterexample is Stripe: it says agents wrote 30% of its code in a week, cut global tax-filing time to one-third of the U.S. version, made sellers 20% more productive, and led it to hire more sellers. The practical split is between AI that unlocks new capacity and AI that makes a narrow incumbent product unnecessary.

Open-model economics are becoming a strategic fault line. Interconnects describes open-model training as highly capital intensive and says Nvidia is reportedly spending $26 billion to create a broad ecosystem of model builders and inference demand, while acknowledging that it is unclear whether the strategy will pay off. Its base case is a bifurcation: closed labs retain the most valuable knowledge-work, drug-discovery, and software-engineering markets, while open models specialize in efficient, modifiable, enterprise-specific agents running on private data; revenue-share licenses are being tested to keep near-frontier open-weight development financeable.

Agent commerce will probably reuse existing payment rails, leaving authorization and exception handling as the wedge. A SaaS discussion favors agents using ordinary checkout rather than requiring every merchant to build a new API, but identifies CAPTCHA/3DS flows, subscription permissions, and fraud or chargeback liability as unresolved problems. Perplexity’s current product direction points to the same control layer: users can set connector tools to Allow, Always Ask, or Deny, with recurring runs following thread-level approvals; CEO Aravind Srinivas frames this as keeping humans able to intervene.

5. Worth Your Time

  • Watch — Will Gaybrick on a16z. The conversation connects Stripe’s “build everything” posture, its 7,000 one-shot PRs per week, disappearing checkout pages, and stablecoin-enabled micropayments.

  • Read — Teaching Everyone to Fish for Tokens. The clearest current framing of Nvidia’s open-model strategy, the financing problem for open-weight labs, and the likely shift toward specialized on-prem agents.

  • Read — Training AI Scientists to Replicate Research. Start with the original abstract for Replica’s rubric-based evaluation and Faraday’s held-out replication result before accepting the broader AI-scientist thesis.

  • Try — GBrain. Garry Tan’s free, MIT-licensed project generates a personalized agent for Codex or Claude Code through a 12-question onboarding, installs 70 skills, and creates a private knowledge wiki.

Agents Are Becoming Both the Customer and the Competitor
Research extraction

Direct answer: The supplied bundle contains only the arXiv abstract page for arXiv:2608.13331 , not the full paper. It corroborates the paper's headline claims but does not contain enough detail to verify benchmark construction, harness setup, numeric replication results, or stated limitations.

  • Motivation/benchmark rationale: Replication "illuminates details that were previously underspecified" and "requires similar hypothesis-driven exploration to open-ended research" .
  • Benchmark construction (summary level): The paper introduces Replica, "a scalable task space for paper replication"; reward comes from "an auto-generated rubric-based judge that has low noise and agrees with human assessment of replication quality" . No task counts, construction steps, or rubric-evaluation details are provided in this source.
  • Model/harness setup (summary level): Faraday is "a 27B-parameter 'AI Scientist' agent that leverages coding agents as tools"; the paper frames it as a step toward "AI agents capable of long-horizon scientific innovation without requiring complex harnesses" . The only harness detail stated is use of coding agents as tools; no training recipe or tool architecture appears in the abstract.
  • Reported replication results: Faraday is reported as "surpassing the performance of Claude Opus 4.8 and GPT-5.5 on held-out replication tasks", and qualitative analysis of individual rollouts "reveals that Faraday adopts a more scientifically-principled approach" . No metrics, evaluation protocol, or comparison details are given in this source.
  • Gaps/uncertainty: Because only the abstract-level text is supplied, the benchmark construction, model/harness setup, exact replication results, and limitations cannot be independently verified from these materials; any granular claims would require the full paper, which is absent from the bundle .
Training AI Scientists to Replicate Research
a16z

Travis Kalanick (founder of Atoms and Uber) joined @davidsenra's podcast, which covers why specialized robots beat humanoids at industrial scale and the "Physical AI stack" in food and mining . The episode also includes what founders get wrong about venture capital and a fundraising playbook ("QED storytelling," "five-room auction") . In the excerpt a16z shared, Kalanick says leaders must find the line between order and chaos, keeping the fewest rules while staying out of chaos; at Uber, early city launches had hours-long pricing calls that became a five-minute playbook by city 20 .

My conversation with [@travisk](https://x.com/travisk), founder of Atoms and Uber. 0:00 Building Atoms & the Meta Problem of Management 3… Travis Kalanick says a leader's job is to find the line between order and chaos: "On one side is order. You have lots of rules, lots of s…
Garry Tan

YC President/CEO Garry Tan open-sourced GBrain, a free, MIT-licensed project for building a personal AGI, with new personalized agent generation for Codex and Claude Code: a 12-question onboarding creates a SOUL.md and installs 70 of his proven agent skills plus a Karpathy-style knowledge wiki ; it is free with existing Claude Code/Codex subscriptions and can bootstrap an agent repo from a pasted image . Repo: https://github.com/garrytan/gbrain.

At YC Startup School 2026, Tan argued the next generation of startups will be built by smaller teams than ever, powered by "personal AGI" — AI agents running on your own infrastructure that compound knowledge over time — and that founders should own their intelligence instead of renting it ; his talk covers "Building a Company of One" and why he open-sourced his stack .

GBrain now supports personalized agent generation/onboarding for Codex and Claude Code It'll generate an agent AI-style SOUL.md and insta… What do you get? A private github repo with 70 of my proven skills and the beginnings of your Karpathy-style knowledge wiki. Read the ful… It's free to try and works with your existing Claude Code or Codex subscription right now. Just make a new directory and start Codex or C… The next generation of startups will be built by smaller teams than ever before. At Startup School 2026, YC's [@garrytan](https://x.com/g… This is my open source to help you create your own Personal AGI as mentioned at my Startup School talk earlier this year [https://x.com/y…
a16z

Stripe President of Technology & Business Will Gaybrick, interviewed by a16z's David George, argues AI agents should be a build-more lever, not a cost cutter: agents wrote 30% of code in a week, global tax filing shipped in 1/3 the time the US version took, and after AI made sellers 20% more productive, Stripe hired more sellers . The interview also covers Stripe's "Minions" agents generating 7K one-shot PRs a week and why software timelines keep compressing .

Gaybrick argues the agentic economy requires microtransactions: agents will navigate the internet like "little hummingbirds," slurping small amounts of data and compute, so humans won't need to create accounts everywhere ; microtransactions "will just be necessary for that economy to exist, the agentic economy" , and stablecoins make them possible — an agent with a budget can convert dollars into stables and check out . Content, he says, won't remain squeezed between subscription and free models . Related claims from the same interview: checkout pages will disappear; agents plus stablecoins make micropayments real; tokens are becoming a currency worth protecting like dollars ; Stripe is building Tempo, "a payments-only blockchain" ; and "tokens are money now" .

Stripe's Will Gaybrick: "Build everything" Against an industry that sees agents as a way to cut costs, Stripe is using them to build more… Stripe's Will Gaybrick says AI agents will navigate the internet like hummingbirds, and that's what will make microtransactions work: "If…
Michael Seibel

AI in government policy research is emerging as an investment-relevant theme: Michael Seibel argues AI will kill 'misguided sacred cows' — much policy has 'great intentions but no results,' is optimized for press and votes, and a 'reasonable amount' of city spending isn't accomplishing its stated goals; this is now easy to research with agents . He pushes politicians to cut ineffective spending, saying SF citizens back it and funds can be redirected , and offers to help pay for polling to overcome backlash fears .

Use of AI in government policy research is going to kill so many misguided sacred cows. It turns out so much policy had great intentions … Politicians should understand that the risks are lower than they fear for cutting ineffective spending. SF citizens have their back. Mone… Political consultants - who monetize winning elections might say risk is not worth potential electoral backlash. They are too self intere…
Vinod Khosla

Khosla advocates for AI-only physicians: he cites "not peer reviewed but some real life indicative proof" and argues that if the FDA allowed "AI physicians as a category with human performance as a predicate," real-world proof would follow quickly; he points to @CuraiHQ for AI intake and full process . He is a co-author of "A case for AI alone — autonomous medical care" (shared via @JAMA_current), with @ZekeEmanuel, @nealkhosla, and @AbeBakerButler . Eric Topol, whose post Khosla quotes, cautions "this remains unproven; none of the studies were in real world medicine" .

Not peer reviewed but some real life " indicative proof" at [https://x.com/nealkhosla/status/2036451113835405313](https://x.com/nealkhosl… A case for AI alone —autonomous medical care [@JAMA_current](https://x.com/JAMA_current), by [@ZekeEmanuel](https://x.com/ZekeEmanuel) [@…
a16z
  • Stripe's President of Technology & Business Will Gaybrick says Stripe uses AI agents to "build everything" rather than cut costs: agents wrote 30% of code in a week, global tax filing shipped in roughly 1/3 the time the US version took, and after AI made sellers 20% more productive, Stripe hired even more sellers .
  • On the same a16z podcast with David George, Gaybrick argues "there's no one left for Stripe to copy" and predicts checkout pages will disappear; he ties agents plus stablecoins to real micropayments and says tokens are becoming "a currency worth protecting like dollars" . Related claims in the episode: stablecoins "solve a political problem" and Stripe is building Tempo, a payments-only blockchain, with "tokens are money now" .
  • Stripe's internal agent system "Minions" generates 7K one-shot PRs a week, a signal of how deeply agentic coding workflows are embedded at a major payments company .
Stripe's Will Gaybrick: "Build everything" Against an industry that sees agents as a way to cut costs, Stripe is using them to build more…
a16z

Stripe President of Technology & Business Will Gaybrick argues AI is accelerating software creation, not commoditizing it: "we are seeing the exact opposite right now, where software creation is exploding. Customers are monetizing faster than ever" . Stripe reports agents wrote 30% of code in a week, global tax filing shipped in 1/3 the time the US version took, AI made sellers 20% more productive (after which Stripe hired more sellers), and an internal "Minions" system produces 7K one-shot PRs a week . In the a16z podcast he also argues checkout pages will disappear, agents plus stablecoins make micropayments real, and tokens are becoming currency worth protecting . Investment signal: Stripe's 2026 customer cohort is 50% larger and growing faster than the 2025 cohort, which was 70% larger and growing faster than the 2024 cohort .

A common thesis is software gets severely commoditized from here on out. Stripe's Will Gaybrick says their data shows the opposite: "It's… Stripe's Will Gaybrick: "Build everything" Against an industry that sees agents as a way to cut costs, Stripe is using them to build more…
@jason

Per @grok's explainer, quote-posted by @Jason: Section 8 of the Clayton Act bars interlocking directorates — no person (or firm via its agents, under the deputization theory) may serve as director/officer of two competing corporations above size thresholds — and DOJ probes these overlaps to stop potential information sharing or coordination between rivals . Enforcement has ramped up since 2022, especially against PE/VC board seats in overlapping portfolio companies; many cases end in quiet resignations and it is civil, not criminal, antitrust. The agencies are now checking whether Databricks and Fivetran — both of which handle enterprise data for AI uses — qualify as competitors . @Jason said the rule was new to him and asked @grok/human experts what the trigger benchmark is (https://x.com/grok/status/2089409681571676247) .

Section 8 of the Clayton Act bars interlocking directorates: no person (or firm via its agents under the deputization theory) may serve a… This is interesting, I had no idea. What’s the benchmark here for triggering Section 8 [@grok](https://x.com/grok) / human experts? [http…
Vinod Khosla

AI-in-healthcare market signal from Vinod Khosla: billing systems and codes determine more of actual care provision than AI's capability, making reimbursement the binding constraint on AI adoption in healthcare; Khosla expanded this thesis in a Khosla Ventures post from early 2025, 'The Opportunity for AI in Healthcare Is Here Now' (https://www.khoslaventures.com/posts/the-opportunity-for-ai-in-healthcare-is-here-now) . In the quoted reply, @chrissyfarr argues the science exists but there is no good way to pay for this in the current system, the FDA route makes no sense, and workable alternatives are cash membership, part of risk-bearing contracts, selling to employers as a vendor, or essentially free + fee-for-service — requiring a rethink of billing and coding .

Unfortunately, billing systems and codes determine more of care provision than actual ability of AI to provide. I covered this extensivel… [@vkhosla](https://x.com/vkhosla) [@ZekeEmanuel](https://x.com/ZekeEmanuel) [@AbeBakerButler](https://x.com/AbeBakerButler) [@nealkhosla]…
Natural Language Processing

An early-stage developer building a Flutter pronunciation-teaching app (~1,600 words across EN/ES/PT/IT/FR) with ElevenLabs TTS and cached audio reports a concrete TTS gap: eleven_multilingual_v2 has no sense/POS/context control; ElevenLabs pronunciation dictionaries are exact-string, case-sensitive, and lack POS scoping; phoneme tags exist only on eleven_flash_v2/v3, so switching models means re-synthesizing the cache and losing voice identity across five languages . Published IPA consistency of 80–90% for v3-class is too low for a teaching product , respelling hacks fail on stress-shift pairs and don't transfer to non-English languages , and STT-based grading can't catch heteronym mispronunciation because STT returns orthography only . The developer asks whether any TTS API accepts sense/POS hints or per-request phonemes on a multilingual model ; commenters note shipped apps either ignore the problem or silently curate datasets around it .

In practice, Cartesia's custom-pronunciation API was found more IPA-reliable than ElevenLabs, supports explicit lexical stress, and supports per-request inline IPA (<<...>>), though its dictionaries are still word-string-keyed . For on-device phoneme evaluation, ZIPA advertises cross-lingual phone models runnable on a phone, but accuracy collapses on isolated words (vowels turn into schwas), making forced alignment with wav2vec2 a more practical path; a commenter suggested force aligners can guide between true minimal-pair vowel options , which the developer reframed as scoring audio against two candidate pronunciations rather than open-vocabulary phone recognition .

Caveat: part-of-speech alone doesn't disambiguate heteronyms — both senses of "windy" are adjectives — and the practical fix is storing pronunciation as a per-entry field, affecting only a handful of ~1,600 entries .

TTS/STT can't tell "wind" from "wind" — how do you handle heteronyms in a pronunciation-teaching app? I’ve been tinkering with related things for a personal project, partly after becoming annoyed that shipped apps either don’t consider thi… Might be English only, but I have had better luck with Cartesia than Elevenlabs for IPA reliability, and you can denote lexical stress: [… This is the most useful reply I've had, thank you. Two things in it I want to separate, because I think they solve different halves. On C… > My word library actually knows which sense is on screen — every entry carries a part of speech — but there's no API surface to hand tha… You're right, and the example is well chosen: both readings of "windy" are adjectives, so POS gives me nothing. No tag distinguishes the …
Interconnects

Nvidia is spending $26B to foster nearly open-source model development, releasing data and training code (e.g., Nemotron) so many companies can build their own 'token machines,' driving demand for Nvidia chips rather than consolidating AI with Anthropic/OpenAI .

The open-source AI ecosystem's future is uncertain: building competitive models remains capital-intensive, and it's unclear if Nvidia's investment will generate enough profitable demand; exits like Databricks and 01.ai are seen as anomalies so far . If open-source training doesn't become financially self-sustaining, open models may diverge into long-tail use cases (enterprise-specific on-prem agents with private data), while closed players keep the most valuable areas (knowledge work, drug discovery, SWE) .

Open model builders are testing revenue-share licenses (e.g., Kimi-K3, Qwen3.8) to keep near-frontier open-weight development viable; success of these experiments is pivotal for Nvidia's demand-growth strategy .

Meta's open-weight strategy (e.g., releasing strong Muse Spark 1.2) is a direct competitive move that would hurt Anthropic/OpenAI's token revenue growth, whereas Nvidia wants a multi-company 'learn to fish' ecosystem .

The open model market currently sees an explosion in post-training/finetuning (DeepSeek V4 Flash, Inkling Small, GLM 5.X via APIs like Tinker), even as base-model training becomes more opaque and expensive, shifting the standard pretraining→midtraining→post-training taxonomy .

Teaching Everyone to Fish for Tokens
Aravind Srinivas

Perplexity's Computer agent product now lets users set each connector tool to Allow, Always Ask, or Deny, approve a single action or allow the tool for the rest of the thread, with recurring runs following thread approvals — available now on web for all Computer users . Perplexity CEO Aravind Srinivas frames the update as preserving human agency in AI agents: users should keep their hands on the wheel and be able to intervene when necessary . The release brings human-in-the-loop agent governance into a mainstream agent product, a small signal for agent-control and safety-layer startups.

Computer now lets you set each connector tool to Allow, Always Ask, or Deny. Approve one action or allow the tool for the rest of the thr… When building AI agents, it is important to still give humans agency to have their hands on the wheel and intervene when necessary. [http…
martin_casado

Cursor launched Origin, its code hosting platform, now live and described as fast, easy to use, and deeply integrated with Cursor, with support for syncing repos from GitHub . a16z general partner Martin Casado publicly cheered the launch as "Finally!! And perfectly timed" , marking Cursor's expansion from AI coding assistant into code-hosting infrastructure — a competitive move against GitHub in the AI dev-tools stack.

Origin, our code hosting platform, is now live. It's fast, easy to use, and deeply integrated with Cursor. Get started by syncing your re… Finally!! And perfectly timed launch 🥳 [https://x.com/cursor_ai/status/2089399057659596847](https://x.com/cursor_ai/status/20893990576595…
Leo Polovets

VC @lpolovets now ignores all inbound decks and pitches that are 90+% AI, despite initially being open-minded; he treats sloppy or lazy fundraising as a high-leverage signal that execution will disappoint elsewhere ("how you do anything is how you do everything") .

I now ignore all inbound decks &amp; pitches that are 90+% AI. Initially I was open-minded, but as the aphorism goes: how you do anything…
Vinod Khosla

Vinod Khosla: whoever gets paid should sign the liability waiver; physicians have insurance, and AI can be similarly insured as in self-driving cars . He said this in response to a joke about making humans sign the liability waiver at the end .

Whoever gets paid should sign the liability waiver. Physicians have insurance and so can AI as it does in self driving cars. [https://x.c… [@vkhosla](https://x.com/vkhosla) nothing says teamwork like making the human sign the liability waiver at the end
Vinod Khosla

Vinod Khosla, responding to praise for his decade-old essay on AI in medicine, says the vision "was laughed at back then" but that 2035 "could be really exciting" for medicine. He forecasts that every ECG will be worth ten times more, every EEG will disclose hidden and prospective future diseases, and that medicine will shift to predicting disease before it happens, rather than after symptoms become dominant and too late to act . The admiring reply calls the old article "the most accurate prediction of the complete transformation of medicine and healthcare in the age of AI" and "prophetic" .

I appreciate it, but it was laughed at back then. 2035 could be really exciting in medicine, looking ten years forward. Every ECG will be… This article from Vinod, written 10 years ago, is more than directionally correct; it is the most accurate prediction of the complete tra…
Aravind Srinivas

Perplexity CEO Aravind Srinivas highlights that dense models are still slow to run on local hardware, but calls a recent demonstration "incredible" and "a sign of things to come soon-ish," signaling accelerating local AI inference progress . The quote references a post by @ggerganov captioned "let that sink in" with an image, underscoring a notable milestone in local model execution .

Dense models are slow to run on local hardware, but this is incredible and a sign of things to come soon-ish. [https://x.com/ggerganov/st… let that sink in ![](https://pbs.twimg.com/media/HP8TMV7XsAA0FYY.jpg)
Vinod Khosla

Khosla and co-authors published in JAMA a comparison of Google's Articulate Medical Intelligence Explorer (AMIE) with physicians across 159 simulated OSCE cases spanning multiple specialties: AMIE was rated significantly better at eliciting complaints (97% vs 50% favorable), systems review (88% vs 35%), medical history (85% vs 50%), family history (50% vs 21%), and medication history (68% vs 45%), all P < .001 . They also report that LLMs with customized architectures markedly outperform naive architectures on diagnosis, test ordering, and guideline-concordant treatment and reduce safety incidents, implying existing AI-performance studies may systematically underestimate AI . Khosla calls on CMS to weigh whether autonomous AI will exceed AI-aided physicians as the best medical care .

Game mostly over for human doctors vs. AI? [@ZekeEmanuel](https://x.com/ZekeEmanuel) [@AbeBakerButler](https://x.com/AbeBakerButler) [@ne…
Vinod Khosla

Asked whether AI alone will beat a doctor plus AI in delivering medical care , Vinod Khosla answered “Absolutely,” citing Stanford studies by Arnie Milstein’s team showing that “Humans degrade the performance of good AI” . This signals Khosla’s backing for autonomous AI over human-in-the-loop clinician workflows in healthcare.

What are your thoughts? Will AI alone beat a doctor plus AI in delivering medical CARE? [https://x.com/vkhosla/status/2089380431347282268… Absolutely. Look at the studies Arnie Milstein and his team have done at Stanford. Humans degrade the performance of good AI. [https://x.…