ZeroNoise Logo zeronoise
Post
AI products hit two limits: the reliability layer and model costs
•
4 min read
• 305 docs
Several sources argue that AI product work now depends on who supplies reliability, how much the model layer costs, and whether evals measure real work. Also covered: agent loops for PMs and how mature products can compete with AI-native demos.

Who supplies the reliability?

Kendra Vant, CPO at Tapi, cites Princeton research finding that models have become more accurate but not more consistent: ask the same task five times and the answers differ . When she uses Claude Cowork, she corrects it and re-prompts. In her words, the user is "providing the reliability layer" . Her own customers, busy property managers, "just want my software" to work the first time .

So she says "can the model do it?" is the wrong question. The right one is where the reliability layer sits and whether you can afford to build it . Her checklist:

  • What are you quietly fixing when you use similar AI yourself?
  • Will your users do that work for you?
  • Who owns the layer: engineering, product, or nobody (in which case it's still a demo) ?
  • What does it cost in trust when the layer fails, and does it fail gracefully ?
  • What does the layer cost to run at scale? She has seen evals, model checkers, and human reviewers that worked at demo scale cost more than the product takes to run .

Model dependence is vendor risk

On r/ProductManagement, one post pointed to reporting on Anthropic's IPO prospectus and asked what PMs will do if the models behind their "wrapper" features become unaffordable . The top advice was to treat this as vendor risk. Put the model behind your own interface so changing providers is a config change, and don't let one vendor's pricing become a single point of failure . One B2B team uses "bring your own API key": the customer owns the provider relationship and pays the token bill, which also makes data privacy easier . Another commenter warned that switching isn't free, because any change of model or configuration changes agent behavior and quality at scale .

An a16z episode adds data on cost. Databricks' router, which picks a model for each task, solved more problems at 35% lower cost than the strongest single model. Elise AI fine-tuned a smaller model that was 60% cheaper and had low enough latency for live audio . The speakers argue the unit that matters is the cost of getting the customer's job done, not the cost of using the top model for every step . Their adoption figures: 69% of S&P 500 companies have live AI deployments, 30% report quantifiable impact, and only 2% track a metric over time .

Evals that look like real work

Hiten Shah traces how Harvey went from wrapper to legal product. It built legal process into the product, then created BigLaw Bench, which grades real legal tasks. That let the team test model, workflow, and retrieval changes against the same tasks, and turn failures lawyers spotted into tests . On Harvey's Legal Agent Benchmark of longer tasks, frontier models completed under 10% end to end under strict scoring. Shah says that kind of result shows whether the bottleneck is context, workflow, tools, or the model .

Aakash Gupta makes the same argument for personal workflows: "A prompt is text. A skill is behavior." He runs 25 skills against 75 tests. He added the tests after an exec-update skill filed a Stripe bug that was corrupting a metric under "wins" . Separately, Shah open-sourced 16 competitive-intelligence skills, covering positioning, pricing, battlecards, win/loss, and more .

Agent loops for PMs

Gupta lists six recurring loops: feedback digest, deal intelligence, competitor brief, metric anomaly flag, onboarding friction, and a customer call list . He adds that one agent should rank the backlog while a second checks the ranking against strategy . Roadmap bets and reading stakeholders stay with the PM . His rules: put every correction into the skill file, not the chat, and retire any loop whose output you've stopped editing .

Also worth noting

  • Competing with AI-native demos. A mature B2B team keeps losing the first impression to demos built on one clean prompt . One suggestion: run the same appealing prompt, then have the buyer change the source data after the output was approved, so your controls show up as part of real work .
  • Outcome-based pricing. HP's Faisal Masud describes pricing that is either seat-based with outcome commitments or purely outcome-based, tied to reductions in support tickets. His reason: CIOs no longer want to pay for dashboards . He also says AI compresses work from idea to PRD to demo, but getting to production is "a whole different ballgame" .
  • Cutting interruptions from QA. One PM added a "what happens if" section to every spec (empty states, no connection, wrong input, permissions), and QA questions mostly stopped after two sprints .
  • Decision-making as the core craft. Lenny Rachitsky highlighted Robby Stein's line that PM craft is "above all else… decision-making" , and Shreyas Doshi endorsed it .
AI products hit two limits: the reliability layer and model costs
Summary
Coverage start
1 day ago
Coverage end
16 hours ago
Frequency
Daily
Published
15 hours ago
Reading time
4 min
Research time
3 hrs 27 min
Documents scanned
305
Documents used
16
Citations
29
Sources monitored
98 / 99
Insights
Skipped contexts
Source details
Source Docs Insights Status
rahulvohra 0 0
Paul Graham 10 4
Tony Fadell 0 0
Patrick Collison 5 1
Daniel Ek 0 0
Gustaf Alströmer 0 0
Stewart Butterfield 0 0
PM Diego Granados 0 0
👨🏻‍💻☕️ 0 0
scott belsky 0 0
Ryan Hoover 6 2
Janna Bastow simplybastow.bsky.social 0 0
Jackie Bavaro 0 0
Sachin Rekhi 0 0
Dan Olsen 0 0
The community for ventures designed to scale rapidly | Read our rules before posting ❤️ 112 8
Will Lawrence 0 0
Product Marketing 12 2
Ami Vora 0 0
PM Interview: Practice Group for Product Manager Case Interviews 0 0
One Knight in Product 0 0
Aakash Gupta 4 1
Shreyas Doshi's Product Almanac | Substack 0 0
Lenny Rachitsky 0 0
Acquired 0 0
a16z 1 1
Exponent 0 0
Product Alliance 0 0
Product Management Exercises 0 0
rocketblocks 0 0
Product Design 0 0
ProductManagementJobs 13 3
Product Management 111 7
Product Management - The place for all things product 10 1
Product Management 0 0
Aspiring and current tech PM's 0 0
Masters of Scale 0 0
Product Science Group 0 0
How I built This 0 0
SaaStr AI 0 0
productized io 0 0
Lenny's Reads 0 0
The Product Folks 0 0
Strategyzer 0 0
Lenny's Podcast 0 0
AJ&Smart 0 0
Y Combinator 0 0
Product School 1 1
Mind the Product 1 1
@andrewchen 0 0
The Looking Glass 0 0
Leah’s ProducTea 0 0
Run the Business 0 0
Product Managers at Work 0 0
The Product Compass 0 0
Ravi on Product 0 0
Productify by Bandan 0 0
Product Thinking with Melissa Perri 0 0
Product Talk Daily 0 0
The Beautiful Mess 0 0
Gibson Biddle's "Ask Gib" Product Newsletter 0 0
Casey Accidental 0 0
Hiten Shah 3 2
Product Growth 0 0
Perspectives 0 0
Lenny's Newsletter 0 0
andrew chen 1 1
Brian Balfour 0 0
Casey Winters 0 0
elena verna 0 0
Kevin Weil 🇺🇸 0 0
April Underwood 0 0
Julie Zhuo 0 0
Marty Cagan 0 0
Lenny Rachitsky 9 4
Christian Idiodi 0 0
John Cutler 0 0
Teresa Torres 0 0
Gibson Biddle 0 0
Shreyas Doshi 6 3
Adam Nash 0 0
Merci Grace 0 0
Jackie Bavaro 0 0
Hunter Walk 0 0
Brian Balfour 0 0
Scott Belsky 0 0
Nir Eyal 0 0
Teresa Torres 0 0
Julie Zhuo 0 0
Andrew Chen 0 0
John Cutler 0 0
Ken Norton 0 0
Gibson Biddle 0 0
Elena Verna 0 0
Casey Winters 0 0
Shreyas Doshi 0 0
Lenny Rachitsky 0 0
Melissa Perri 0 0
Marty Cagan 0 0