ZeroNoise Logo zeronoise
Post
AI products hit two limits: the reliability layer and model costs
•
4 min read
• 305 docs
Several sources argue that AI product work now depends on who supplies reliability, how much the model layer costs, and whether evals measure real work. Also covered: agent loops for PMs and how mature products can compete with AI-native demos.

Who supplies the reliability?

Kendra Vant, CPO at Tapi, cites Princeton research finding that models have become more accurate but not more consistent: ask the same task five times and the answers differ . When she uses Claude Cowork, she corrects it and re-prompts. In her words, the user is "providing the reliability layer" . Her own customers, busy property managers, "just want my software" to work the first time .

So she says "can the model do it?" is the wrong question. The right one is where the reliability layer sits and whether you can afford to build it . Her checklist:

  • What are you quietly fixing when you use similar AI yourself?
  • Will your users do that work for you?
  • Who owns the layer: engineering, product, or nobody (in which case it's still a demo) ?
  • What does it cost in trust when the layer fails, and does it fail gracefully ?
  • What does the layer cost to run at scale? She has seen evals, model checkers, and human reviewers that worked at demo scale cost more than the product takes to run .

Model dependence is vendor risk

On r/ProductManagement, one post pointed to reporting on Anthropic's IPO prospectus and asked what PMs will do if the models behind their "wrapper" features become unaffordable . The top advice was to treat this as vendor risk. Put the model behind your own interface so changing providers is a config change, and don't let one vendor's pricing become a single point of failure . One B2B team uses "bring your own API key": the customer owns the provider relationship and pays the token bill, which also makes data privacy easier . Another commenter warned that switching isn't free, because any change of model or configuration changes agent behavior and quality at scale .

An a16z episode adds data on cost. Databricks' router, which picks a model for each task, solved more problems at 35% lower cost than the strongest single model. Elise AI fine-tuned a smaller model that was 60% cheaper and had low enough latency for live audio . The speakers argue the unit that matters is the cost of getting the customer's job done, not the cost of using the top model for every step . Their adoption figures: 69% of S&P 500 companies have live AI deployments, 30% report quantifiable impact, and only 2% track a metric over time .

Evals that look like real work

Hiten Shah traces how Harvey went from wrapper to legal product. It built legal process into the product, then created BigLaw Bench, which grades real legal tasks. That let the team test model, workflow, and retrieval changes against the same tasks, and turn failures lawyers spotted into tests . On Harvey's Legal Agent Benchmark of longer tasks, frontier models completed under 10% end to end under strict scoring. Shah says that kind of result shows whether the bottleneck is context, workflow, tools, or the model .

Aakash Gupta makes the same argument for personal workflows: "A prompt is text. A skill is behavior." He runs 25 skills against 75 tests. He added the tests after an exec-update skill filed a Stripe bug that was corrupting a metric under "wins" . Separately, Shah open-sourced 16 competitive-intelligence skills, covering positioning, pricing, battlecards, win/loss, and more .

Agent loops for PMs

Gupta lists six recurring loops: feedback digest, deal intelligence, competitor brief, metric anomaly flag, onboarding friction, and a customer call list . He adds that one agent should rank the backlog while a second checks the ranking against strategy . Roadmap bets and reading stakeholders stay with the PM . His rules: put every correction into the skill file, not the chat, and retire any loop whose output you've stopped editing .

Also worth noting

  • Competing with AI-native demos. A mature B2B team keeps losing the first impression to demos built on one clean prompt . One suggestion: run the same appealing prompt, then have the buyer change the source data after the output was approved, so your controls show up as part of real work .
  • Outcome-based pricing. HP's Faisal Masud describes pricing that is either seat-based with outcome commitments or purely outcome-based, tied to reductions in support tickets. His reason: CIOs no longer want to pay for dashboards . He also says AI compresses work from idea to PRD to demo, but getting to production is "a whole different ballgame" .
  • Cutting interruptions from QA. One PM added a "what happens if" section to every spec (empty states, no connection, wrong input, permissions), and QA questions mostly stopped after two sprints .
  • Decision-making as the core craft. Lenny Rachitsky highlighted Robby Stein's line that PM craft is "above all else… decision-making" , and Shreyas Doshi endorsed it .

Want personalized briefs on the topics you care about?

Create your own agent and get daily or weekly cited briefs on the topics you care about, shaped by the people, newsletters, podcasts, blogs, and channels you choose.