ZeroNoise Logo zeronoise
Post
AI products hit two limits: the reliability layer and model costs
•
4 min read
• 305 docs
Several sources argue that AI product work now depends on who supplies reliability, how much the model layer costs, and whether evals measure real work. Also covered: agent loops for PMs and how mature products can compete with AI-native demos.

Who supplies the reliability?

Kendra Vant, CPO at Tapi, cites Princeton research finding that models have become more accurate but not more consistent: ask the same task five times and the answers differ . When she uses Claude Cowork, she corrects it and re-prompts. In her words, the user is "providing the reliability layer" . Her own customers, busy property managers, "just want my software" to work the first time .

So she says "can the model do it?" is the wrong question. The right one is where the reliability layer sits and whether you can afford to build it . Her checklist:

  • What are you quietly fixing when you use similar AI yourself?
  • Will your users do that work for you?
  • Who owns the layer: engineering, product, or nobody (in which case it's still a demo) ?
  • What does it cost in trust when the layer fails, and does it fail gracefully ?
  • What does the layer cost to run at scale? She has seen evals, model checkers, and human reviewers that worked at demo scale cost more than the product takes to run .

Model dependence is vendor risk

On r/ProductManagement, one post pointed to reporting on Anthropic's IPO prospectus and asked what PMs will do if the models behind their "wrapper" features become unaffordable . The top advice was to treat this as vendor risk. Put the model behind your own interface so changing providers is a config change, and don't let one vendor's pricing become a single point of failure . One B2B team uses "bring your own API key": the customer owns the provider relationship and pays the token bill, which also makes data privacy easier . Another commenter warned that switching isn't free, because any change of model or configuration changes agent behavior and quality at scale .

An a16z episode adds data on cost. Databricks' router, which picks a model for each task, solved more problems at 35% lower cost than the strongest single model. Elise AI fine-tuned a smaller model that was 60% cheaper and had low enough latency for live audio . The speakers argue the unit that matters is the cost of getting the customer's job done, not the cost of using the top model for every step . Their adoption figures: 69% of S&P 500 companies have live AI deployments, 30% report quantifiable impact, and only 2% track a metric over time .

Evals that look like real work

Hiten Shah traces how Harvey went from wrapper to legal product. It built legal process into the product, then created BigLaw Bench, which grades real legal tasks. That let the team test model, workflow, and retrieval changes against the same tasks, and turn failures lawyers spotted into tests . On Harvey's Legal Agent Benchmark of longer tasks, frontier models completed under 10% end to end under strict scoring. Shah says that kind of result shows whether the bottleneck is context, workflow, tools, or the model .

Aakash Gupta makes the same argument for personal workflows: "A prompt is text. A skill is behavior." He runs 25 skills against 75 tests. He added the tests after an exec-update skill filed a Stripe bug that was corrupting a metric under "wins" . Separately, Shah open-sourced 16 competitive-intelligence skills, covering positioning, pricing, battlecards, win/loss, and more .

Agent loops for PMs

Gupta lists six recurring loops: feedback digest, deal intelligence, competitor brief, metric anomaly flag, onboarding friction, and a customer call list . He adds that one agent should rank the backlog while a second checks the ranking against strategy . Roadmap bets and reading stakeholders stay with the PM . His rules: put every correction into the skill file, not the chat, and retire any loop whose output you've stopped editing .

Also worth noting

  • Competing with AI-native demos. A mature B2B team keeps losing the first impression to demos built on one clean prompt . One suggestion: run the same appealing prompt, then have the buyer change the source data after the output was approved, so your controls show up as part of real work .
  • Outcome-based pricing. HP's Faisal Masud describes pricing that is either seat-based with outcome commitments or purely outcome-based, tied to reductions in support tickets. His reason: CIOs no longer want to pay for dashboards . He also says AI compresses work from idea to PRD to demo, but getting to production is "a whole different ballgame" .
  • Cutting interruptions from QA. One PM added a "what happens if" section to every spec (empty states, no connection, wrong input, permissions), and QA questions mostly stopped after two sprints .
  • Decision-making as the core craft. Lenny Rachitsky highlighted Robby Stein's line that PM craft is "above all else… decision-making" , and Shreyas Doshi endorsed it .
AI products hit two limits: the reliability layer and model costs
Hiten Shah
  • Harvey’s product progression suggests building domain workflow into the application before relying on deeper model specialization. Its BigLaw Bench evaluated real legal work—including facts and sources, analysis, and finished-work structure—and let the team test model, workflow, and retrieval changes against the same tasks; lawyer-identified failures could become tests for later versions.
  • Long-horizon, end-to-end evaluations can reveal failures that polished partial outputs hide: Harvey reported that frontier models completed less than 10% of its Legal Agent Benchmark tasks in aggregate under strict all-pass scoring. The article recommends using such failures to determine whether the bottleneck is context, workflow, tools, or the model.
  • Shah’s AI product roadmap puts specialized intelligence after workflow, context, evaluations, and reliable execution, followed by learning from expert corrections and adapting to organizational knowledge. Harvey’s August 2026 Tenet launch—a post-trained open-weight model for long-horizon legal work—illustrates specialization against capabilities the product had already learned to measure.
The Wrapper That Kept Going
Product School
  • HP product teams are encouraged to use AI where it fits and keep humans in the loop. AI has compressed work from ideation through PRDs and early demos, but getting products into production remains a separate challenge; HP’s goal is to shorten the full cycle from ideation to market.
  • HP’s view is that AI should support product strategy, not replace it: it can inform and consult, but copy-paste use is insufficient and human judgment and original ideas remain important. Adoption targets alone have not ensured gains, and tools can add work; token costs also make ROI a necessary check.
  • HP made its Workforce Experience platform device- and OS-agnostic and sells it around customer problems through CIOs and RFPs, rather than bundling software with hardware. HP dogfooded it in its 80,000-device IT environment, used a long alpha, beta, and extended-release process before GA, and says its customer pitch includes a 30–40% reduction in ticket volume.
  • HP has an AI Command Center alpha in its labs, intended as a future Workforce Experience elite capability to track employees’ AI adoption, usage, and yield and help customers manage budgets and efficiency.
  • To account for fewer human seats and focus on realized value, HP sees options including seat-based pricing with outcome commitments or fees tied directly to reductions in support tickets; the interview argues customers increasingly expect demonstrable ROI rather than paying for dashboards alone.
HP President on Replacing Middle Management with AI Agents | Faisal Masud | E314
  • Enterprise AI adoption remains shallow despite broad deployment: 69% of S&P 500 companies have live AI deployments, 30% report quantifiable impact, and 2% track a metric over time; the opportunity is to turn isolated deployments into influential recurring workflows with reliable applications that bridge general model capabilities and company-specific knowledge.
  • Product teams can improve AI economics through task-appropriate model routing and fine-tuning: Databricks’ smart router solved more problems at 35% lower cost than the strongest individual model, while Elise AI’s fine-tuned smaller model was 60% cheaper and had lower latency, making live audio use cases feasible. Optimizing the cost of completing a customer’s job can avoid using the most expensive model for every step.
  • Reported customer outcomes include Chime reducing cost to serve by over 10% a year for four years (described as almost a 50% overall reduction) and Shopify’s AI Sidekick increasing by 8% the share of new customers reaching five orders within 15 days of onboarding.
  • For SaaS products, applying AI to an existing customer base and distribution can support revenue growth; one proposed benchmark was 10% or greater growth acceleration over the next 12–18 months. The discussion frames revenue acceleration—not efficiency improvements alone—as evidence that AI investments are strengthening the business.
  • Consumer agents make conventional engagement measurement less complete: background tasks may not register as screen time, even as assistants aim to become persistent and proactive.
AI, Infrastructure, and the Next Investment Cycle
Shreyas Doshi

Shreyas Doshi endorsed the view that product management’s central craft is decision-making—not simply doing many things reasonably well—and that PMs should study what makes decisions great.

This is the way (and always has been). Great point by [@rmstein](https://x.com/rmstein) [https://x.com/lennysan/status/210538881381629973… One of my favorite lessons from [@rmstein](https://x.com/rmstein)'s talk: "People say the PM's main craft is that you do lots of things p…
Mind the Product
  • AI reliability is part of the product, not a demo detail: Research discussed in the episode suggests models have become more accurate without becoming more consistent, and users can get useful results by correcting and guiding them. That makes power users who help fix outputs a different test case from customers who expect a task-focused product to work correctly the first time.
  • Assess the reliability layer before committing to an AI feature: Identify what users currently do to compensate for AI gaps, whether target users will do that, and who will build and own the layer. Also define failure costs, graceful recovery, and the ongoing cost of maintaining checks or human review at scale; a system that works at demo scale may not be economically viable in production.
  • Product managers need software literacy, not necessarily coding ability: Being comfortable with Git and learning from a codebase can improve product work, while confusing a convincing prototype with production-ready software risks overlooking substantial remaining work.
Four questions to ask before building AI into your product—Kendra Vant (Chief Product Officer, Tapi
Shreyas Doshi

Shreyas Doshi announced a private playlist of 23 candid videos answering career questions and dilemmas, available free by request through Maven; the announcement does not specify that the videos are product-management-specific.

Announcing a brand new private resource: this is a playlist of 23 super candid videos of mine, with answers to your real career questions…
Shreyas Doshi

Shreyas Doshi added three YouTube videos aimed at ambitious, advanced-career product professionals; his career-video list is sorted by views and begins with “How to talk to executives.”

PSA: in case you missed this list of videos when I shared them last week — I’ve since added 3 new videos on the channel. These videos are… I’ve recently released a bunch of career-related videos on YT but haven’t shared any of them here on X. So here they are (sorted by views…
Lenny Rachitsky

A Meta PM who had contacted someone for help with an upcoming OpenAI interview loop decided to stay at Meta to work on Muse; the source says this would not have happened two months earlier. Lenny Rachitsky framed the anecdote as “Meta ascendant,” though it is a single-person signal, not evidence of a broader hiring trend.

interesting datapoint for yall: Meta PM pinged me for help w/ upcoming OAI loop a week ago, and today told me they decided to stay at Met… Meta ascendant [https://x.com/viableben/status/2105465935934939644](https://x.com/viableben/status/2105465935934939644)
andrew chen

Andrew Chen contrasts the PC’s rise, which he says made $20k+ 1980s workstations such as SGI, NeXT, and Sun a distant memory, with today’s premium setups: a PC paired with an RTX PRO 6000 or a Mac Studio; he says prices are rising. This is a hardware-positioning and pricing signal, not evidence of sales or demand levels.

when the personal computer won, it made $20k+ workstations in 1980s a distant memory - SGI, Next, Sun, etc. I remember thinking, huh why …
Product Management
  • Product managers report using AI for deep research, SQL queries, note-taking, personal task tracking, language cleanup, decks and one-pagers, interactive documentation, prototyping, and working through difficult concepts.
  • One team says AI can produce mockups and prototypes that used to take a full day during a meeting or within 30 minutes, enabling faster feedback and better requirements; recording demos can also be turned into knowledge-transfer articles. AI document search can surface older project files, including executed SOWs, service schedules, presentations, and recordings.
  • AI is also used to draft weekly and ad-hoc leadership updates by scanning Gmail, Drive, Slack, Jira, and Confluence. For project plans, it can provide a useful starting structure—especially for novel work—but one commenter says they do not yet trust the results.
This sub is full of examples of you search "AI" Deep research, writing SQL queries, note taking, personal task track, language clean up, … What does your team composition look like? If you’re going the builder route, then you can just build. Not recommended unless you have a … The biggest time save is writing the various LT updates each week, including ad-hoc ones. It'll scan Gmail, Drive, Slack, Jira, Confluenc…
Product Management
  • For AI-related development, one PSPO I holder said the credential is not essential and advised showing AI tools built or tasks automated to free time for product strategy. A highly upvoted reply recommends storytelling training over another PM certification as a way for PMs to maintain an edge.
  • Product-owner certifications have mixed, context-dependent value: one manager sends new PM hires through a two-day CSPO course for Scrum basics, while another practitioner says a PO certification helped with value framing and breaking epics into workflows. Other commenters say such training is relevant only where Scrum is used and does not replace practical experience; one hiring manager considers PO certifications on a PM candidate’s resume a red flag.
  • One respondent who held PSM, PSPO, SAFe Agilist, and PMP credentials preferred PMP for learning about stakeholder, team, and sponsor interactions.
I have PSPO I but I don't think it's a must have. When it comes to AI, show what you built or automated tasks using AI to be able to full… Whenever I get this question nowadays from PMs, I recommend getting a course in storytelling instead. That's where PMs can keep their edge. Yes, I have sent every new Product Manager hire through the 2-day **Certified Scrum Product Owner** course. Not because I believe product… I have a scrum product owner cert and a pragmatic institute product manager cert. both have been pretty helpful but they don’t excuse hav… Only relevant when there is scrum in use. If I see a resume of PM candidate with product owner certifications that’s actually a red flag for me I got PSM, PSPO, Safe Agilist and PMP. Out of them PMP covers much more valuable things around stakeholder, team and sponsor interactions…
Lenny Rachitsky

Lenny argues that Google has an unusually valuable data asset, while AI assistants already use Google data to deliver useful capabilities and Google is “MIA” in the space—an opportunity for Google to build AI products around that asset. He says the intent is to inspire Google to act, not to take a shot at it.

Is there a more valuable dataset than your Google data? Wild that every AI assistant slurps it up and does magical things with it while G… Not saying this throw shade at Google, but to inspire them to action
Aakash Gupta
  • PM agents can automate support and interview feedback digests, sales-deal intelligence, competitor updates, metric anomalies, onboarding friction, and customer outreach; a separate maker and checker can rank backlog against strategy and rerun rankings that fail the check.
  • For discovery, agents can build divergent prototypes from a team’s component library for a synthetic customer to review, while roadmap bets and stakeholder reads stay with the PM. Persistent loops need a trigger, skill file, maker, checker, gate, and state file; write corrections into the skill and retire loops whose output is no longer edited.
  • Aakash’s proposed product-team operating model uses skills for recurring work, parallel subagents (including a skeptic), direct access to Zendesk, Gong, Amplitude, and Linear, and git-versioned memory. It classifies customer signals as observations, interpretations, or hypotheses and flags conflicts with past bets; the PM decides what to promote before the system generates a PRD, review panel, prototype, evals, and draft code, with another human gate before merge.
  • Skills encode behavior that can be tested and graded, unlike prompts that produce outputs without a test; this makes it possible to check whether changes worsen results and preserve work beyond its creator. Aakash says he uses skills 20+ times a day and runs 25 against 75 tests.
  • Aakash cites hiring data in which 60% of recent AI PM hires lacked a technical background, among 12,000+ people who moved into the role in under two years. His learning roadmap covers AI fundamentals and context engineering, prototyping before writing the PRD, agents (while noting a good UI can win for simple tasks), manually labeling outputs before using an LLM judge for evals, observability and cost, and AI-specific product sense; he cautions that the best users can also be the most expensive.
PM is now the slowest part of an AI-native team. Here's how to run it with agents instead: I've spent months building and testing PM loop… This is what Zapier's CEO rated Transformative. I showed Wade Foster an operating model for a whole product team: Skills hold the team's … A prompt is text. A skill is behavior. That distinction sounds academic until you try to hand your work to someone else. I have used skil… You don't need a CS degree to become an AI PM. 60% of recent AI PM hires didn't have a technical background. (That's from the hiring data…
Hiten Shah

Hiten Shah open-sourced 16 AI competitive-intelligence skills covering market briefings, positioning, pricing, launches, battlecards, deal prep, and win/loss; each teaches AI a different method. He says he will later show a product built around them that connects those methods with market knowledge, but gives no further product details here.

I just open-sourced 16 competitive intelligence skills. They cover market briefings, positioning, pricing, launches, battlecards, deal pr…
Product Management
  • For a complex, fast-changing product where dense docs go unread and videos become outdated, pair release-time live demos or walkthroughs with recordings and documentation; recruit a strong learner to organize broader team training, then hold a follow-up after staff have tried the changes.
  • For larger features, involve a Support or Implementation point person in customer field trials or pilots and have them shadow the training as a train-the-trainer approach.
  • As a proposed supplement, use detailed product and troubleshooting material to generate a support diagnostic guide, then regenerate it as the product or issue-resolution guidance changes; flowcharts and visual references may help.
Best practices regarding customer support training Do a live demo before every release, not just a recording. Ask their manager for help in finding the strongest learner in the group, and … Some folks (the minority, IME) are fine to learn on their own by reading through documentation. I’m a big fan of live walkthroughs myself… These days I would take a modern paid AI and load your linear detailed documentation and any existing troubleshooting documentation and h…
Lenny Rachitsky

Ramp CPO Geoff identifies three long-term paths for product managers: technical PM, tastemaker PM, and GM.

CPO of Ramp [@geoffintech](https://x.com/geoffintech): I see 3 paths for product managers long-term: 1. The technical PM 2. The tastemake…
Product Management - The place for all things product

A crosspost asks whether AI-driven speed gains for one engineer make the whole team faster; the supplied post does not provide an answer or evidence.

When one engineer gets much faster with AI, does the team actually get faster?
Lenny Rachitsky

Rob Stein argues that product management’s core craft is decision-making: PMs should study what enables great decisions, beyond simply being broadly competent across many tasks.

One of my favorite lessons from [@rmstein](https://x.com/rmstein)'s talk: "People say the PM's main craft is that you do lots of things p…
Product Management
  • For PMs whose strategy work is repeatedly interrupted by release reviews, QA questions, design feedback, and stakeholder requests, batch routine reviews into fixed windows or office hours and protect focus blocks; set an explicit urgent-bug/P1 or release-blocker exception so delivery work is not left waiting.
  • Add a “what happens if” section to specs for cases such as empty states, lost connectivity, invalid input, and permissions; one PM reported that QA pings mostly stopped after two sprints, and described batching build reviews at 11 and 4 except when a release was blocked that day.
  • Create an internal page with answers to common requests and links to relevant PDFs to reduce repeated “call the PM” requests.
How do you stop yourself getting distracted by colleagues' requests? I run a daily 9am drop-in session for devs to get feedback and ask questions. Protects the rest of my day from this particular distractio… Part of this is the job, but someone sending you a DM is not always “drop everything and answer me” - you’re the one treating it that way… Create a JIRA form with 25 mandatory fields, demanding market validation and sources for each request, otherwise it is not considered. Gu… What you need to do is funnel these requests. Create an office hour or dedicated design review time. Be very vocal that this time is for … The QA ping is the one I'd attack first. Every time QA asks me what the expected behavior is, it means the spec missed a case, so I start… It's amazing how few people will spend even a moment to find information! "Just call the product manager" is not a good practice. These i…
Product Management

In one large-org model, initiatives span multiple squads and are led by a PM/UX/engineering trio while PMs retain squad responsibilities; one squad PM reports that an initiative lead leaves teams to work out details and route decisions upward, creating problems on a complex effort, and asks who should own coordination, defining the work, and trade-offs. A respondent recommends the initiative PM act as the key coordinator, with squad PMs raising specific, actionable gaps such as missing requirements or unresolved finance input; clarifying squad autonomy, concurrent work, and outcome ownership can help set the boundary. Another commenter suggests using a TPM when more than one other team is involved to coordinate teams, collate updates, and track progress.

Running an initiative I worked in similar structure before, so I can related a little bit. Based on the background you provided in the first half, it looks lik… Is there a program team in your org? Usually if there’s more than 1 other team involved, I’ll get a TPM to support bringing teams togethe…