We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
Big Ideas
Claire Vo argues that AI has shifted the PM bottleneck from deciding what can be built to having conviction about what is worth building: her own shipping capacity grew faster than her ability to find commercially meaningful products. She warns that clearing backlogs, shipping competitor parity, or abandoning ideas on noisy signals can create motion without meaningful progress on customer problems or business goals. Her alternative to a feature-and-date roadmap: set a durable conviction, define evidence that would confirm or disprove it, test quickly with real customers, and label each release a probe, experiment, or promise so commitment matches evidence.
Practitioner reports suggest this shift is already visible in daily work: a senior PM at a European fintech says AI handles product and code questions, documents, benchmark updates, customer-feedback summaries, analytics, and ticket creation; another PM identifies discovery as the bottleneck before engineers run out of refined tickets. Dan Shipper recommends separating frontier exploration from product delivery with a lab of one or two people. He suggests expecting to discard about 90% of its experiments and moving the promising work through internal use, early customers, and product review; gates include repeat usage, a sustained 10× improvement, and affordability at scale.
Tactical Playbook
The Beautiful Mess calls the judgment built into old handoffs “positive friction”: rereading, reshaping, and challenging customer evidence forced attention, while AI can turn a few calls into a flood of derivative opportunities and tasks. Keep original feedback atomic and linked to its source; distinguish signals from interpretations, hypotheses, choices, and actions; at each transformation record what carried forward, what changed, and why. Label AI’s contribution, avoid summary-on-summary chains, and focus changes enough that customer feedback can tell you what worked.
For a consequential decision, imagine it has failed six months later and ask what went wrong. One practitioner says the exercise surfaced worries people might otherwise hold back and increased confidence to make bigger bets. Make the prompt concrete: a commenter found that a real measure—30 new escalations a week with two people and no extra hours—elicited more specific risks than an abstract failure story; also ask which failure would be hardest to reverse.
Case Studies & Lessons
A Run the Business retrospective describes a vertical B2B release that drew negligible uptake despite strong marketing. It assisted experts in a critical workflow but did not fit how they worked in the field; its output was not meaningfully better than a human-in-the-loop, and the productivity gain was too small to drive adoption. The less obvious failure was “passive dissent”: some embedded experts thought about the risks but did not tell product or UX directly, partly because they did not want to derail the release or were unsure their input was welcome. The author’s remedies are to explain the reasoning behind a solution, keep dialogue with field experts ongoing, explicitly seek dissent and show a willingness to change course—and not ship the wrong call just to meet a date.
Career Corner
For senior PM roles, Shreyas Doshi’s stage-fit heuristic pairs Explore with a visionary or craftsperson, Expand with a craftsperson and later an operator, and Extract with an operator; he says he did not pursue Stripe’s head-of-product role when the company needed an operator rather than his preferred craftsperson archetype. In interviews, use cognitive empathy to target the evidence each person needs: boards look for impact, execution, and complementary skills; cross-functional peers want understanding of their challenges and alignment; prospective reports want clarity, empowerment, and growth. Even an awkward question can be steered toward the underlying concern without dismissing it.
- For consumer AI, Zhuo proposes four tests: the “mom test” (fit familiar mental models), the “spare key test” (request sensitive access when the user has context and value in return), the “putting-it-off test” (solve disliked, deferred tasks rather than merely speeding up tasks that are already easy), and the “bored-in-line test” (offer low-effort engagement without dead ends).
- Muse illustrates approachability with a friendly agent, human-like responses, short conversational language, and one default main thread; for a website signup, it asked whether Zhuo wanted to enter an email verification code herself or connect Gmail so it could do so. A high-value use case was finding a rare childhood cassette across language and marketplace barriers: Muse located a listing and a shipping workaround, though a captcha still required user intervention. Zhuo cautions that week-one retention is too early to trust, despite early-adopter stickiness.
- Muse’s Ideas suggestions help address the blank-slate problem, but Zhuo found its separate content feed less relevant than existing social feeds and some recommendations irrelevant. She recommends making suggestions more prominent for new users and organizing parallel tasks as “active work” with separate follow-up threads; a useful test is whether someone returning after a day can resume the right task without rereading everything.
- Use Kent Beck’s 3X framework to match product leadership to the product’s stage—Explore, Expand, or Extract—because goals, prioritization, and optimization should change by stage rather than copy another company’s playbook. Explore typically calls for a visionary or craftsperson; Expand calls for a craftsperson, then an operator in its later phase; Extract generally calls for an operator, though a major turnaround may call for a visionary. Doshi says he passed on Stripe’s head-of-product role when the company needed an operator rather than his own craftsperson archetype, and an operator was hired.
- In startup and mid-size company interviews, a candidate with genuine understanding of senior product leadership can help a founder or executive clarify what the role requires rather than only pitching themselves. For senior-level conversations, ask what product areas the interviewer will continue to own, what single “superpower” they want in the hire, and what one or two things the person must excel at; the ownership question can expose potential role conflict and signal that company needs come first. These questions are aimed at founders and senior executives, not a PM3 interviewing with a GPM, though they can apply when a VP is hiring for their team.
- If circumstances allow, practice interviewing before approaching top-choice companies: candidates are often rusty in their first two or three interviews. Treat this as a guideline, not a rule, and use judgment.
- In a review based on a week and a half of use, Julie Zhuo said Muse was the first consumer agent she had tried that made her think people might change their habits; it reached #1 on the App Store, but she cautioned that week-one retention was too early to trust and reflected early adopters.
- Zhuo’s consumer-agent product tests are to match familiar mental models (her “mom test”), earn sensitive access only when users have opted in and see a concrete benefit, solve tasks people dread or defer rather than merely speeding up already-easy tasks, and offer relevant next actions instead of a blank slate. Muse illustrated progressive permission requests by asking about Gmail access only when email verification was needed, and presenting purchase details before the user connected payment. Its agent also found a rare item on a Chinese auction site and a shipping workaround, though captchas still required user intervention.
- For discovery, Zhuo found Muse’s generated story feed less relevant than existing social feeds, while its personalized “Ideas” suggestions helped users see what to ask an agent and offered useful next steps. She proposed bringing Ideas closer to the default experience for newcomers, and replacing hard-to-track side chats with an “active work” view that lets users branch into task-specific conversations; a useful test is whether someone can return after a day and resume the right task without rereading everything.
- For senior PM leadership interviews, use cognitive empathy: infer what each interviewer actually needs, then steer even weak or awkward questions toward evidence that answers those needs, without showing disdain for the question. Cross-functional interviewers may not know how to assess product leaders directly, so standard interview scripts alone may not distinguish a candidate.
- Tailor evidence to the audience: boards and VCs look for a record of impact, execution ability, and skills complementary to the founder or hiring manager; cross-functional peers look for understanding of their function’s challenges, alignment rather than total agreement, and evidence of impact; prospective reports value clarity, empowerment, and growth; hiring managers look for product insight, rapport, and complementary strengths.
- For a high-stakes interview, reduce unfamiliarity by watching the interviewer’s videos or listening to their podcasts beforehand; the speaker says this can ease nerves and make rapport more natural. Around three-quarters through the process, ask the hiring manager professionally what concerns remain about your candidacy, conveying confidence in your fit and care about the company’s success. Keep former managers informed early in a job search, without contacting them excessively, to get advice and build stronger references.
- Treat take-home assignments pragmatically: the speaker says exercises described as taking 3–4 hours may require 10–40 hours to do fully, and that candidates may be judged comparatively on presentation as well as substance; he recommends investing the time needed for top target companies and trying to improve the process after joining. He cites an earlier Stripe exercise as a practical example: analyze data, then email a customer with experiment results and a recommendation.
- At Angel City, Jess Smith’s team used calling-based research to estimate that no more than 10% of fans crossed over between co-located men’s and women’s teams; they chose to build for the other 90% and invite the crossover audience, while defining the product’s identity from day one for an underserved audience.
- The Valkyries launched their brand a year before taking the court; despite having no players yet, merchandise sold within 60 days in all 50 U.S. states and 70 countries. Smith also held off on partnership deals until the team had defined what a Valkyries partnership meant and was ready to go to market, warning that launching before knowing what to ask for could do more harm than good.
- For early-stage product teams, Smith’s staffing lesson was to hire in anticipation of revenue rather than wait for revenue first, because delayed hiring can hurt the product; she wished she had moved faster on resources in year one. The Valkyries kept revenue and fans as their two main annual OKR priorities.
- The Valkyries balanced pricing against access: $15 tickets were available when bought early, and the team set aside and donated tickets for every game. To serve fans beyond the venue—given that successful sports brands may see only about 1% of fans attend a game—they created a 40-bar Northern California network committed to showing every game.
- Market and beachhead selection: Compare competitors’ features, free plans, reviews, customer segments, and distribution; use a Value Curve scoring factors important to the host from 0–5, and compete on being distinctive rather than best at everything. When the ideal customer is uncertain, choose a small beachhead and score 5–7 candidate segments on burning pain, willingness to pay, winnable share, and referral potential, showing evidence and marking unverified scores. AskOne’s author chose product people who run live sessions, skipped interviews to learn from first users, and acknowledged that the decision might need reversing.
- Positioning and business-model example: AskOne identified a moderation gap: Slido’s free plan had no moderation, while Mentimeter offered Q&A moderation only on Pro ($24.99); AI moderation cost about $0.00002 per question, making 1,000 questions per day roughly $0.02 and enabling moderation in free rooms up to that daily limit per organization. The paid plan was priced at $20 monthly or $10 monthly billed annually for larger rooms, more rooms, and custom branding.
- Focus and distribution trade-offs: Of roughly 20 candidate features, AskOne built those supporting its revised value proposition plus a demo room, deferring the rest until it had first users; its additions included surveys, free AI moderation, Zoom and Meet apps, and an agent API/MCP server. The team treated the Zoom/Meet apps and API/MCP server as distribution channels; their listings were still pending, and AskOne intentionally remained weaker on presentation design.
- Agent-assisted delivery and validation: The “artifact > plan > go” workflow is: have an agent research and present an interactive artifact for PM challenge, let it inspect the codebase and write a plan, then have it build, test, review, and capture screenshots while the PM verifies key scenarios before release. For moderation, the author tested 100 questions across dimensions and adjusted the prompt; in that test, rude and spam questions were rejected or sent to a human moderator, while automated evaluations and human-model agreement measures were deferred.
- Learn through early use: AskOne used analytics to understand use, drop-off, and behavior, and let prospective users try a demo room without signing up. Its September 24, 2026 baseline was 20 host users, 17 organizations, zero paying organizations, and 33 rooms opened.
- The review reports Muse reached #1 on the App Store, but cautions that week-one retention is too early to trust; chart position is an early interest signal, not proof of durable use.
- For consumer agents, make the experience approachable by matching familiar mental models: Muse builds on chat, uses a friendly personalized agent and short, text-like language, and defaults to one main conversation. Ask for sensitive access only when the user can see the value: Muse offered Gmail connection as an optional shortcut for an email verification code, and requested payment after presenting the product, amount, and receipt.
- Agent value is strongest when it removes an aversive task, not merely makes an already-easy task faster. In the reviewer's example, Muse searched Chinese auction sites for obscure Mandarin-dubbed childhood tapes, found a seller and a third-party shipping workaround, but still needed the user to handle a CAPTCHA.
- Reduce blank-slate friction with proactive, contextual suggestions: the reviewer found Muse's Ideas useful, but its generated-content feed less relevant than existing social feeds. The reviewer recommends making Ideas closer to the default, starting with clear benefits such as saving money, and keeping any social-app data import opt-in. For users juggling tasks, organize ongoing work into distinct active threads; test whether someone returning after a day can resume the right task without rereading everything.
- A release for expert users in a vertical B2B product drew negligible uptake despite strong marketing because it did not fit users’ real-world workflow, its output was not meaningfully better than a human-in-the-loop, and the productivity gain was too small to prompt adoption. Although cross-functional partners initially felt product had ignored known warnings, retrospective conversations showed some embedded experts had not shared feedback directly; the author linked the breakdown to a wasted learning cycle and squandered customer goodwill.
- To reduce repeat failures, product teams should explain the reasoning behind a solution, maintain ongoing dialogue with field and domain experts, actively invite dissent and demonstrate a willingness to change course, and prioritize the right product decision over ship-date commitments. Experts should engage directly, ground their perspective in long-term customer outcomes, and bring their disagreement and proposed alternative to the senior decision-maker with a stake in the outcome.
- Zapier’s AI-fluency rubric v2 assesses mindset, strategy, building, and accountability rather than AI usage alone; its stated principle is that people may delegate work to AI but remain accountable for it.
- In Wade Foster’s PM examples, a capable submission was an AI-assisted PRD, but he expected a prototype and clear customer evidence; an adoptive submission paired a working prototype with an evidence-backed PRD.
- A transformative product-team model connected agents to customer and product systems, used parallel reviews (including a skeptic) and versioned memory, and turned customer signals into classified observations, interpretations, or hypotheses with contradictions flagged. The PM decided what to promote; an automated chain then produced a PRD, review, prototype, evaluations, and draft code, with a human gate before merge.
- For day-to-day practice, the article recommends disclosing how much AI-generated work has been reviewed rather than passing unchecked output off as finished. Foster says a transformative rating should be rare—about one PM on a team of 40—and constant system-tweaking may mean the PM is not shipping; the suggested progression is to ship a first pull request, automate repetitive tasks, add and version system memory, then build team- and company-level operating systems.
- When frontier exploration competes with roadmap execution, separate the work: a lab explores new capabilities while the product team scales proven work and serves existing customers. A lab can be as small as one person; expect it to discard about 90% of its experiments, with the product team adopting only a small share (the speaker estimates about 10%).
- Keep the lab small—one or two people—and pair rapid, messy experimentation (“pirates”) with people who can turn promising prototypes into useful, extensible systems (“architects”). Dogfood experiments for a tight feedback loop; if that is not possible, test with a few early-adopter customers. Run competing approaches in parallel to explore an uncertain frontier.
- Move experiments through a research pipeline: lab trials, internal use, early-customer testing and product-team evaluation, then scaling and release. Review the pipeline regularly and gate progression on repeat usage, whether the solution remains substantially better over time (the speaker’s bar is 10× better), and whether it is affordable to serve at scale.
- At Every, an AI copy-editing agent was built using editor Kate’s edits from the prior three years. After internal use began, a dashboard tracked accepted suggestions and remaining work; it showed Kate doing 12% less work on those types of edits than the month before, prompting consideration of early-customer testing.
- AI can remove workflow handoffs that used to force PM judgment. The author warns that summary-on-summary pipelines blur signals, insights, intent, and actions, encourage “ship, ship, ship,” and can create more change than customers can absorb or give feedback on.
- For AI-assisted product work, keep original feedback atomic and link each derived issue, observation, theme, or action to its source. Represent signals, interpretations, hypotheses, options, choices, intent, actions, changes, and impact as distinct but related items; at each transformation, record what was carried forward, what changed, and why. Label AI’s contribution and avoid repeatedly summarizing summaries without a path back to the source.
- Use Linear issues to state bounded team intent—whose need it represents, why it matters, the problem, and enough of an approach to make work verifiable—not as a transcript repository; keep personal work breakdowns separate when they add team noise. Treat code analysis as one evidence source, not proof of customer experience or problem fit, and favor focused improvements over breadth so customers can give clearer feedback and teams can better identify what caused an effect.
Shreyas Doshi posted a new video about how senior PM leadership interviews work; the post provides no interview tactics or framework details.
A useful product-testing principle: “Design like you’re right, test like you’re wrong”—approach a design with conviction, but test it as an unproven hypothesis.
Meta says Muse can shop across Shopify and is adding connectors for other retailers; at checkout it is supposed to show the exact purchase, while its built-in wallet can use a single-use card limited to one merchant and amount for a short period, without exposing the underlying card number . Amazon has blocked Muse, showing that user approval alone does not guarantee access; Shah says users should know the permissions before letting an agent spend money or keep working across devices . For agentic-commerce PMs, the post highlights the need to design for both user-side authorization and merchant-side access, and to make the agent’s scope clear upfront .
A user criticized Uber’s map for repeatedly recentering on their location instead of preserving a zoomed-out view that helps them judge how far away their car is, and asked Uber to stop the behavior.
- As AI makes more ideas buildable, the product bottleneck shifts from engineering capacity to deciding what is genuinely worth building; feature-and-date roadmaps and effort-based prioritization can accelerate backlog-clearing, competitor parity, and shipping churn without producing meaningful customer or business progress.
- Replace feature lists with a conviction-driven process: set a longer-term direction, define in advance what evidence would confirm or disprove it, use rapid builds to test against real customers and data, then allocate effort based on what is learned. Keep convictions durable but features disposable—stay with the problem, revise the solution, and avoid shipping repeatedly without learning.
- Make the commitment level explicit for each feature: a probe, an experiment the team will pursue against a conviction, or a promise customers can rely on. Shape the roadmap around ambition, evidence of progress, what would prove the team wrong, and what merits more investment; consider tracking the number of ambitious experiments rather than raw PR volume or revenue per headcount.
Hiten Shah argues that institutional problems can arise when organizations lose sight of the customer; he illustrates this with education leaders focusing on adults rather than children and politicians talking to other politicians rather than voters. For product managers, the takeaway is to keep product decisions anchored in the people the product is meant to serve, rather than only in internal stakeholders.
- AI fluency is becoming a hiring and performance-review signal: Aakash reports that hundreds of companies have used Zapier’s four-level rubric—Unacceptable, Capable, Adoptive, Transformative—to screen candidates or assess teams. Zapier’s v2 evaluates mindset, strategy, building, and accountability rather than usage alone; the article contrasts this with Duolingo dropping AI use as a formal review metric. For PM work, a capable artifact is a solid AI-assisted PRD; an adoptive one adds a working prototype and customer evidence; a transformative system clusters and classifies customer signals, flags conflicts with prior bets, leaves promotion decisions to the PM, and automates PRD, review, prototype, evaluation, and draft-code production behind human gates. AI can do delegated work, but the PM retains accountability: disclose whether AI output was only skimmed or fully reviewed, rather than passing unfiltered work off as finished. The article describes Transformative as a rare rating and warns that pursuing it constantly can mean tinkering instead of shipping; its suggested progression is to ship a first pull request, automate repetitive work, add system memory, version PM tools, then build team- and company-wide operating systems.
- AI is bringing PMs closer to the three core parts of the job: customer understanding, technical possibility, and business value. Direct connections to customer-data tools can reduce dependence on intermediaries and free time for customer conversations; prototypes on real code and design systems can replace early wireframes; and querying the data warehouse lets PMs build revenue-by-segment cases without waiting for an analyst. The PM still decides what ships, supported by agentic tools.
- To reduce the risk of building for stakeholders rather than users, use continuous discovery to move from opportunity to problem to solution, test prototypes with users before engineering investment, and keep stakeholder discussions anchored in business goals and user problems rather than design choices. The source highlights this approach particularly for large engineering investments.
- Product leaders can replace micromanagement with systems, coaching, and context: improve processes or cross-team goal alignment to address root causes, position people to use their strengths and develop skills, and clarify what matters and why the work is meaningful.
- Early-career PM applicants: Google’s US 2027 new-grad and intern APM application window closes October 6; Meta’s RPM is an 18-month program with a bootcamp and three rotations, starting in March or September 2027. The post says Meta RPM takes no referrals, and advises candidates to prioritize their application; for Google, build a working app with real users and usage evidence to demonstrate applied AI/ML, and ask the recruiter which interview rounds to expect before preparing.
- A senior PM at a European fintech reports using AI for product and code questions, product documents and research protocols, benchmark updates, NPS and customer-service feedback analysis, analytics queries, Slack-to-Linear tickets, and security/legal requests; developers mainly review generated code. Strategy, idea generation, and user interviews remain, but the PM feels stripped of parts of the role. A separate practitioner says they inspect locally imported code with Cursor and an LLM, refresh local docs periodically, and use it to find undocumented capabilities such as bulk actions.
- One PM from an AI-first company says faster build-test-ship-learn cycles made choosing the right product more critical, and their team urgently hired senior PMs. Another PM describes discovery as the bottleneck when engineers may run out of refined tickets before discovery is complete.
- Suggested adoption tactics: build hands-on, choose three real problems to solve, and assess the cost of improving AI tools, skills, or process rather than treating adoption as a binary AI/no-AI choice. To manage the pace, another practitioner advises focusing on the most critical work and automating the rest to limit burnout.
After Replit Agent deleted database data during development, Replit separated development and production databases so the Agent cannot change production data during development. Hiten Shah highlights this product-level permission boundary because it avoids relying on the model to remember a rule.