We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
Big Ideas
Claire Vo argues that AI has shifted the PM bottleneck from deciding what can be built to having conviction about what is worth building: her own shipping capacity grew faster than her ability to find commercially meaningful products. She warns that clearing backlogs, shipping competitor parity, or abandoning ideas on noisy signals can create motion without meaningful progress on customer problems or business goals. Her alternative to a feature-and-date roadmap: set a durable conviction, define evidence that would confirm or disprove it, test quickly with real customers, and label each release a probe, experiment, or promise so commitment matches evidence.
Practitioner reports suggest this shift is already visible in daily work: a senior PM at a European fintech says AI handles product and code questions, documents, benchmark updates, customer-feedback summaries, analytics, and ticket creation; another PM identifies discovery as the bottleneck before engineers run out of refined tickets. Dan Shipper recommends separating frontier exploration from product delivery with a lab of one or two people. He suggests expecting to discard about 90% of its experiments and moving the promising work through internal use, early customers, and product review; gates include repeat usage, a sustained 10× improvement, and affordability at scale.
Tactical Playbook
The Beautiful Mess calls the judgment built into old handoffs “positive friction”: rereading, reshaping, and challenging customer evidence forced attention, while AI can turn a few calls into a flood of derivative opportunities and tasks. Keep original feedback atomic and linked to its source; distinguish signals from interpretations, hypotheses, choices, and actions; at each transformation record what carried forward, what changed, and why. Label AI’s contribution, avoid summary-on-summary chains, and focus changes enough that customer feedback can tell you what worked.
For a consequential decision, imagine it has failed six months later and ask what went wrong. One practitioner says the exercise surfaced worries people might otherwise hold back and increased confidence to make bigger bets. Make the prompt concrete: a commenter found that a real measure—30 new escalations a week with two people and no extra hours—elicited more specific risks than an abstract failure story; also ask which failure would be hardest to reverse.
Case Studies & Lessons
A Run the Business retrospective describes a vertical B2B release that drew negligible uptake despite strong marketing. It assisted experts in a critical workflow but did not fit how they worked in the field; its output was not meaningfully better than a human-in-the-loop, and the productivity gain was too small to drive adoption. The less obvious failure was “passive dissent”: some embedded experts thought about the risks but did not tell product or UX directly, partly because they did not want to derail the release or were unsure their input was welcome. The author’s remedies are to explain the reasoning behind a solution, keep dialogue with field experts ongoing, explicitly seek dissent and show a willingness to change course—and not ship the wrong call just to meet a date.
Career Corner
For senior PM roles, Shreyas Doshi’s stage-fit heuristic pairs Explore with a visionary or craftsperson, Expand with a craftsperson and later an operator, and Extract with an operator; he says he did not pursue Stripe’s head-of-product role when the company needed an operator rather than his preferred craftsperson archetype. In interviews, use cognitive empathy to target the evidence each person needs: boards look for impact, execution, and complementary skills; cross-functional peers want understanding of their challenges and alignment; prospective reports want clarity, empowerment, and growth. Even an awkward question can be steered toward the underlying concern without dismissing it.
- For consumer AI, Zhuo proposes four tests: the “mom test” (fit familiar mental models), the “spare key test” (request sensitive access when the user has context and value in return), the “putting-it-off test” (solve disliked, deferred tasks rather than merely speeding up tasks that are already easy), and the “bored-in-line test” (offer low-effort engagement without dead ends).
- Muse illustrates approachability with a friendly agent, human-like responses, short conversational language, and one default main thread; for a website signup, it asked whether Zhuo wanted to enter an email verification code herself or connect Gmail so it could do so. A high-value use case was finding a rare childhood cassette across language and marketplace barriers: Muse located a listing and a shipping workaround, though a captcha still required user intervention. Zhuo cautions that week-one retention is too early to trust, despite early-adopter stickiness.
- Muse’s Ideas suggestions help address the blank-slate problem, but Zhuo found its separate content feed less relevant than existing social feeds and some recommendations irrelevant. She recommends making suggestions more prominent for new users and organizing parallel tasks as “active work” with separate follow-up threads; a useful test is whether someone returning after a day can resume the right task without rereading everything.
- Use Kent Beck’s 3X framework to match product leadership to the product’s stage—Explore, Expand, or Extract—because goals, prioritization, and optimization should change by stage rather than copy another company’s playbook. Explore typically calls for a visionary or craftsperson; Expand calls for a craftsperson, then an operator in its later phase; Extract generally calls for an operator, though a major turnaround may call for a visionary. Doshi says he passed on Stripe’s head-of-product role when the company needed an operator rather than his own craftsperson archetype, and an operator was hired.
- In startup and mid-size company interviews, a candidate with genuine understanding of senior product leadership can help a founder or executive clarify what the role requires rather than only pitching themselves. For senior-level conversations, ask what product areas the interviewer will continue to own, what single “superpower” they want in the hire, and what one or two things the person must excel at; the ownership question can expose potential role conflict and signal that company needs come first. These questions are aimed at founders and senior executives, not a PM3 interviewing with a GPM, though they can apply when a VP is hiring for their team.
- If circumstances allow, practice interviewing before approaching top-choice companies: candidates are often rusty in their first two or three interviews. Treat this as a guideline, not a rule, and use judgment.
- In a review based on a week and a half of use, Julie Zhuo said Muse was the first consumer agent she had tried that made her think people might change their habits; it reached #1 on the App Store, but she cautioned that week-one retention was too early to trust and reflected early adopters.
- Zhuo’s consumer-agent product tests are to match familiar mental models (her “mom test”), earn sensitive access only when users have opted in and see a concrete benefit, solve tasks people dread or defer rather than merely speeding up already-easy tasks, and offer relevant next actions instead of a blank slate. Muse illustrated progressive permission requests by asking about Gmail access only when email verification was needed, and presenting purchase details before the user connected payment. Its agent also found a rare item on a Chinese auction site and a shipping workaround, though captchas still required user intervention.
- For discovery, Zhuo found Muse’s generated story feed less relevant than existing social feeds, while its personalized “Ideas” suggestions helped users see what to ask an agent and offered useful next steps. She proposed bringing Ideas closer to the default experience for newcomers, and replacing hard-to-track side chats with an “active work” view that lets users branch into task-specific conversations; a useful test is whether someone can return after a day and resume the right task without rereading everything.
- For senior PM leadership interviews, use cognitive empathy: infer what each interviewer actually needs, then steer even weak or awkward questions toward evidence that answers those needs, without showing disdain for the question. Cross-functional interviewers may not know how to assess product leaders directly, so standard interview scripts alone may not distinguish a candidate.
- Tailor evidence to the audience: boards and VCs look for a record of impact, execution ability, and skills complementary to the founder or hiring manager; cross-functional peers look for understanding of their function’s challenges, alignment rather than total agreement, and evidence of impact; prospective reports value clarity, empowerment, and growth; hiring managers look for product insight, rapport, and complementary strengths.
- For a high-stakes interview, reduce unfamiliarity by watching the interviewer’s videos or listening to their podcasts beforehand; the speaker says this can ease nerves and make rapport more natural. Around three-quarters through the process, ask the hiring manager professionally what concerns remain about your candidacy, conveying confidence in your fit and care about the company’s success. Keep former managers informed early in a job search, without contacting them excessively, to get advice and build stronger references.
- Treat take-home assignments pragmatically: the speaker says exercises described as taking 3–4 hours may require 10–40 hours to do fully, and that candidates may be judged comparatively on presentation as well as substance; he recommends investing the time needed for top target companies and trying to improve the process after joining. He cites an earlier Stripe exercise as a practical example: analyze data, then email a customer with experiment results and a recommendation.
- At Angel City, Jess Smith’s team used calling-based research to estimate that no more than 10% of fans crossed over between co-located men’s and women’s teams; they chose to build for the other 90% and invite the crossover audience, while defining the product’s identity from day one for an underserved audience.
- The Valkyries launched their brand a year before taking the court; despite having no players yet, merchandise sold within 60 days in all 50 U.S. states and 70 countries. Smith also held off on partnership deals until the team had defined what a Valkyries partnership meant and was ready to go to market, warning that launching before knowing what to ask for could do more harm than good.
- For early-stage product teams, Smith’s staffing lesson was to hire in anticipation of revenue rather than wait for revenue first, because delayed hiring can hurt the product; she wished she had moved faster on resources in year one. The Valkyries kept revenue and fans as their two main annual OKR priorities.
- The Valkyries balanced pricing against access: $15 tickets were available when bought early, and the team set aside and donated tickets for every game. To serve fans beyond the venue—given that successful sports brands may see only about 1% of fans attend a game—they created a 40-bar Northern California network committed to showing every game.
- Market and beachhead selection: Compare competitors’ features, free plans, reviews, customer segments, and distribution; use a Value Curve scoring factors important to the host from 0–5, and compete on being distinctive rather than best at everything. When the ideal customer is uncertain, choose a small beachhead and score 5–7 candidate segments on burning pain, willingness to pay, winnable share, and referral potential, showing evidence and marking unverified scores. AskOne’s author chose product people who run live sessions, skipped interviews to learn from first users, and acknowledged that the decision might need reversing.
- Positioning and business-model example: AskOne identified a moderation gap: Slido’s free plan had no moderation, while Mentimeter offered Q&A moderation only on Pro ($24.99); AI moderation cost about $0.00002 per question, making 1,000 questions per day roughly $0.02 and enabling moderation in free rooms up to that daily limit per organization. The paid plan was priced at $20 monthly or $10 monthly billed annually for larger rooms, more rooms, and custom branding.
- Focus and distribution trade-offs: Of roughly 20 candidate features, AskOne built those supporting its revised value proposition plus a demo room, deferring the rest until it had first users; its additions included surveys, free AI moderation, Zoom and Meet apps, and an agent API/MCP server. The team treated the Zoom/Meet apps and API/MCP server as distribution channels; their listings were still pending, and AskOne intentionally remained weaker on presentation design.
- Agent-assisted delivery and validation: The “artifact > plan > go” workflow is: have an agent research and present an interactive artifact for PM challenge, let it inspect the codebase and write a plan, then have it build, test, review, and capture screenshots while the PM verifies key scenarios before release. For moderation, the author tested 100 questions across dimensions and adjusted the prompt; in that test, rude and spam questions were rejected or sent to a human moderator, while automated evaluations and human-model agreement measures were deferred.
- Learn through early use: AskOne used analytics to understand use, drop-off, and behavior, and let prospective users try a demo room without signing up. Its September 24, 2026 baseline was 20 host users, 17 organizations, zero paying organizations, and 33 rooms opened.
- The review reports Muse reached #1 on the App Store, but cautions that week-one retention is too early to trust; chart position is an early interest signal, not proof of durable use.
- For consumer agents, make the experience approachable by matching familiar mental models: Muse builds on chat, uses a friendly personalized agent and short, text-like language, and defaults to one main conversation. Ask for sensitive access only when the user can see the value: Muse offered Gmail connection as an optional shortcut for an email verification code, and requested payment after presenting the product, amount, and receipt.
- Agent value is strongest when it removes an aversive task, not merely makes an already-easy task faster. In the reviewer's example, Muse searched Chinese auction sites for obscure Mandarin-dubbed childhood tapes, found a seller and a third-party shipping workaround, but still needed the user to handle a CAPTCHA.
- Reduce blank-slate friction with proactive, contextual suggestions: the reviewer found Muse's Ideas useful, but its generated-content feed less relevant than existing social feeds. The reviewer recommends making Ideas closer to the default, starting with clear benefits such as saving money, and keeping any social-app data import opt-in. For users juggling tasks, organize ongoing work into distinct active threads; test whether someone returning after a day can resume the right task without rereading everything.
- A release for expert users in a vertical B2B product drew negligible uptake despite strong marketing because it did not fit users’ real-world workflow, its output was not meaningfully better than a human-in-the-loop, and the productivity gain was too small to prompt adoption. Although cross-functional partners initially felt product had ignored known warnings, retrospective conversations showed some embedded experts had not shared feedback directly; the author linked the breakdown to a wasted learning cycle and squandered customer goodwill.
- To reduce repeat failures, product teams should explain the reasoning behind a solution, maintain ongoing dialogue with field and domain experts, actively invite dissent and demonstrate a willingness to change course, and prioritize the right product decision over ship-date commitments. Experts should engage directly, ground their perspective in long-term customer outcomes, and bring their disagreement and proposed alternative to the senior decision-maker with a stake in the outcome.
- Zapier’s AI-fluency rubric v2 assesses mindset, strategy, building, and accountability rather than AI usage alone; its stated principle is that people may delegate work to AI but remain accountable for it.
- In Wade Foster’s PM examples, a capable submission was an AI-assisted PRD, but he expected a prototype and clear customer evidence; an adoptive submission paired a working prototype with an evidence-backed PRD.
- A transformative product-team model connected agents to customer and product systems, used parallel reviews (including a skeptic) and versioned memory, and turned customer signals into classified observations, interpretations, or hypotheses with contradictions flagged. The PM decided what to promote; an automated chain then produced a PRD, review, prototype, evaluations, and draft code, with a human gate before merge.
- For day-to-day practice, the article recommends disclosing how much AI-generated work has been reviewed rather than passing unchecked output off as finished. Foster says a transformative rating should be rare—about one PM on a team of 40—and constant system-tweaking may mean the PM is not shipping; the suggested progression is to ship a first pull request, automate repetitive tasks, add and version system memory, then build team- and company-level operating systems.
- When frontier exploration competes with roadmap execution, separate the work: a lab explores new capabilities while the product team scales proven work and serves existing customers. A lab can be as small as one person; expect it to discard about 90% of its experiments, with the product team adopting only a small share (the speaker estimates about 10%).
- Keep the lab small—one or two people—and pair rapid, messy experimentation (“pirates”) with people who can turn promising prototypes into useful, extensible systems (“architects”). Dogfood experiments for a tight feedback loop; if that is not possible, test with a few early-adopter customers. Run competing approaches in parallel to explore an uncertain frontier.
- Move experiments through a research pipeline: lab trials, internal use, early-customer testing and product-team evaluation, then scaling and release. Review the pipeline regularly and gate progression on repeat usage, whether the solution remains substantially better over time (the speaker’s bar is 10× better), and whether it is affordable to serve at scale.
- At Every, an AI copy-editing agent was built using editor Kate’s edits from the prior three years. After internal use began, a dashboard tracked accepted suggestions and remaining work; it showed Kate doing 12% less work on those types of edits than the month before, prompting consideration of early-customer testing.
- AI can remove workflow handoffs that used to force PM judgment. The author warns that summary-on-summary pipelines blur signals, insights, intent, and actions, encourage “ship, ship, ship,” and can create more change than customers can absorb or give feedback on.
- For AI-assisted product work, keep original feedback atomic and link each derived issue, observation, theme, or action to its source. Represent signals, interpretations, hypotheses, options, choices, intent, actions, changes, and impact as distinct but related items; at each transformation, record what was carried forward, what changed, and why. Label AI’s contribution and avoid repeatedly summarizing summaries without a path back to the source.
- Use Linear issues to state bounded team intent—whose need it represents, why it matters, the problem, and enough of an approach to make work verifiable—not as a transcript repository; keep personal work breakdowns separate when they add team noise. Treat code analysis as one evidence source, not proof of customer experience or problem fit, and favor focused improvements over breadth so customers can give clearer feedback and teams can better identify what caused an effect.
Shreyas Doshi posted a new video about how senior PM leadership interviews work; the post provides no interview tactics or framework details.
A useful product-testing principle: “Design like you’re right, test like you’re wrong”—approach a design with conviction, but test it as an unproven hypothesis.
Meta says Muse can shop across Shopify and is adding connectors for other retailers; at checkout it is supposed to show the exact purchase, while its built-in wallet can use a single-use card limited to one merchant and amount for a short period, without exposing the underlying card number . Amazon has blocked Muse, showing that user approval alone does not guarantee access; Shah says users should know the permissions before letting an agent spend money or keep working across devices . For agentic-commerce PMs, the post highlights the need to design for both user-side authorization and merchant-side access, and to make the agent’s scope clear upfront .
A user criticized Uber’s map for repeatedly recentering on their location instead of preserving a zoomed-out view that helps them judge how far away their car is, and asked Uber to stop the behavior.
- As AI makes more ideas buildable, the product bottleneck shifts from engineering capacity to deciding what is genuinely worth building; feature-and-date roadmaps and effort-based prioritization can accelerate backlog-clearing, competitor parity, and shipping churn without producing meaningful customer or business progress.
- Replace feature lists with a conviction-driven process: set a longer-term direction, define in advance what evidence would confirm or disprove it, use rapid builds to test against real customers and data, then allocate effort based on what is learned. Keep convictions durable but features disposable—stay with the problem, revise the solution, and avoid shipping repeatedly without learning.
- Make the commitment level explicit for each feature: a probe, an experiment the team will pursue against a conviction, or a promise customers can rely on. Shape the roadmap around ambition, evidence of progress, what would prove the team wrong, and what merits more investment; consider tracking the number of ambitious experiments rather than raw PR volume or revenue per headcount.
Hiten Shah argues that institutional problems can arise when organizations lose sight of the customer; he illustrates this with education leaders focusing on adults rather than children and politicians talking to other politicians rather than voters. For product managers, the takeaway is to keep product decisions anchored in the people the product is meant to serve, rather than only in internal stakeholders.
- AI fluency is becoming a hiring and performance-review signal: Aakash reports that hundreds of companies have used Zapier’s four-level rubric—Unacceptable, Capable, Adoptive, Transformative—to screen candidates or assess teams. Zapier’s v2 evaluates mindset, strategy, building, and accountability rather than usage alone; the article contrasts this with Duolingo dropping AI use as a formal review metric. For PM work, a capable artifact is a solid AI-assisted PRD; an adoptive one adds a working prototype and customer evidence; a transformative system clusters and classifies customer signals, flags conflicts with prior bets, leaves promotion decisions to the PM, and automates PRD, review, prototype, evaluation, and draft-code production behind human gates. AI can do delegated work, but the PM retains accountability: disclose whether AI output was only skimmed or fully reviewed, rather than passing unfiltered work off as finished. The article describes Transformative as a rare rating and warns that pursuing it constantly can mean tinkering instead of shipping; its suggested progression is to ship a first pull request, automate repetitive work, add system memory, version PM tools, then build team- and company-wide operating systems.
- AI is bringing PMs closer to the three core parts of the job: customer understanding, technical possibility, and business value. Direct connections to customer-data tools can reduce dependence on intermediaries and free time for customer conversations; prototypes on real code and design systems can replace early wireframes; and querying the data warehouse lets PMs build revenue-by-segment cases without waiting for an analyst. The PM still decides what ships, supported by agentic tools.
- To reduce the risk of building for stakeholders rather than users, use continuous discovery to move from opportunity to problem to solution, test prototypes with users before engineering investment, and keep stakeholder discussions anchored in business goals and user problems rather than design choices. The source highlights this approach particularly for large engineering investments.
- Product leaders can replace micromanagement with systems, coaching, and context: improve processes or cross-team goal alignment to address root causes, position people to use their strengths and develop skills, and clarify what matters and why the work is meaningful.
- Early-career PM applicants: Google’s US 2027 new-grad and intern APM application window closes October 6; Meta’s RPM is an 18-month program with a bootcamp and three rotations, starting in March or September 2027. The post says Meta RPM takes no referrals, and advises candidates to prioritize their application; for Google, build a working app with real users and usage evidence to demonstrate applied AI/ML, and ask the recruiter which interview rounds to expect before preparing.
- A senior PM at a European fintech reports using AI for product and code questions, product documents and research protocols, benchmark updates, NPS and customer-service feedback analysis, analytics queries, Slack-to-Linear tickets, and security/legal requests; developers mainly review generated code. Strategy, idea generation, and user interviews remain, but the PM feels stripped of parts of the role. A separate practitioner says they inspect locally imported code with Cursor and an LLM, refresh local docs periodically, and use it to find undocumented capabilities such as bulk actions.
- One PM from an AI-first company says faster build-test-ship-learn cycles made choosing the right product more critical, and their team urgently hired senior PMs. Another PM describes discovery as the bottleneck when engineers may run out of refined tickets before discovery is complete.
- Suggested adoption tactics: build hands-on, choose three real problems to solve, and assess the cost of improving AI tools, skills, or process rather than treating adoption as a binary AI/no-AI choice. To manage the pace, another practitioner advises focusing on the most critical work and automating the rest to limit burnout.
After Replit Agent deleted database data during development, Replit separated development and production databases so the Agent cannot change production data during development. Hiten Shah highlights this product-level permission boundary because it avoids relying on the model to remember a rule.
TBM 441: AI, the Loss of Positive Friction, and What to Do About It
Imagine a stereotypical product workflow.
Customer call → synthesize insights → decide what to do about it → do it → ship it → measure impact
Pre-AI, this would involve 1) reading/reviewing a call transcript and your notes, 2) synthesizing that into a set of takeaways either in a document or presented somehow, 3) considering a set of options by talking about those options, writing about them, and narrowing those down to some tangible actions/investments, 4) taking those actions (e.g., building something) using whatever collaborating techniques and “pointers to the work” (e.g., tickets), 5) shipping it, and 6) ideally measuring impact by going into an analytics tool, watching feedback channels, reaching out to customers, etc.
At each stage, you were, in effect, asked to process, think about, discuss, reshape, restate, copy-paste, “migrate” information down the line. There might have been meetings, routines, gates, swapping between tools. With only N hours in the week, you were constantly kicking yourself for not being as collaborative, inclusive, comprehensive, or rigorous/disciplined enough. Yes, this was a pain in the ass, but it was also “positive friction.” Each of these information migration moments was an opportunity, in HCI parlance, to apply a forcing function: to focus conscious attention on something, to avoid automating your thinking, to pay attention, discern, say “no” or “what if.”
The friction was judgement and discretion’s friend. We tend to remember the crappy parts: the meeting prep, the “I have a document with 35 feedback themes and a Miro board. Where the hell do I put this?” We remember that sense of never doing enough.
But it was also valuable.
New Reality (Unless You Push Back)
Say you wanted to just automate everything now with AI. You could effectively:
Do your expenses during the call. Just ask the occasional question.
Take that three-thousand-word transcript and ask an LLM to “summarize this into key pain points” (the average call can generate dozens if not hundreds of these)
Plop those pain points into a “skill” that looks at the current code, maybe even checks analytics, and writes up “specs”
Run the “Strategy alignment prioritization” skill (alongside your secret “what I need for my promotion” skill) and sequence the bets
Shuttle that into some kind of harness to build the thing
Have some other agents review the PR
Ship it…
Half-heartedly verify the features work, but hey, Playwrite has been doing a good job lately, so automate that
Trigger a downstream analysis skill that asks if it “worked”
And you could effectively do dozens of these calls, or just monitor customer “voice” channels and not even show up. Promotion! Bingo! What a time to be alive!
Notice all the opportunities for thinking, pruning, shaping, designing, judgement-applying that are lost. AND, noticing all the drudgery that is lost. Noticing how very quickly nothing you are looking at is “real” (original, atomic, as written by a human). You have tens of thousands of words of original feedback, yes, but very quickly all you have is pointers (if you’re lucky and/or smart). Yes, the code is real. That’s why everyone has the utmost confidence in that step. But how did the code get that way? Sure, if it is all obvious slop, that’s one problem. But what if it isn’t? That’s almost worse.
Prior to AI, you might focus on one product area, one opportunity, and methodically work through it. You did the same calls, but those calls didn’t come with hundreds of things you could/can do. The blast radius of any particular thing you were working on was narrower. Very quickly, you amassed more feedback than you could meaningfully act on, support, market, communicate to customers, and keep in your own head.
Now you can go from five customer calls (~8000 words) to 80 “opportunities” to 80 “tickets,” 400 “tasks,” and hundreds of stuff to show customers, to hundreds of bits of feedback on those things, to 50 “feature adoption dashboards” (hidden behind MCP obviously)…. Not to mention all the internal call recordings, status updates, PRs, stuff, fluff, slop, 1000 interface variations, etc. And after doing ALL of that, your customer is going to say, “I can’t keep track of everything you’re doing.”
All About Relationships
I recently did an activity where I looked at hundreds of artifacts at my current gig and looked at the “job” each one of them played. Many artifacts contain LOTS of atomic bits of job-doing information.
Signals: What is observed, recorded, written, or measured
Insights / Interpretation: What we infer from those signals
Models / Hypotheses: How we believe things work and connect, including assumptions
Options: Possible paths, configurations, or interventions
Choice: The selected option and resulting commitment
Intent: The future state or direction being pursued
Actions: What actors do to advance the intent
Mechanical Change: The specific thing that is modified
Effects / Impact: The resulting change
Actors: Who observes, interprets, chooses, acts, or is affected
Materials: Systems, tools, technologies, artifacts, and infrastructure
Constraints: Conditions that limit, shape, or enable the system
Whenever I see a simplistic “SDLC” diagram, I laugh because real knowledge work is much more than a couple stages. The ontology is much richer.

Yes, there might be a sort of physics in the flow (e.g., observe, orient, decide, act), but imagine everything above as a graph. Every edge in that graph is a transition point, even when we lump together ideas. A PRD might contain signals, insights, choices, etc., and itself serve as a model, but you are really dealing with a discrete mix of more enduring/lasting ideas and more transient ideas.
Of course, the problem was there pre-AI as well. Very few people thought in this level of detail because “it was too complicated” and “you do this naturally when you’re solving problems.” But the active participation was key.
We tend to think of the nouns. But it is the edges where the work happens. Pre-AI this happened whether we liked it or not. Now…we have excuses.
At work, I was trying to untangle an initiative recently. What I observed was as follows:
TONS of AI-generated summaries on top of summaries
Lots of surface area in an attempt to be customer-centric. We had a lot of change in the mix, probably far outpacing the customer’s ability to absorb or provide feedback on that change
AI is very lazy at telling the difference between things (e.g., signals and insights, intent and actions, etc.)
Unless we are very deliberate, we end up with Linear tickets being a dumping ground for synthesized feedback (AI) and code review (AI), and you couldn’t help but be trigger-happy and just ship, ship, ship.
Principles
Here are some rough principles I’m using to tackle this problem:
Keep original feedback atomic, and preserve its path back to the source. Transcripts and feedback are records of the original context and the start of a loop. Whenever we break that material into smaller items, such as an issue, observation, theme, or action, each item should link back to the original transcript or feedback. That traceability lets us check what was said and recover the surrounding context. If you can’t link back to something real, stop doing it.
Preserve both atomicity and coherence. Product artifacts often mash together distinct ideas, such as Intent, Options, and Change, into a single issue or document. AI needs those ideas represented as separate, atomic pieces while keeping their relationships and context intact. Atomicity without coherence leaves fragments; coherence without atomicity leaves ideas mashed together. We need both. Miss either and we’re out.
Add positive friction when content is copied and recontextualized. At each transformation, make it clear what is being carried forward, what changed, and why. The friction belongs at that moment: enough to make the transformation intentional and keep the source links intact.
Separate the underlying job from its presentation to different audiences. Different audiences may need different framing, but that doesn’t change the underlying job or intent. Keep the source clear, label each audience-specific version as a variation, and link it back to the source.
Label AI-generated content by degree of AI contribution. Make the label visible on every AI-generated artifact, with a variation such as “100% AI-generated,” “HITL curated,” or “AI used for editing/grammar.” Keep the label with the content when it is copied or recontextualized.
Avoid compounding AI flattening. The biggest anti-pattern is AI flattening layered on AI flattening, with no thread back to reality. Each round can strip away more of the original context while making the resulting abstraction look authoritative. Summaries are most useful when they get close to a specific problem, opportunity, or action, and remain linked to the source material.
Don’t let AI capability turn a tool into a dumping ground. The fact that a tool can process information with AI doesn’t mean everything should be put there and sorted out later. Give each item a clear purpose, and keep it connected to its source and context.
Treat code analysis as one source of evidence. It can be very helpful to see “what the code says,” but seeing X in the code doesn’t establish that customers experience X, that we’re solving the right problem, or why the code came to be that way. Don’t let the concreteness of code analysis create confidence about the broader product reality that the code alone can’t support.
Use Linear for bounded team intent. An issue should make clear whose need it represents, why it matters, what problem we’re addressing, and enough of how we’ll address it to make the work verifiable by people or machines. It should capture intent, not serve as a transcript repository.
Go for depth over breadth in product improvements. If everything changes at once, customers can become confused, their feedback can become scattered and nonspecific, and the data can conflate the effects of different changes. Focus improvements enough that customers can respond clearly and we can learn what made a difference. Being trigger-happy with new features puts us at risk of taking on too much breadth.
Recognize Linear has two jobs. One is team-visible product and feedback intent: the problem or opportunity we’ve decided to pursue. The other is personal planning and decomposition as work unfolds. Both help people track progress and preserve context, but personal breakdowns can create noise for the team and may fit better in personal notes or another workspace.
Hope sharing this helps you think about how you work!
As a last word, I’d leave with you the question:
Where was positive friction a good thing? And how can you embrace new technologies without losing the good things?
- AI can remove workflow handoffs that used to force PM judgment. The author warns that summary-on-summary pipelines blur signals, insights, intent, and actions, encourage “ship, ship, ship,” and can create more change than customers can absorb or give feedback on.
- For AI-assisted product work, keep original feedback atomic and link each derived issue, observation, theme, or action to its source. Represent signals, interpretations, hypotheses, options, choices, intent, actions, changes, and impact as distinct but related items; at each transformation, record what was carried forward, what changed, and why. Label AI’s contribution and avoid repeatedly summarizing summaries without a path back to the source.
- Use Linear issues to state bounded team intent—whose need it represents, why it matters, the problem, and enough of an approach to make work verifiable—not as a transcript repository; keep personal work breakdowns separate when they add team noise. Treat code analysis as one evidence source, not proof of customer experience or problem fit, and favor focused improvements over breadth so customers can give clearer feedback and teams can better identify what caused an effect.