We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
Big Ideas
The product surface is moving beyond the UI. A document-platform founder reports that about half of all documents ever created now arrive through its API/MCP, generated by users’ coding agents rather than users themselves; the shift was nearly invisible in app metrics. The team found its most engaged users prompting agents to publish, edit, and analyze, turning the app into a read-only view of work done elsewhere. It responded with an agent-native CLI featuring JSON output, stable aliases, and safe defaults. For PMs: segment key actions by initiator—human versus agent—and treat interface stability and documentation as product surfaces, not implementation details.
Tactical Playbook
Give AI prototypes a lifecycle, not production status. PMs can reach clickable demos in hours, but the handoff often leaves screenshots, a Loom, a repo, and several URLs while stakeholders want iteration and engineering needs context. A workable flow:
- Label the artifact before the demo. One PM describes it as a “design mockup” for touch and feel, asks for feedback, then links it from the PRD and relevant tickets rather than presenting it as the build.
- Bring design in before commitment. A useful sequence is prototype → design one-to-one → refine → share after sign-off. AI prototypes often look finished at first glance, so explicitly mark what is changeable and what is fixed.
- Separate M1 scope from implementation effort. One PM reports that prototypes mask system complexity and create false delivery-speed expectations; another says prototype code commonly lacks error handling, security, and scaling, while the PRD and Figma file survive handoff.
- Use real endpoints only when the foundation exists. One API-backed prototype reached production in two months rather than an expected year, but the supporting backend was already in place; the thread’s estimate was that mock data can add two months.
Case Studies & Lessons
Note2Tabs: retention needs a return reason. The guitar-transcription SaaS is attracting users and delivering value, but transcription is transactional: users get tabs and leave. Its proposed shift is from transcription as the product to transcription as the entry point for editing, playback, practice, creation, and sharing. A practical extension is to keep tabs, notes, and practice history in the product; for early CAC/LTV, use cohort retention from small paid tests instead of trusting a blended number built on little history.
Matic: simple interaction can require a deep product. Matic’s launch post describes Cues, where users point and say “clean this,” ask the robot to follow them, or send it to a mapped room. The post says the robot understands 75 languages, is used by 13,000 families, and reflects nine years and $115 million of work. The PM lesson is to optimize around the user’s goal-level interaction while validating the invisible system that makes the apparent simplicity trustworthy.
Career Corner
Don’t hire for a coaching fantasy. Shreyas Doshi says he suppressed his hiring intuition when a role-relevant flaw looked “coachable”; in the majority of cases, the flaw blocked next-level impact within months and coaching did not work. People grow on their own timeline and toward directions they are naturally attracted to. Separate a development gap from a must-have capability, test the latter directly, and do not hire expecting your coaching plan to change it.
AI-facing roles are selecting for judgment and agency. LangChain’s deployed-engineer interviews test whether candidates can improve an initially generated agent, choose features tied to retention or spend, make assumptions with sparse context, manage demo time, show AI interest, and take ownership. The team reports successful candidates from software, MBA, and consulting backgrounds. For PMs, the signal is to build evidence of customer discovery, prioritization, and shipped agent workflows—not only model fluency.
Tools & Resources
Interview-to-PRD automation remains a gap. One PM says ChatPRD can write the document but lacks customer context; BuildBetter came closest by reading calls and threading quotes into spec sections, but its templates needed tuning, leaving Dovetail plus manual writing as the actual workflow. If evaluating tools, make quote provenance and customer-specific context hard acceptance criteria, not cosmetic output quality.
a16z podcast: Joe Schmidt (author of the "Lighthouse or Land Grab" piece) and Andy, who built what the hosts call the best sales organizations at Samsara and Meraki (transcribed "Moroi"), lay out which sales playbook enterprise AI startups should run . PM-relevant takeaways:
- Framework: plot opportunities on two axes — buyer exposure/risk (risk of buying wrong, plus whether the solution is exposed to the buyer's end customers) and whether proof travels in the market. High exposure + traveling proof = Lighthouse (regulated industries, few logos, category creation, social proof is the currency); low exposure + proof doesn't travel = Land Grab (established budgets, replacing existing workflows, "math" is the currency) .
- Choosing: if buyers already budget for the category and you're replacing an existing tool, lean Land Grab; if the product is brand new, needs market education, and has no existing budget, it's Lighthouse missionary work . Practical test: are prospects willing to get on the phone and buy now? If not, you need deep forward-deployed Lighthouse arrangements .
- Examples: Harvey (legal AI) = Lighthouse — won critical high-risk law-firm logos and proof traveled, making exposed buyers feel safe . STO (AI accounts-receivable collections) = Land Grab — showed mid-market buyers the math: more effective than current software/human teams, improved working capital, saved money . Pylon (AI-native customer support) = Land Grab, starting at modest ACVs and climbing . Decagon commits publicly to support benchmarks it promises to hit . Further AI (insurance) = Lighthouse — governance-first AI plus forward-deployed teams to win the world's biggest insurers .
- Samsara case: the 2016–2019 US ELD mandate forced a category to happen overnight; as a new entrant against AT&T/Verizon, Samsara sold mid-market first — less social proof required, and short sales cycles produced fast product feedback that accelerated product innovation .
- Meraki (transcribed "Moroi") land-grab tactic: webinars that shipped a free access point so mid-market buyers could experience that cloud-managed networking was simpler than Cisco/HP — "get them to try it" .
- ACV: a deal class must only top the hurdle of healthy unit economics; then sign as many deals as possible at that ACV and inch up the ACV ladder over time .
- Trial discipline: AI POCs often turn into endless "science projects" because capabilities change daily; guard against this with a fixed end date (30/45/60 days) and success criteria defined up front .
- Sequencing: most successful companies eventually run both playbooks — start with whichever yields the earliest, easiest sales, then verticalize into Lighthouse pushes for the top logos per vertical once mature (e.g., school districts, public sector), because procurement and sales motions differ .
- Sales org & careers: hire sales/revenue operations earlier than typical — even one person owning territory alignment, comp design, and a "sales constitution" ; early-stage teams should aim for ~100% quota attainment by setting reasonable goals that attract winners, since cost of sales matters less than winning the market . Early-career advice: join the best company that will grow (a "career elevator") rather than chase title or commission .
- Market context: the product-led-growth era was an artifact of the last adoption cycle (2000–2010 SaaS platforms left only wedge-product openings); AI agents are reopening core platforms like CRM/HR, so "there's a moment right now to go sell big software again," and buyers are more educated than ever, making sales easier and more self-serve .
- Common early mistake: analysis paralysis — spend ~1% of time on strategy and 99% executing; talk to customers and chase whoever will buy; "no bonus points for hard-earned revenue" .
Grok Bot was 48 hours old in public beta at research time, and early usage receipts sort into eight jobs people hire it for: office manager, chief of staff, life admin (calendars/reservations/tickets), teach-by-demo, inbox/Slack/morning briefs, sales ops/CRM rebuild, watchers/snipers, and content pipeline . Adoption already mimics company building — "people are already writing job descriptions, training new hires, adding managers, and firing bad fits with Grok Bot" and "software adoption is starting to look a lot like company building" . A plumbing company with zero engineers went from download to automated dispatch and office chores in 24 hours, then the owner fired his first bot ("Slow, very slow, and would make mistakes") and blamed his own job description .
Patterns worth borrowing for agentic products:
- Onboard by teaching, not prompting: the community converged on recording yourself doing the task once in the browser; "nobody wrote a 500-word system prompt" . xAI's frame: skill = the how, routine = the when — teach the skill by demonstration, then attach a schedule . A teacher turned syllabus + textbook + Google Classroom access into a homework-posting assistant in one hour . The no-model-selector chat interface is credited with making agents approachable for non-engineers .
- Fleet of narrow bots beats one generalist: power users run one bot per job with handoffs, usually routed through a chief-of-staff bot; the single do-everything bot is the beginner mistake .
- Trust = draft-and-approve by default: credible workflows keep a human gate at money-or-send moments — queued email drafts (36 LinkedIn drafts queued, zero sent), chat confirmation before booking, approval before spending .
- Security is scoped by connections, not by bots: all bots share one cloud VM, filesystem, and credential pool; xAI's docs reportedly say not to use separate bots as a security boundary, so scope risk by what you connect and keep beta off payments, production, and customer data .
- Lock-in is a pre-investment decision: Grok Bot is the zero-setup path; OpenClaw and Hermes are the portable ones — decide before sinking weeks of teaching into it .
Concrete early benchmarks: a 90-minute sales play now ~95% automated, a custom CRM rebuilt by a small bot crew in ~1.5 days, ~200 recruiting applications triaged to strong/mid/reject in one pass, a Salesforce end-of-quarter question answered in ~10 seconds, and competitor pricing pages monitored via Slack ping on change .
Launch traction: a reservation-booking demo hit 937 likes and ~4.6M reach via Elon's quote boost ; Alex Finn's setup playbook drew 2,734 likes / 3,385 bookmarks / 287K views ; Eric Zakariasson's hundred-item use-case catalog drew 1,918 likes / 2,928 bookmarks / 399K views . Cross-platform first-48h totals: X 61 posts / 14,884 likes; Reddit 31 threads / 6,440 upvotes; Instagram 20 reels / 994,396 views; YouTube 11 videos / 229,523 views . Most-requested missing feature: "multiplayer" — groceries and social plans shared with a partner .
LangChain's agent platform spans open-source agent development packages (LangChain, LangGraph), a commercial platform (LangSmith) for observing, evaluating, and monitoring agents plus infrastructure to deploy them, and a recently released no-code agent builder aimed at less technical users . LangChain's deployed engineers work across the full agent development lifecycle — understanding how to make an agent better after an initial version is generated — rather than one-off algorithm or function work . These engineers own the technical account from prospect through post-sales and double as the platform's biggest dogfooders, pushing fixes into the production codebase and carrying the strongest signal on customer-blocking issues as the interface between customers, product, and engineering . Hiring for this AI-facing customer role weighs business judgment (which features would drive retention or spend), comfort making and explaining assumptions from limited information, demo time management, agency, and demonstrated interest in AI more heavily than deep technical credentials; successful candidates come from software, MBA, and consulting backgrounds, and LangChain is compressing its interview loop as it scales the team .
PMs can now go from idea to a clickable prototype in a few hours; the awkward part starts after the first demo, when stakeholders revisit, engineering needs context, and the creator is left juggling screenshots, a Loom, a repo, and several URLs .
- Handoff patterns that work: use the prototype as a concept/validation tool and PRD reference, then demote it once the PRD is signed off so design produces proper UI/UX specs ; explicitly frame client-facing mockups as "design mockup… before I hand it off to dev" and link the prototype from the PRD and tickets . In enterprise settings, position it as "how your user journey may look", align stakeholders, then convert to user stories with the prototype attached to the PRD . One PM goes further: convert requirements docs into a pure front-end interactive product, let users refine it, then submit docs plus prototype to engineers .
- The prototype code itself rarely survives handoff — the PRD and Figma file do; engineering discards vibe-coded code because it lacks error handling, security, and scaling, while stakeholders only see a button that works .
- Expected complication: prototypes mask system complexity and create a false sense of delivery speed; a PM had to tell stakeholders "that is the m1 scope," and after two years of that mindset no system was finished enough to extend .
- Counterexample: one prototype reached production in ~2 months instead of an estimated year because it was built on real API endpoints with proper front/middle/back separation — engineers added auth and multitenancy and took it to prod ; the catch: the supporting backend already existed, and mock-data prototypes can add ~2 months .
- Design friction is the recurring tension: PMs are told to prototype with the design team rather than treat design as cleanup ; AI prototypes look "finished" and imply finality, which pushed one team to have design prototype first . Quality lapses (one PM's prototype repeated the same data value 11 times on a page) and ambiguity about which parts are changeable feed distrust .
- Career signal is contested: "continue vibe coding prototypes and you will stay stuck at junior levels" vs. value "lies in how to identify the best application scenarios… and enable maximum value" .
- Don't hand UI-only prototypes to users for independent testing — it implies the product is near production and can be counterproductive .
- Shreyas Doshi, former Stripe/Twitter/Google PM leader, shares the hiring mistake that cost him most: early in his career (and into PM leadership) he firmly believed everyone can change and grow to a tremendous extent in nearly every PM skill, and later realized this was wrong — most of his hiring mistakes were rooted in this belief .
- The belief was driven by ego (I can fix people), not genuine belief in potential: he would see a clear flaw in a candidate, suppress his hiring instinct, and rationalize it as coachable, committing to coach them on the missing skill .
- The outcome: the identified flaw would significantly block the person's path to next-level impact within a few months; coaching worked in only a minority of cases .
- Key realization: people do not grow on your timeline or in the direction you push them; they grow on their own timeline and in directions they are naturally attracted to .
- He also found the people never truly opted into coaching (despite agreeing when asked) and often treated it as a checkbox before asking for promotion .
- Practical advice: listen to the hiring intuition — if a candidate has a significant flaw in an area relevant to the job, don't hire them expecting growth; it's unfair to the individual, the company, and yourself . He attributes the mistake to a lack of empathy at a fundamental human level — when the ego (look how great I am as a manager) dominates, you can't empathize with others' true situations .
- In r/ProductManagement's weekly rant thread, PMs discussed roadmap-control failures: a CPO presented an all-hands roadmap with delivery dates for none of the work the product team was actually doing ; another PM said their roadmap was "rewritten in a hallway chat" they weren't in, leaving them owning an unagreed deadline . One suggested response: assert ownership directly — "I own the roadmap and you don't decide when it changes. I do."
- Multiple PMs reported fallout from removing QA: after leadership laid off all QA staff, the last two weekly releases were "trash" — a predictable outcome to the PM ; another company made product and designers do QA for two months until they pushed back, then hired QA contractors and later SDETs "at much larger salaries than the original QA team" . A PM inheriting an app with no designated QA (devs and domain SMEs doing all QA) found ~20 bugs within two hours . Advice: stop releasing until a release can actually be QA'd
- On role scope, a PM asked why they were doing the jobs of sales engineers, implementation managers, and customer support ; delegating is a suggested fix, but one PM warned that refusing to do pre/post-sale work got them viewed as obstructionist and hurt leadership's view of their performance
- A discovery failure example: a PM spent a year building and testing a voice-based interview-prep web app but can't get users; feedback was that it's unique and useful, but people aren't landing interviews in the first place . A peer challenged whether any research was done before building, noting people are still landing interviews, so the problem may be distribution, not demand
- Two PMs described role/agency gaps: one was "hired to do product" but does project management, with internal users setting priorities and a boss overriding everything ; another reported 70% of bandwidth going to tech debt and support while leadership promised four features for seven-figure deals by end of quarter, leaving the PM to deliver bad news
- A PM described a UX redesign driven by stakeholders and executives discussing their own preferences as if they are the users
Hiten Shah argues that AI is forcing more of our judgment to become reusable ; the gap is that AI produces faster than teams can decide if work is good, yet the quality bar is rarely written down — teams judge by instinct until disagreements surface . The fix: before testing AI work, define what doing the task well actually means; for tasks with fuzzy success (research, writing, support responses, code), a useful eval makes implicit judgment explicit enough to reuse, forming the quality bar . Implementation steps: assemble representative examples, judge them against fixed criteria, make a change, and compare — so you can answer "what actually got better?" and catch trade-offs like a prompt improving accuracy while adding verbosity, or a model helping hard cases but hurting common ones . Every eval encodes an opinion about what matters — accuracy, preference, task completion — and that choice steers what teams improve . Disagreements and failures are useful because they expose missing criteria (evidence quality, customer intent), so sharpen the standard whenever work exposes something unspecified . The eval set compounds: each failure becomes an example, each disagreement sharpens a criterion, and customer feedback adds detail, until the eval set is a record of what the team has learned about doing the job well . As AI scales output, more judgment must become reusable via examples, criteria, tests, known failure cases, and rules for when the AI should ask, retry, escalate, or stop . Shah is hosting a free live session, "Evals 101: How to Know If Your AI Actually Works," Friday, August 14, 10 AM PT .
The best interface for a home robot is one nobody has to learn — Hiten Shah's product principle, illustrated by the Matic robot: point at a mess, say "clean this," and it figures out the rest, an interaction it took nine years of work to make feel obvious .
Matic launched Cues, voice & gesture control for its home robot, after raising $115M and spending 9 years in development . The interaction model: "Hey Matic, clean this" while pointing at a spill makes the robot hear, see, locate the spill in 3D, and clean on its own; "Hey Matic" makes it turn and look at the speaker; "follow me" makes it follow behind; "go clean the living room" uses its house map to navigate there . Usability targets: so easy a 5-year-old and an 80-year-old can use it, and it understands 75 languages .
Outcomes so far: 13,000 families use Matic, and WIRED scored it 10/10 — the only hardware to receive that rating in a decade .
Shreyas Doshi argues that if you’re excellent at execution, consistent, and super-reliable but are not advancing as a leader, it is likely because you are not simulating before communicating . Simulation means anticipating how others will react to your default communication and modifying your communication so it produces the effect you desire; some do it instinctively . A concrete practice: take 5 minutes every morning to simulate each upcoming meeting — visualize how you want it to go, the role you will play, what could derail it, and how you want people to feel during and at the end of the meeting .
- Consumer-first strategy: MTV's business required juggling advertisers, cable operators, artists, and record companies, but the key was connecting with the consumer, building a research enterprise for insights, and earning loyalty — "consumer first, second, and third" — which made distribution, ads, and other deals fall into place .
- Constraint-driven innovation: MTV deliberately hired people with no TV experience, and the combination of no money and a passionate team forced totally new approaches — "there's nothing like having no money to force people to innovate" .
- Demand-pull distribution: To bypass local cable monopolists who refused to pay 10 cents/month, MTV ran the "I want my MTV" campaign to pull consumer demand through distributors — a deliberate, non-legal rule-breaking move in the cable business .
- Scaling culture: To preserve an insurgent culture beyond ~100-300 people, Freston kept the organization flat, made sure opinions were heard, de-emphasized politics, encouraged and tolerated risk, reinforced company values, and kept a casual, fun vibe (no dress code) — treating a creative, innovative culture as a competitive advantage .
- Hiring bar: "A players hire A players and B players hire C players"; substandard hires tend to hire similar people, eroding the company's vibe from within, so bad actors were not tolerated .
- Measurement and incentives: MTV made diversity hiring a goal on people's bonus plans, eventually reaching ~50% women managers; lesson: "you get what you measure and you get what you incentivize." However, raw hiring numbers were insufficient — retention required making people comfortable, which took false starts .
- Incumbent innovator's dilemma (MTV/Viacom case): MTV saw teens shifting attention to the internet early and knew they lacked the internal DNA, so they looked to acquire. They were the first to bid ~$1.5B for Facebook (half as earnout) but Zuckerberg declined; the Viacom board viewed YouTube as a "copyright infringement machine" and later sued for $1B and lost after 10 years, while YouTube grew to be worth ~$600B — a cautionary tale for big-company product decisions constrained by liability fear .
- Talent and point of view: "Have a point of view" and build relationships with key creative actors/talent, since talent sits at the center of all change .
- Experimentation model (A24): Freston highlights A24 as a new model: independent, low-cost projects, great instincts/taste, self-financed, and a live venue where they can test and pilot ideas, plus experimentation with short-form internet programming — a "mini conglomerate outside the mainstream" .
- Career advice: Freston advises stepping off the "conveyor belt" and embracing uncertainty; travel is "the world's greatest classroom," building empathy and a better-formed existence that makes you more attractive to recruiters. He credits "What Color Is Your Parachute?" exercises (evaluating skills, doing what you love) with leading him from a failed apparel business to MTV, and says living abroad taught humility, improvisation, risk-taking, and how to bet on and tolerate unusual people — a "perfect resume" for leading an unconventional media company .
Lenny Rachitsky observes that routine dental visits are a useful way to track AI adoption over time: six months ago his dentist had zero AI, and today AI is analyzing x-rays and transcribing the conversation; he wonders what the next six months will bring .
Hiten Shah argues that as AI gets better at producing work, leverage shifts to the people who can define what good work looks like; teams now need explicit, reusable quality bars because AI outputs arrive faster than teams can judge them .
Core concept: an eval is not just test infrastructure — a useful eval makes a team's judgment about quality explicit and reusable. Before testing work, write down what 'doing the task well' means: which sources matter, when the AI should ask instead of assume, what 'sounds like us' means, and which mistakes make output unusable .
Implementation: define the quality bar, pick representative examples, judge them against the same criteria, then change prompts/models and compare. A single score isn't necessary; what matters is visibility into what changed (e.g., accuracy vs. verbosity, hard vs. common examples) .
Use disagreements and failures to sharpen the bar: each exposes a missing criterion (evidence quality, customer intent), and the definition improves with every example. Over time the eval set becomes a record of what the team has learned about doing the job well .
Pitfall: grading on one dimension (e.g., factual accuracy or instruction-following) optimizes toward that dimension and can produce outputs customers hate; good always means good at a particular job per a particular definition of success .
Shreyas Doshi says he long believed that with the right coaching, anyone can build any product management competency at any time — a belief he now attributes to ego and a lack of true empathy, having learned otherwise the hard way; he discusses this in a new video (https://www.youtube.com/watch?v=yjWagfR69k8) .
In a r/ProductManagement thread on using AI for discovery/planning and handing off to devs , PMs shared concrete workflows:
-
One workflow: drop all project materials (emails, meeting transcripts, drafts) into a folder, use
claude coworkor Copilot to create a per-feature context file, turn those into user stories, and for large projects build a working mockup with the internal design system as the dev handoff . - Another PM built a Claude Code skill that, after exploring/designing/mocking, auto-creates an epic, stories, spec doc, and handoff prompt; when prioritized, devs point Cursor or VS Code at the epic and tickets auto-move across the Kanban board until testing, then auto-reassign to QA for test case, e2e, and manual review . A related setup connects a "second brain" to the repo and auto-creates user stories/epics/tasks for review after PRD and RFC validation, stressing the loop should be as closed as possible .
- A lighter option: keep standard user stories but add references to relevant context so the AI agent can fill gaps .
- Counter-signals: "AI generated stories are awful. Full of fluff and false rigor" , though training/refining the AI and human review can mitigate ; another comment advises using AI to enhance the current process rather than reinvent what works .
- A forward-looking view: dev/product handoff may not exist in five years because it will be the same person, and user stories may disappear if PMs and devs mostly review agent-generated output .
Hiten Shah asserts that anyone who reviews AI work is already doing evals, just one output at a time; this happens wherever people decide whether AI-generated output is good enough to use. The next step is making that judgment reusable so teams can test whether the work is improving . He argues evals aren't just for engineers anymore .
MitchellH called out low-effort AI-designed web pages (thin lines, glowy styles, inconsistent fonts, lots of monospace) as an immediate turn-off: he bounces off new products with that look before ever judging the product itself . Hiten Shah reposted it as a pet peeve of his too .
A founder building a tool for sports/trading-card collectors that reads and parses eBay purchase emails (card, price, purchase date) found no Plaid-for-email provider: Nylas and 'a couple others' require self-hosting plus Google security/compliance checks that are expensive and time-consuming, and Unipile (which claims you can leverage its license) didn't work well and is pricey . The current workaround is a personal mailbox where users auto-forward eBay emails; direct Gmail OAuth would be easier on setup but requires rigorous Google testing, and no established provider was found to bypass it pre-PMF . Commenters attribute this to Gmail read scopes sitting in Google's restricted category: apps need an annual CASA security assessment plus reviewer walkthrough before real users can grant access, and again whenever the app changes meaningfully; Nylas/Unipile pricing reflects having already completed that process, while auto-forwarding avoids requesting the scope entirely — which is why forwarding keeps winning despite feeling like a downgrade . One counter-signal: third-party primary-email access is a real security risk, forwarding puts the user in control, and direct eBay integrations may be a better path .
A marketer posted that after engineering quoted 3 sprints for a simple landing page + lead form, they "vibecoded" the same thing overnight on Emergent + Claude and shipped it the next morning — an example of AI coding tools letting non-engineers build and ship around the engineering queue .
A top commenter pushed back with a PM-relevant framing: marketing teams often control the front page because they're constantly testing variants, small changes are deliberately kept off engineering because engineers work on more complex problems, and "3 sprints doesn't mean 3 sprints of effort. It means that is where it is in priority" — a useful lens for PMs explaining to stakeholders why small requests take long .
On r/prodmgmt, a PM asked about Google's post-interview "fit loop" (pre-hiring committee), wondering whether it is a casual chat or more intense . A commenter advised treating the fit loop as a real interview that screens for red flags, recommending preparation of product-sense stories and collaboration/conflict-handling examples, and noted that Google hiring is currently very slow .
A product manager describes still copying customer interview quotes into PRD specs by hand despite testing AI tools . ChatPRD writes PRD text but lacks customer context; BuildBetter was the closest fit, reading call transcripts and threading quotes into spec sections, though its templates required tuning . The PM’s current workflow remains Dovetail plus manual spec writing, highlighting a perceived tooling gap for automated interview-to-PRD generation .
r/ProductManagement comment by u/unmanaged_chaoas
PMs who vibe-code prototypes: what happens after the first demo?
I keep seeing PMs get from an idea to a clickable prototype in a few hours now.
The awkward part seems to start after the first demo: stakeholders want to revisit it, engineering needs context, and the creator ends up with screenshots, a Loom, a repo, and several slightly different URLs.
For people doing this at work, what actually survives the handoff?
- a live prototype
- a recorded walkthrough
- Figma or a PRD
- the source repo
- some combination of these
I’m especially curious what context prevents a prototype from being mistaken for production-ready software.
I built out a prototype sitting on top of real API endpoints for various calculation engines. The prototype was built to see feasibility and market validation. I had built it as an app, proper front end, middle tier and back end separation.
When the decision was made to make it a full scale product, the engineers took my Git, bolted on Authentication and authorization, made the storage more robust, took the single tenant prototype to be a multitenant framework, updated the internal APIs for the changes, and pushed to Prod.
What should have taken a year to stand up was in Production in 2 months.
PMs can now go from idea to a clickable prototype in a few hours; the awkward part starts after the first demo, when stakeholders revisit, engineering needs context, and the creator is left juggling screenshots, a Loom, a repo, and several URLs .
- Handoff patterns that work: use the prototype as a concept/validation tool and PRD reference, then demote it once the PRD is signed off so design produces proper UI/UX specs ; explicitly frame client-facing mockups as "design mockup… before I hand it off to dev" and link the prototype from the PRD and tickets . In enterprise settings, position it as "how your user journey may look", align stakeholders, then convert to user stories with the prototype attached to the PRD . One PM goes further: convert requirements docs into a pure front-end interactive product, let users refine it, then submit docs plus prototype to engineers .
- The prototype code itself rarely survives handoff — the PRD and Figma file do; engineering discards vibe-coded code because it lacks error handling, security, and scaling, while stakeholders only see a button that works .
- Expected complication: prototypes mask system complexity and create a false sense of delivery speed; a PM had to tell stakeholders "that is the m1 scope," and after two years of that mindset no system was finished enough to extend .
- Counterexample: one prototype reached production in ~2 months instead of an estimated year because it was built on real API endpoints with proper front/middle/back separation — engineers added auth and multitenancy and took it to prod ; the catch: the supporting backend already existed, and mock-data prototypes can add ~2 months .
- Design friction is the recurring tension: PMs are told to prototype with the design team rather than treat design as cleanup ; AI prototypes look "finished" and imply finality, which pushed one team to have design prototype first . Quality lapses (one PM's prototype repeated the same data value 11 times on a page) and ambiguity about which parts are changeable feed distrust .
- Career signal is contested: "continue vibe coding prototypes and you will stay stuck at junior levels" vs. value "lies in how to identify the best application scenarios… and enable maximum value" .
- Don't hand UI-only prototypes to users for independent testing — it implies the product is near production and can be counterproductive .