We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
Big Ideas
Make AI evals a discovery habit, not a late QA step. Teresa Torres defines evals as methods for measuring whether an AI product or workflow performs as intended; they help teams maintain quality, catch issues before users do, and create a feedback loop alongside interviews and assumption testing. For PMs, the practical shift is to specify expected behavior and test it while the product hypothesis is still changing.
For agents, measure ownership rather than engagement. Hiten Shah argues that a more useful agent may generate fewer messages and sessions because the customer is spending less time supervising it. A successful first run proves capability, not ownership: run a real job for at least three cycles, log every intervention, and separate valuable human judgment from accidental handbacks such as rebuilding context, restarting work, or checking whether the task finished. The target is declining accidental operator work while consequential decisions and exceptions remain visible.
Tactical Playbook
Score evidence per assumption before scaling. Strategyzer’s readiness framework separates desirability (do customers need and want it?), feasibility (can it be built and delivered?), and viability (can it create business value profitably). A board-approved business case is still a hypothesis; score the evidence for each aspect rather than claiming the whole idea is validated.
Use the evidence ladder to choose the next test: business plans and high-level research are level 0; statements and reactions are levels 1–2; low-stakes actions such as signups or sales-call requests are level 3; pilots, letters of intent, deposits, and pre-orders are level 4; real market behavior or a Wizard-of-Oz test reaches level 5. More interviews do not strengthen evidence if prospects never commit. In one example, Fireflies.ai charged $100 per month while founders manually joined meetings and sent summaries: payment and continued use tested desirability and viability, but the non-scalable delivery left feasibility unproven.
Case Studies & Lessons
Murmur productized the bottleneck around coding agents. Macroscope says an internal tool orchestrated 90% of the code it shipped over two months, leading it to release Murmur as an early preview. Its diagnosis was not primarily model weakness: engineers were bottlenecked by local development environments and the post-PR lifecycle of review comments, failed CI, rebases, and merge conflicts.
Murmur gives each agent a dedicated cloud VM, lets it test and verify its work, and keeps it moving through GitHub events until human approval. A local Claude Code or Codex session can act as a director for a fleet, while Slack, Linear, GitHub, REST, and MCP make existing work systems entry points for bounded agent tasks; engineers stay focused on work requiring deeper context or judgment. The product lesson is to own the workflow around the model—not just expose model capability—and make access profiles, deployment choices, and audit logs part of the product for enterprise use.
Career Corner
AI fluency is becoming a hiring signal for PMs. Aakash Gupta reports that 76% of 113 PM job postings he reviewed asked for AI knowledge, and lists evals, RAG, agents, MCP, observability, context engineering, and LLM-as-judge among the emerging vocabulary. A community discussion raises the unresolved consequence for entry-level roles: AI can now accelerate documentation, research, analysis, SQL, prototyping, and workflow creation, potentially shifting the bar toward independently identifying problems and showing judgment. Build evidence of both: a small AI-enabled product workflow with explicit evals and instrumentation, plus a clear explanation of trade-offs and failure modes.
Tools & Resources
Explore local AI-assisted product analytics. A community-built open-source MCP server lets an agent run data-science analysis on tabular event data to examine retention, churn differences, underused features, segment behavior, drop-off points, and pre-conversion actions; computation stays local and the model receives analysis results rather than the full CSV. Treat it as a hypothesis generator, validating event definitions and conclusions against raw data and customer conversations.
| Source | Docs | Insights | Status |
|---|---|---|---|
| rahulvohra | 0 | 0 | |
| Paul Graham | 0 | 0 | |
| Tony Fadell | 0 | 0 | |
| Patrick Collison | 0 | 0 | |
| Daniel Ek | 0 | 0 | |
| Gustaf Alströmer | 2 | 0 | |
| Stewart Butterfield | 0 | 0 | |
| PM Diego Granados | 0 | 0 | |
| 👨🏻💻☕️ | 0 | 0 | |
| scott belsky | 4 | 2 | |
| Ryan Hoover | 1 | 0 | |
| Janna Bastow simplybastow.bsky.social | 0 | 0 | |
| Jackie Bavaro | 0 | 0 | |
| Sachin Rekhi | 2 | 0 | |
| Dan Olsen | 0 | 0 | |
| The community for ventures designed to scale rapidly | Read our rules before posting ❤️ | 129 | 17 | |
| Will Lawrence | 0 | 0 | |
| Product Marketing | 0 | 0 | |
| Ami Vora | 0 | 0 | |
| PM Interview: Practice Group for Product Manager Case Interviews | 0 | 0 | |
| One Knight in Product | 0 | 0 | |
| Aakash Gupta | 1 | 1 | |
| Shreyas Doshi's Product Almanac | Substack | 0 | 0 | |
| Lenny Rachitsky | 0 | 0 | |
| Acquired | 0 | 0 | |
| a16z | 1 | 1 | |
| Exponent | 1 | 1 | |
| Product Alliance | 0 | 0 | |
| Product Management Exercises | 0 | 0 | |
| rocketblocks | 0 | 0 | |
| Product Design | 0 | 0 | |
| ProductManagementJobs | 28 | 7 | |
| Product Management | 48 | 9 | |
| Product Management - The place for all things product | 16 | 6 | |
| Product Management | 14 | 5 | |
| Aspiring and current tech PM's | 0 | 0 | |
| Masters of Scale | 1 | 1 | |
| Product Science Group | 0 | 0 | |
| How I built This | 0 | 0 | |
| SaaStr AI | 0 | 0 | |
| productized io | 0 | 0 | |
| Lenny's Reads | 0 | 0 | |
| The Product Folks | 0 | 0 | |
| Strategyzer | 1 | 1 | |
| Lenny's Podcast | 0 | 0 | |
| AJ&Smart | 0 | 0 | |
| Y Combinator | 1 | 1 | |
| Product School | 0 | 0 | |
| Mind the Product | 1 | 1 | |
| @andrewchen | 0 | 0 | |
| The Looking Glass | 0 | 0 | |
| Kyle Poyar’s Growth Unhinged | 0 | 0 | |
| Leah’s ProducTea | 0 | 0 | |
| Run the Business | 0 | 0 | |
| Product Managers at Work | 0 | 0 | |
| The Product Compass | 0 | 0 | |
| Ravi on Product | 0 | 0 | |
| Productify by Bandan | 0 | 0 | |
| Product Thinking with Melissa Perri | 0 | 0 | |
| Product Talk Daily | 0 | 0 | |
| The Beautiful Mess | 0 | 0 | |
| Gibson Biddle's "Ask Gib" Product Newsletter | 0 | 0 | |
| Casey Accidental | 0 | 0 | |
| Hiten Shah | 12 | 5 | |
| Product Growth | 0 | 0 | |
| Perspectives | 0 | 0 | |
| Lenny's Newsletter | 0 | 0 | |
| andrew chen | 8 | 4 | |
| Brian Balfour | 0 | 0 | |
| Casey Winters | 0 | 0 | |
| elena verna | 0 | 0 | |
| Kevin Weil 🇺🇸 | 7 | 1 | |
| April Underwood | 4 | 1 | |
| Julie Zhuo | 2 | 1 | |
| Marty Cagan | 0 | 0 | |
| Lenny Rachitsky | 5 | 0 | |
| Christian Idiodi | 0 | 0 | |
| John Cutler | 0 | 0 | |
| Teresa Torres | 1 | 1 | |
| Gibson Biddle | 0 | 0 | |
| Shreyas Doshi | 0 | 0 | |
| Adam Nash | 0 | 0 | |
| Merci Grace | 0 | 0 | |
| Jackie Bavaro | 0 | 0 | |
| Hunter Walk | 0 | 0 | |
| Brian Balfour | 0 | 0 | |
| Scott Belsky | 0 | 0 | |
| Nir Eyal | 1 | 1 | |
| Teresa Torres | 0 | 0 | |
| Julie Zhuo | 1 | 1 | |
| Andrew Chen | 0 | 0 | |
| John Cutler | 0 | 0 | |
| Ken Norton | 0 | 0 | |
| Gibson Biddle | 0 | 0 | |
| Elena Verna | 0 | 0 | |
| Casey Winters | 0 | 0 | |
| Shreyas Doshi | 0 | 0 | |
| Lenny Rachitsky | 0 | 0 | |
| Melissa Perri | 0 | 0 | |
| Marty Cagan | 0 | 0 |