We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
Big Ideas
Agents are becoming a first-class product user. Scott Belsky argues that an enterprise cloud product without a roadmap for agents is already behind: agents charged with workflow design will favor efficient, accurate, agent-ready services. Muse makes the direction concrete—an always-on assistant that uses a browser and connected apps, with a 1Password partnership for existing logins. Belsky’s product test is proportionality: users will demand trust and utility commensurate with the data access they grant. PM implication: access, identity, and trust signals belong in the core experience, not a later integration.
Consumer AI may reset the retention playbook. Andrew Chen’s thesis is that workplace AI wins on patterned, verifiable drudgery, while consumer products depend on novelty, authenticity, and parasocial trust; visible AI slop undermines those advantages. He argues that if apps become as cheap to make as content, average retention could collapse and greater-than-20% D30 may stop being the right yardstick. The operating response is a fast cultural feedback loop plus human taste and editing: move quickly, but do not copy-paste AI into customer-facing work.
Tactical Playbook
Separate task conflict from relationship conflict. Teresa Torres and Petra Wille recommend using retrospectives and explicit team charters to surface how a trio works together. Bring disagreement about the work into shared discovery; handle interpersonal friction one-on-one first. When it is opinion versus opinion, run an experiment rather than escalating a debate, and use joint escalation when leadership must intervene.
Instrument the economics, not the whole company. In a low-maturity SaaS organization, first identify the customers driving revenue, renewal timing, gross retention, and upsells; then choose a thin slice of telemetry to improve a specific decision. Existing APIs or open-source tools, one willing developer, and privacy sign-off can be enough to start. A dashboard for one team is a better first move than waiting for a company-wide transformation, especially when “just getting changes out” is crowding out measurement.
Case Studies & Lessons
Grok Bot: isolate, observe, unship. The team built from scratch in a small, physically isolated group; it reached an internal prototype about a month after the first line of code, then went from internal beta to public launch in roughly three weeks. It manually onboarded a couple hundred users over two weeks, including a coffee-shop owner, and watched users develop workflows rather than prescribing them. Before launch, it removed experimental and developer-facing features, focused on backend failures that blocked real work, and reframed the roadmap from “what has been added?” to “what can the product now do?” The team says 99% of automations are created in natural language. The lesson is not simply speed: tight decision loops, unusual users, and aggressive subtraction made capability—not feature volume—the launch standard.
Legora: pair usage with economics. Legora’s scorecard combines 95% gross retention, 300%+ NRR, DAU/MAU above 50%, 17 hours of monthly usage per active user, a 78% competitive-pilot win rate, and positive, improving gross margin. For B2B AI products, usage is an early signal that the work matters; retention, expansion, win rate, and margin test whether that importance becomes a durable business.
Career Corner
Make strong product work legible upward. One Apollo PM scored 8.75/10 at moving metrics but 6/10 at stakeholder management; the CEO noticed the second gap and lacked visibility into a high-performing area. The fix was to map the real stakeholder network, publish discovery takeaways where the trio could see them, and secure an all-hands slot. Two quarters later, the CEO raised the area unprompted at a leadership offsite. The career lesson: every meeting is an interview, and metrics alone do not make impact visible.
Tools & Resources
Local-first reference: Desert Ant Labs launched 18 on-device models across audio, vision, and text with Swift, Kotlin, and JavaScript SDKs; its pitch is no tokens, no logins, and no data leaving the device. It is a concrete reference point for evaluating when latency, privacy, and per-call cost justify moving intelligence closer to the workflow.
- Use a tightly scoped incubation model for ambiguous new products. The team started from a blank page with a small, physically isolated group and private communication channels; the first functional internal prototype took about one month, allowing rapid daily micro-decisions. They then rolled it out company-wide as a reality test before preparing for scale, reaching public launch roughly three weeks after the internal beta.
- Make early access a hands-on discovery loop, not a passive feedback channel. Over about two weeks, the core team manually onboarded a couple hundred users, joined the calls, and fixed confusing or broken experiences immediately. They deliberately recruited nontraditional users, including a coffee-shop owner, to expose use cases and integration bugs that internal dogfooding and a Silicon Valley AI user base would miss. The team avoided prescribing an internal “chief of staff” bot pattern, waited to see whether external users independently adopted it, and only then lightly encouraged the pattern in the product.
- Prioritize reliability and unlocked workflows over visible feature volume. During beta, the team removed experimental features and developer-oriented observability from the main surface, then focused on backend problems that caused real tasks to stall. They catalogued the tasks users actually attempted, quantified performance across task categories, and used concrete failures from sales workflows to guide fixes; each improvement was validated by previously blocked work becoming usable.
- Use an outcome-based roadmap test: “What can the product now do?” The team asked what the launch announcement for each change would be and rejected work that users would not directly feel. This reframed roadmap decisions from adding buttons, tabs, or integrations to delivering capabilities, such as defining automations in natural language; the team reports that 99% of platform automations are created this way.
- Choose a new product surface when the existing product creates audience and vision constraints. Rather than placing knowledge-work functionality inside Cursor, the team judged that a coding-product brand could intimidate nontechnical users and that adding tabs would create a cluttered surface with multiple competing visions. Starting from scratch enabled a simpler, more consistent experience for general knowledge work, despite the cost of building a separate product.
- Design AI products around the human teammate model. The team made the agent cloud-based and persistent, with its own computer, so users could access the same state from different devices and the agent could complete work in systems without reliable APIs. They also favored long-lived, role-based agents with memory and broad tool access over one-off chat sessions. For ambiguous product debates, they use a “what would I want from a human teammate?” test to clarify the desired behavior and guide the product toward conversational intent with fewer exposed controls.
- Treat product strategy as continuous adaptation. The team argues that AI products must reinvent priorities and user experience as model capabilities change, rather than assuming that what worked six months ago remains correct. Its operating principles are to delete scaffolding that models can eventually handle themselves and empower people to fix problems without waiting for permission.
- Sequence distribution from individual aha moments to organizational adoption. The team expects early adopters to discover transformative use cases personally and then demand similar capabilities at work; its next focus is bots operating inside business teams, company systems, and organizational context rather than only serving one user. For first-run activation, it recommends connecting everyday tools, asking the agent to identify five tasks it could remove, and immediately delegating the genuinely useful ones; power users can centralize the resulting outputs in a readable store.
Small, isolated incubation: To expand beyond developer tooling, the team created a very small group, isolated it from the company, and built from scratch; the first code-to-useful-internal-prototype cycle took about one month, followed by roughly three weeks from internal beta to public launch. Extending the coding product was considered but rejected because its technical associations, intimidating feel, and increasingly cluttered shared surface risked an inconsistent product vision; starting fresh gave the team control of the entire experience.
Manual discovery and validation: The core team manually onboarded a couple hundred users over about two weeks, joined 20-minute sessions, and fixed observed friction before the next onboarding. They deliberately included unconventional users, such as a coffee-shop owner, to uncover non-developer workflows and counter their Silicon Valley AI bubble. Rather than prescribing a workflow, they watched users independently develop a multi-bot/chief-of-staff pattern, then made the product slightly opinionated while keeping the choice reversible.
Prioritize reliability and simplicity over feature volume: During beta, the team removed experimental and developer-facing visibility features, then focused on making meaningful user tasks work by grouping task categories, measuring progress week over week, and fixing backend failure points such as browser-control and login problems. Their launch test was essentially: what is the user-facing launch post, and will users directly feel the capability? They reframed work from adding buttons and integrations to giving the bot capabilities, removed unnecessary UI, and used natural-language instructions for automations; 99% of platform automations were reportedly created that way.
Use a human-teammate test for ambiguous product decisions: The team’s north star is to ask how a human teammate should behave in the same situation, then derive the necessary product, model, and infrastructure changes. This steers the experience toward autonomous, low-micromanagement collaboration and familiar patterns such as short synchronous huddles when async communication is insufficient.
Treat architecture as product strategy: Two early choices were to run work entirely in the cloud, preserving the same state across devices, and to give each bot its own computer so it could interact with browser interfaces and pixels where APIs or MCPs were unavailable. The resulting model is a set of long-lived role agents that learn over time, rather than one-off chat sessions.
Adapt continuously and take the product into the organization: The team avoids demoware, uses early adopters to push products to their limits, and expects personal aha moments to pull AI into workplace adoption; its enterprise focus is shifting from single-user use cases toward team workflows, real company systems, and organizational memory. The broader operating lesson is to keep translating capability changes into major product reinventions, delete scaffolding as models improve, and build a useful product today rather than planning backward from an abstract moat strategy.
- Grok Bot went from a tiny team building separately in the office to internal obsession four weeks after the first line of code, then launched globally three weeks later; the source says millions of bots were helping people get work done less than a month after launch.
- The team chose to start Grok Bot fresh rather than build it into Cursor, making product architecture and organizational independence an explicit launch decision.
- The launch process emphasized manually onboarding the first 200–300 users and deliberately “unshipping” features during the weeks before launch—useful examples of high-touch discovery and scope reduction before public release.
- Nir Eyal says the Hook model’s four stages—trigger, action, reward, and investment—remain unchanged because human psychology has not changed. For AI-enabled products, he identifies the investment stage as the main opportunity: use information supplied by users to make the product better with use, creating “stored value” and enabling personalization toward a “market of one.”
- For AI product strategy, automate low-effort work that AI handles well while preserving work that depends on human originality and surprise; Eyal argues current AI is poor at generating genuinely novel ideas.
- Grok Bot launch case: Roman Ugarte incubated and now leads product for Grok Bot, which began with a tiny team working separately from the rest of the office and building from scratch rather than adding the product to Cursor.
- The product-development process was highly compressed: the post says the company was obsessed four weeks after the first line of code, launch followed three weeks later, and the team manually onboarded its first 200–300 users before launch.
- The team also unshipped features in the weeks before launch, while the post identifies two non-obvious bets as major success factors but does not describe those bets in the excerpt.
- The post claims that millions of Grok Bots were already helping people get work done less than a month in, though the excerpt provides no supporting usage metrics beyond that claim.
- For AI product and business health, use a five-metric scorecard: gross retention, net revenue retention, DAU/MAU and usage, competitive pilot win rate, and gross margin. Legora reports 95% gross retention, 300%+ NRR, DAU/MAU above 50% with 17 hours of monthly usage per active user, a 78% competitive-pilot win rate, and positive, improving gross margin.
- Apply usage as an early signal of whether the product is becoming important before that impact appears in retention or expansion metrics; treat gross margin as the sustainability check, since scaling a negative-margin model magnifies the problem. Paul Graham linked this metric profile to Legora’s reported 9x annual growth, specifically highlighting its 78% pilot win rate and 95% gross retention.
- Use problem decomposition as a core PM interview signal. Evaluate whether a candidate can identify the true problem, model it as data, define repeatable manipulations, and establish a feedback loop to a system or user-facing visualization. The specific framework matters less than demonstrated critical thinking and the ability to explain decisions; also probe how candidates have developed weaker business or technical skills and whether they show curiosity and lifelong learning.
- Treat AI as acceleration for discovery, not a substitute for strategy. Use it to automate unstructured-data work and create clickable prototypes earlier for fast feedback, while keeping human judgment and initial problem framing in the process. A practical guardrail is to exclude AI from early-stage thinking except when it serves discovery: a prototype can test whether something works, but not whether the team should pursue it or whether it addresses the right problem.
- Build executive communication through deliberate rehearsal. PMs who are shy or ramble in meetings can practice with a mirror, family, a partner, or a trusted colleague to build muscle memory, then develop speaking endurance and prepare more heavily when needed.
- Advance by combining communication with proactive ownership. Strong senior PMs stand out through clear communication grounded in real subject-matter understanding, proactive ownership of a roadmap area, and initiative in finding revenue opportunities; organizational politics may become more influential as company size and seniority increase.
- Product trios should distinguish task conflict—disagreement about the work—from relationship conflict—friction in how people work together—because the two require different responses. Surface task conflict within the team and use shared discovery to resolve it; address relationship conflict one-on-one first, with HR or mediation as an escalation path rather than the starting point.
- Use retrospectives to examine how the team is working, and treat facilitation as a learnable product skill rather than something to outsource by default. Make team norms explicit in a charter, including working hours, communication preferences, and decision-making styles; choose facilitators deliberately based on the context and the need for neutrality.
- When disagreement is opinion versus opinion, design an experiment instead of debating indefinitely. If escalation is necessary, have the parties document the conflict jointly so leadership receives a concrete account.
- Grok Bot was built from scratch by a tiny team working separately from the rest of the company, rather than being added to Cursor. The team reportedly reached broad internal adoption four weeks after writing the first code, launched three weeks later, and had millions of bots helping people within its first month.
- The launch process included manually onboarding the first 200–300 users and deliberately “unshipping” features in the weeks before launch; the case study also centers on two non-obvious product bets, though their details are not provided here.
Cursor’s product strategy: Despite competing with well-funded, fast-growing companies while building on top of those same competitors, Cursor focused less on defensible moats and more on making the product useful immediately. Roman Ugarte, Cursor’s employee #15, identified this as the key to pulling it off.
- PM career insight: @tarstarr says the scarce PM skill “right now” is “elevating other people’s ambition.”
- Manage up by translating execution into strategic language. Reframe operational blockers as strategic or product risk, use the manager’s preferred framework vocabulary, and ask specific “in practical terms” follow-ups instead of debating theory abstractly.
- Make theoretical advice testable. Ask the manager to co-own one concrete decision or experiment, with trade-offs and success criteria documented; alternatively, present a decision that requires a yes/no answer by a defined deadline.
- Use evidence that can travel upward. For complex products, package the specific workflow step, affected accounts, support cost, and delivery data; detailed, repeatable evidence is more useful than arguing about frameworks.
- For willingness-to-pay discovery, assess three dimensions: whether the buyer and user are the same and whether procurement is involved; whether the problem is a must-have, a costly workaround, or merely nice-to-have; and whether prospects are already seeking a solution or first need problem awareness.
- Build evidence incrementally instead of trying to transform the whole organization. In a low-maturity SaaS environment, basic measures such as active users, adoption, retention, and feature outcomes may be unavailable, so start with the team you control: create a minimal dashboard or mission, clean up one neglected process, and lead by example.
- Instrument a thin slice with business and privacy constraints in mind. First understand which customers and revenue mechanics drive success—including renewals, gross retention, and upsells—then identify where telemetry would improve decisions. A practical implementation can use existing internal services or APIs, a friendly developer, and open-source or self-hosted tools; one reported team used PostHog’s free tier after GDPR sign-off to avoid collecting identifying data. Expect privacy and data-security objections, especially when instrumentation competes with a backlog that rewards shipping output over measuring outcomes.
- Use metrics as decision support, not as universal benchmarks. Commenters cautioned that NPS comparisons across different products are unreliable, while an internal relative-health view can still have limited value; telemetry and business/customer outcomes are stronger evidence. When data is incomplete, the PM should triangulate what is available and narrate the evidence rather than wait for perfect analytics.
- Create a career evidence trail around outcomes. Capture customer conversations, usability tests, in-product surveys, feedback themes, and the improvements they influenced, then frame interview stories around customer and business impact—including revenue-related results—rather than NPS alone.
AI is shortening the feedback loop around technology: cheaper, improving models let individuals attempt work that once required teams or years of accumulated skill, while new capabilities change user behavior and expectations, feeding the next round of software and ways of working. For PMs, the recommended operating posture is to stay oriented as conditions shift, test ideas early enough to learn, identify what changed and matters, spend enough time with problems to develop judgment, and update beliefs without rebuilding their worldview every week.
- Match the build sequence to the economics of being wrong. Use more preparation for decisions that are costly, consequential, or hard to reverse; use Fire, Ready, Aim when an initial attempt is cheap, reversible, and informative. A/B testing still requires measurement infrastructure before the test can teach the team, while prototypes can often be built before extended debate.
- Use prototypes as a product-thinking tool, not just an execution artifact. When a working artifact costs less than the discussion required to specify it correctly, build a sandboxed version immediately; interacting with something real can expose unnecessary features, resolve abstract disagreements, and reveal new ideas.
- Optimize experiments for learning, not output volume. A first attempt should have a small blast radius and a large information radius. Teams must distinguish meaningful signal from random results, define how they will validate the result through evaluations, measurement, or expert judgment, and prioritize the time from an important uncertainty to a trustworthy change in belief.
- Scale caution with consequences and commit when evidence is sufficient. Use prototypes, sandboxes, models, or simulations to create cheaper learning loops when real-world attempts are too risky; stop iterating once evidence earns conviction and commit to the decision.
Keep the people crafting customer-facing messages closely connected to the people hearing customer feedback; separating those groups risks producing generic messaging.
Explicit task scoping controls how far an AI system explores. The researchers describe models as highly task-oriented: when asked to improve bounds for codes, the model produced an improvement but did not continue pushing; a follow-up request led it to develop more sophisticated representation-theory work. For open-ended research or optimization, use staged prompting: define an initial task, inspect the method and result, then explicitly ask the system to extend or deepen the work rather than assuming it will optimize beyond the original brief.
Separate strategic direction from long-horizon execution. The discussion notes that both humans and models can become “pigeonholed” in an unproductive path, while a fresh session, parallel agent, or outside reviewer can reset the context and challenge the approach. One proposed architecture assigns one model responsibility for judgment about direction (“taste”) and another responsibility for persistent execution, reducing the risk that exploration and execution contaminate each other’s context.
Design AI products for comprehension and knowledge organization, not only output generation. The speakers argue that as models make producing sophisticated work easier, understanding, absorbing, explaining, and structuring that work become more important constraints. AI workflows should therefore pair generation with explanation of strategy, review, and reusable knowledge structures so users can build on outputs rather than merely receive them.
- At the pre-launch MVP stage, the founder reports being the bottleneck across user interviews, product development, marketing, and strategy, with unproven demand and tight cash.
- Use a discovery-first loop: test whether the target industry resonates, ask interested prospects to try the MVP and give honest feedback, and use tester responses to inform both product decisions and the eventual marketing/sales motion.
- Keep early staffing lean: retain the founder and developer, use freelancers for the specific bottleneck slowing MVP progress, and hire around needs that recur after testing.
- Turn staffing into a workload audit rather than a headcount plan: separate founder-only, repeatable/delegable, and temporary specialist work; select one outcome delaying the MVP, define decision rights and a 2–4-week deliverable, and test collaboration through paid contract work. Evaluate whether the person ships, communicates, and decides independently before choosing a contractor, employee, or co-founder; equity should accompany a clearly defined long-term role.
- Sequence execution to reduce overload: complete user interviews and validation before product development, then product development before marketing; limited early design exploration can run in parallel only when it does not compete with the founder’s core validation work.
- When a formal arrangement is premature, one founder describes a defined one-month trial with no promises as a way to distinguish people who ship from people who only express enthusiasm; if equity is later used, they recommend four-year vesting with a one-year cliff.
Hiten Shah argues that AI will make mediocre software “incredibly cheap,” potentially increasing the relative value of great software—a product-strategy signal for PMs to emphasize meaningful quality and differentiation.
- Consumer AI product thesis: AI performs well on patterned, verifiable workplace tasks such as forms, process-following, updates, and boilerplate code, but consumer products depend on novelty, authenticity, and parasocial trust—areas where pattern-matched AI output can feel like “slop” and undermine engagement.
- Retention may become the wrong north-star: As AI makes apps nearly as cheap and fast to produce as content, consumer products may generate short-lived spikes rather than durable usage; the author argues average retention could collapse and that metrics may shift from conventional measures such as D30 retention toward how many apps reach how many users, alongside the repeatable system, brand, or IP that produces hits.
- Operating implication for consumer PMs: In an oversupplied, zero-sum attention market, teams should use a fast feedback loop to identify the cultural zeitgeist, ship something novel slightly ahead of competitors, and balance speed with the risk of becoming random or alienating. AI can assist ideation, but human judgment and editing remain important because copy-pasted or visibly inauthentic AI messaging can rapidly destroy trust.
“The scarce PM skill right now is elevating other people’s ambition.” — @tarstarr (opens in new tab)
- PM career insight: @tarstarr says the scarce PM skill “right now” is “elevating other people’s ambition.”