We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
Big Ideas
Evaluate AI against the work, not the leaderboard. The current frontier-evaluation debate shows why public scores are insufficient: the interview cites Llama 4 underperforming on held-out private benchmarks while appearing highly capable on public ones. For PMs, the implication is to build an evaluation set from real enterprise workflows, with rubrics that account for capability, cost, latency, and flexibility. Longer-running agents also require stable, retryable evaluation infrastructure and fewer tasks assessed against richer criteria.
The next personalization layer is workflow ownership. A current product essay argues that software should become malleable rather than expose a fixed settings menu: let users rearrange the interface, extend preferences beyond predefined toggles, turn feedback into product changes, and treat the seams between tools as part of the experience. The author’s half-hour rebuild of an eight-agent workspace illustrates the wedge: remove repeated, personal friction before adding broad feature breadth.
Tactical Playbook
Test pull with friction before reducing friction. A fast AI prototype proves feasibility for its builder, not viability for a market; the Visory interview explicitly separates “pain” from demand and recommends testing whether people have actually acted on the problem.
- Ask for evidence tied to a real job. Before building, Visory requested two sensitive board packs plus an audio recording of the customer reviewing one; ten people willing to clear that hurdle became the threshold to start.
- After launch, compare stated and revealed preference. Cross-pack search won enthusiastic reactions and demos, but time-poor users went straight to a key-signals view. Ask for three recent instances of the problem and what the user did instead.
- Model time to trust, not just time to value. Visory users checked the product manually for two cycles and only relied on it around the third; infrequent-use products may need a simulated second-cycle experience rather than a standard 30-day trial.
Case Studies & Lessons
Bolt’s pivot paired focus with commercial adaptation. Its team describes a cloud-IDE market with strong hype but weak willingness to pay; after several failed 2024 directions, Bolt was the last attempt before a planned shutdown, and reported ARR rose from $0.5M to $5.5M in 30 days. The company then narrowed toward professional product builders—especially PMs, designers, and engineers—and says B2B revenue grew 10× year over year. The PM lesson is to go deep on an ICP and its workflows: detached prototyping becomes more useful when it uses production components and has a clean developer handoff. When customers exhausted a $9 subscription in under a day, the team shipped usage-based pricing within 72 hours.
Career Corner
Optimize for operating language, not tool names. A sample of 421 PM postings across 62 company job boards found “roadmap” in 85.5% of listings, while cross-functional work, prioritization, and stakeholder management each exceeded 55%. Jira appeared in 1.7%, Figma in 2.4%, and Productboard in none; “experimentation” appeared in 24.2% versus 2.1% for “A/B testing.” Use the terminology of the target role when translating experience, but treat the study as directional: it covers tech companies on one ATS and excludes agencies and regulated employers.
Tools & Resources
Workflow-grounded evaluation: The interview describes ValSmith, which turns a company’s GitHub codebase into an internal benchmark for comparing coding agents on performance and ROI. It is a useful model for PMs evaluating AI vendors: the best general model may not be best for a specific repository, and token-efficient choices can be unintuitive.
Treat distribution and channel fit as product discovery. Visory had paying customers and was viable, but its intended direct-to-director B2B motion shifted to enterprise when boards requested multi-seat access and IT, AI-transformation, and legal involvement; Kirsten chose an existing channel rather than build that enterprise channel because it was not the company she wanted to run. The product continued under another name after the founders exited. AI reduced the cost of building but not the cost of reaching trusted users, so distribution should be treated as a separate hypothesis with its own tests and stop criteria, while adoption, willingness to pay, and retention are tested alongside product discovery rather than afterward.
Validate pull and viability, not the founder’s conviction. Being the target user gives useful context but is dangerous; real demand is another person trying to make progress and being blocked, so founders should treat themselves as one data point, validate jobs rather than features, and test against people who are not themselves. An agent or synthetic-user test can provide early usability input when grounded in real discovery interviews, but real people are needed to test viability and demand; pain alone is not demand. Visory deliberately added friction before building: directors were asked to provide two sensitive board packs and an audio recording of their own review so the team could compare the product with their manual process. Ten people willing to complete that demanding test was the threshold for starting development, providing stronger evidence than positive comments.
Use revealed behavior to prioritize features. Cross-pack search received the strongest research reactions and demo praise, but users barely touched it after launch; time-poor directors went instead to a key-signals view that triaged risks, execution gaps, and positive findings. Discovery should therefore ask for three recent instances of a problem and what the user did instead, distinguishing recurring jobs from Kano-style “delighter” or rainy-day features. Visory retained some infrequently used features because they remained valuable in demos and the product was still early, while keeping the core experience focused on primary jobs.
Model time-to-trust for infrequent-use products. Visory users manually compared the product with their own work in cycle one, gained confidence in cycle two, and only began relying on it around cycle three. Because board meetings could be monthly, quarterly, or even six-monthly, a 30-day trial and a model assuming two cycles understated the time needed to prove value; trials, pricing, churn, and financial assumptions should reflect trust-building, with a simulated second-cycle experience on the customer’s own data used to accelerate proof.
Price pilots to avoid creating a damaging anchor. ChatGPT and Claude established a roughly $20 reference point, so Visory positioned value against the cost of governance failure and director time rather than feature-by-feature comparisons. It modeled $300 per board seat per month but piloted at $150, which became the customer anchor and made the later move to $300 feel like a 100% increase; pilots should use the full price or a price close enough that the step-up does not feel punitive. Early payment infrastructure is also part of validating real willingness to pay.
Treat AI reliability as ongoing product work. Rapid AI development does not remove QA, regression testing, evaluation, or monitoring: models drift, providers retire models, and system behavior can change without a conventional software release. Commercial AI products should establish an evaluation and testing framework, visibility into model calls and data, and detection for provider changes from the outset; Visory’s multi-model analysis and adjudication system began selecting one model after another was retired overnight, illustrating the operational risk.
- Market-entry experimentation: 11 Labs defined a market-specific distribution mix across direct sales, resellers, and cloud channels; each launch had a thesis covering why to enter, how to sell, and the results expected within 3–6 months. The team then iterated through experiments rather than assuming one channel would work everywhere. The speaker attributed a reported revenue progression from $0 to $100M in 20 months, then to $200M, $330M, and more than $600M over progressively shorter periods, to repeating this distribution-and-experimentation loop.
- Activation and growth: Reduce both financial and setup friction, optimize time-to-value, and test variants such as grants or credits, usage-based/pay-as-you-go pricing, outcome-based pricing, and hands-on onboarding; the speakers emphasized that experiments should be adapted by market because what works in the US may not work in Germany or Brazil. 11 Labs operationalized this through a grants program that gave qualifying startups three months of free product access; it issued tens of thousands of grants, and the interview says participants later contributed more than 10% of enterprise revenue through upsell.
- AI adoption in existing teams: Position agents as productivity assistants rather than replacements, introduce role-specific agents such as AI SDRs, account executives, and customer-success managers, and prove value through faster inbound response, improved conversion, and long-tail upsell revenue. Compensating the human owners of accounts for AI-generated upsells was recommended to reduce internal friction while preserving human involvement in relationship-driven selling.
- Human/agent task allocation: Assign agents to data-heavy workflows such as TAM construction, account scoring, signal monitoring, research, CRM updates, and forecasting; reserve human capacity for buyer relationships and creative campaigns or experiments, where the speakers said people still have an advantage.
- Framework — hyperpersonalization: Julie Zhuo argues that software is moving beyond fixed settings and one-size-fits-all experiences toward letting users tune both interface and content to their own workflows and identity. The value comes from workflow utility—compounding friction in repetitive, personal, cross-tool tasks—and expression beyond bland utility.
- Product implementation: For high-frequency or highly personal workflows, make the interface malleable, let preferences extend beyond predefined toggles, use user feedback to change the product quickly, and treat the seams between tools as part of the experience.
- Case evidence: Zhuo built a mostly working alternative to her six-click remote terminal setup in half an hour, then optimized it for her eight-agent workflow, Claude/Codex usage, and personal friction points; she reports that it worked better for her because she was the best customer and could maintain a tight feedback loop. In a personalized learning app for her children, she used stories based on their interests and life events, added per-lesson feedback, and adapted rewards to requests. She also cites a study of 145 ninth graders in which interest-based algebra problems were solved faster and more accurately, with the largest gains among struggling students and benefits persisting after personalization was removed.
- Roman Ugarte’s Grok Bot team started as a small group working separately from the rest of the company and built the product from scratch; the team reportedly reached company-wide obsession four weeks after the first code and launched three weeks later. Early product development included manually onboarding the first 200–300 users and spending the weeks before launch unshipping features.
- A useful product framing is to optimize for what the product can now enable—“Our product can now…”—rather than cataloging what it contains with “Our product now has…”. This shifts prioritization toward user-relevant capabilities instead of feature inventory.
- Product architecture decision: Grok Bot’s product lead said the team chose a standalone app rather than adding it to Cursor because competitor products with a new tab for each form factor could feel like “shipping your org chart” instead of expressing one consistent vision; they therefore started from scratch.
- Launch execution: The product was incubated by a small team working separately, reached company-wide enthusiasm four weeks after the first code, and launched three weeks later; the launch recap claims millions of bots were helping users within less than a month. The team manually onboarded its first 200–300 users and deliberately “unshipped” features before launch.
- Grok Bot was built from scratch by a tiny team working separately from the rest of the company, rather than being added to Cursor; the team reached company-wide excitement four weeks after writing the first line of code and launched three weeks later. The post says millions of Grok Bots were helping people work within their first month.
- The launch process emphasized hands-on discovery and scope reduction: the team manually onboarded the first 200–300 users and deliberately “unshipped” features during the weeks before launch.
- Lenny’s product thesis is that an AI system completing 100% of a job feels categorically different from one that gets users 90% of the way there, highlighting completion quality as a meaningful product threshold.
- Hyperpersonalization framework: Treat hyperpersonalization as serving two value drivers: workflow utility—reducing friction in high-frequency, cross-tool work—and expression, letting users tune UI or content to their identity and preferences. It is most applicable to workflows repeated many times a day and to users whose needs vary by brain, lifestyle, or tool stack; implementation steps are to make interfaces malleable, let preferences extend beyond finite toggles, turn user feedback into product changes, and treat cross-tool seams as part of the experience.
- Dogfooding as a rapid product-development loop: Faced with a remote AI-agent workflow that was six clicks deep and could not combine Claude and Codex in the official apps, the author built a mostly working version in 30 minutes, then iterated on mobile and desktop for her eight-agent setup. Because she was the best customer, she could remove friction and add ideas immediately; she reported that the result worked better for her and avoided features she did not use.
- Personalized learning case: To motivate children without relying on streaks or short-term boosts, the author built lessons around each child’s interests, real-life events, trips, and family values, then added end-of-lesson feedback and collectible rewards in response to requests. Separately, a study of 145 ninth-graders found that students solved interest-based algebra problems faster and more accurately, with the largest gains among struggling students, and retained the benefit after personalization was removed—supporting evaluation of personalization against learning outcomes, not only engagement.
- StackBlitz found that the cloud IDE market had substantial hype but weak willingness to pay, as developers remained satisfied with local environments. After testing multiple product directions in 2024, Bolt became the final pivot before a planned shutdown; it grew ARR from $0.5 million to $5.5 million in 30 days. The case underscores validating willingness to pay and pivoting quickly when the original market is not commercially viable.
- Bolt narrowed its focus from a broad user base to professional product builders—PMs, designers, and engineers—and prioritized specific B2B workflows. The company reports that B2B revenue grew 10× year over year, while workflow depth and a clearly defined ICP became key sources of competitive differentiation.
- For PM and design experimentation, a detached environment can enable faster iteration and feedback than working directly in production. To make prototypes production-ready, Bolt is connecting them to the exact production components and creating a seamless developer handoff that pulls changes into code for engineering sign-off.
- AI product pricing is shifting from fixed per-seat models toward usage- and value-based pricing because agents can reduce the number of seats needed for some workflows. After users exhausted Bolt’s $9 subscription within hours of launch, the team shipped usage-based pricing within 72 hours, allowing customers to pay for the amount of usage they needed.
- Manage products against their lifecycle: validate during exploration, increase investment when growth accelerates, and use metrics and market shifts to identify a declining business. Senior product leaders and executives should assess the market, company position, and meaningful metrics, then decide when to double down or move on—even while an existing product remains profitable.
- Make portfolio reviews more objective by scoring each product on strategic fit, growth potential, current traction, and maintenance cost; this exposes assumptions and keeps discussion focused on trade-offs. Reviews should also identify what to stop or deprioritize, not only where to fund growth.
- Before approving another quarter of investment, ask what the additional spend will change and assess retention, maintenance cost, support and engineering burden, profitability, and whether customer losses reflect a problem the team can realistically fix. Agree in advance on the results that would justify stopping investment, particularly when internal attachment to a project may distort judgment.
- One organization described prioritizing products that drive revenue without increasing burn, while also favoring products that could help attract investment.
A local 8B model’s evaluation score improved from 0/15 to 14/15 after changing one setting, with the model weights and machine unchanged; the author attributed the original failure to thinking mode consuming the output budget. For AI product evaluations, this is a reminder to inspect runtime settings and output-budget constraints before concluding that the underlying model is inadequate.
- Launch-case framework: For a PM interview scenario asking how to launch in the US and reach $10K within one month, clarify whether the target means revenue or profit, the pricing, and whether US users already exist; then work backward into customer volume, conversion rates, acquisition channels, and explicit assumptions. The candidate’s initial competitor-positioning and acquisition research was a reasonable starting point but did not connect tightly enough to the revenue target.
- Feasibility and cross-functional ownership: One respondent advised acknowledging the ambition of immediate profitability, asking about ramp-up time, defining the steps leading to launch, and agreeing on post-launch follow-up metrics. Another commenter challenged treating first-month revenue as solely a PM responsibility, distinguishing product discovery from marketing distribution, sales capture, and growth optimization.
- Interview dynamics: A respondent considered asking for a few moments to think reasonable and screen sharing understandable given AI misuse concerns. The candidate reported being accused of using AI despite explaining the pause, after which the interview ended without substantive feedback.
Forward-deployed engineers (FDEs) are most valuable when they sit within the product team—not go-to-market—stay closely connected to the roadmap, solve hard customer problems, and feed those learnings back into the product. Customer-specific domain expertise can then be abstracted into a reusable product experience.
The operating test is that FDE solutions should have an expiration date: productize recurring solutions so future customers get them by default; if the same problem keeps requiring an FDE, the product has failed to learn.
AI-era product strategy should not treat the economy as a fixed backlog: Hiten Shah argues that once intelligence and execution become dramatically cheaper, the harder question is what humans choose to do, because AI may automate today’s backlog while expanding the amount of work considered worth doing.
- For AI messaging products, a trust-building interaction is to draft the message, show users exactly what will be sent, and wait for explicit approval before sending; Hiten Shah argues that this small pause makes the product easier to trust.
- Trust is built through many such product details, but committees, roadmaps, and arbitrary ship dates can delay them; Shah’s recommendation is to ship when the experience is ready and continue improving it afterward.
- Evaluate local AI models by assigning them a small, concrete job you already do every day rather than comparing their chat responses with the best cloud model. Repeatedly using the model for that job helps identify whether it delivers a meaningful, durable use case.
- For ambiguous product or strategy decisions, structure the analysis into three buckets: market/revenue opportunity, cost dynamics, and operational capacity; prioritize the largest opportunity before optimizing execution details. Relevant inputs include demand, sponsorship or monetization potential, cannibalization, upfront and ongoing costs, payback timing, staffing, infrastructure, and external constraints.
- Make data-heavy comparisons auditable by stating elimination or prioritization rules before applying them, using the data to narrow choices rather than discuss every datapoint, and ending with a clear recommendation and next steps.
- Make investment trade-offs explicit across the full time horizon: compare upfront capital, time to launch or revenue, annual run-rate economics, and cumulative break-even. The case feedback notes that the intended Atlanta-versus-Nashville crossover was nine years after accounting for Nashville’s two-year head start, illustrating how a short-term choice can differ from the stronger long-term option.
- For products or initiatives dependent on external stakeholders, map customers, employees, sponsors, residents, and government bodies; identify which groups can actually delay execution, engage the highest-risk authorities early, communicate concrete community benefits, and validate local demand to reduce opposition.
- For career development, use deliberate interview practice: pause at each prompt, answer aloud before reviewing the solution, and build repetition through structured drills and feedback.
Repeatedly shipping mediocre work lowers the bar for approving subsequent mediocre work, causing quality debt to compound.
- Senior product leaders are navigating AI-driven uncertainty, shifting expectations, shorter windows, and pivot fatigue; many are seeking roles that let them apply prior experience for roughly 80% of their time while using the remaining 20% to upskill, while some are moving into IC roles to become AI-native builders.
- A practical career-advocate exercise is to list more than 10 people who hired, promoted, supported, or championed you, then reconnect by revisiting the problems solved and pivots navigated together. This both clarifies the strengths you should bring to market and activates a passive pipeline of opportunities. Warm introductions can surface roles before they are public and give candidates a chance to shape the role around their profile; often, a direct contact remembers or recommends you into a role held by someone one degree removed.
- A data/analytics engineer with about two years of experience across two companies is considering a PM opportunity at a small startup after enjoying product work, but is unsure whether to switch now or remain in data for another 2–3 years.
- The transition questions center on entry-level PM demand, whether two years of technical/data experience is sufficient, and how much a data background helps when applying for PM roles.
- A count of 421 open PM postings across 62 company job boards found roadmap in 85.5% of listings; cross-functional, prioritization, and stakeholder management each appeared in more than 55%. None of the four most common requirements was a tool, method, or metric.
- Tool and methodology keywords were much less common: Jira appeared in 1.7% of postings, Figma in 2.4%, Productboard in none, Agile in 8.6%, and Scrum in 2.1%.
- Resume wording may affect matching: experimentation appeared in 24.2% of postings versus A/B testing in 2.1%, despite the post treating them as the same practice. The sample is limited to tech companies using one ATS and excludes agencies and regulated employers where Scrum certifications may be explicitly requested.
r/ProductMgmt comment by u/bookninja717
Accused of using AI during a PM intern interview — looking for honest feedback
I recently interviewed for a Product/PM intern role and wanted to get some honest opinions from people who have experience interviewing for PM roles.
At the beginning of the interview, they asked me to share my screen, which I did. I answered their questions about my introduction, background, and projects.
Towards the end, they gave me a question along the lines of:
“Our company is launching in the US, and within one month we have a goal of making $10K. What would be your approach?”
I found the question challenging because there were many unknowns, so I asked for some time to think before answering.
My initial answer was that I would research competitors in the US, understand how they are positioning themselves, what approaches they are using to acquire customers, and then discuss those findings with the team to decide our approach.
Looking back, I realize my answer wasn’t complete. I should have clarified things such as whether the $10K meant revenue or profit, what the pricing was, whether we already had US users, etc. I also could have broken the $10K goal down into customers, conversion rates, acquisition channels, and so on.
After I answered, the interviewer told me that it seemed like I was using AI.
They then asked me to show my desk. My phone was there because I was using it for the internet hotspot connection, but I genuinely did not use my phone or AI to answer the question.
I explained that I wasn’t using AI and that I had simply taken some time to think because I found the question difficult. They still believed that I had used AI.
At that point, I didn’t want to argue with them. I said something along the lines of:
“I’m sorry if from your perspective it seems like I used AI to get the answer. I understand, and I apologize if I wasted your time in the interview.”
They then ended the interview suddenly without really concluding it or giving any feedback.
Honestly, this really demotivated me.
I’m not posting this to attack the interviewer or company. I genuinely want to understand what I could have done differently.
For people who interview PM candidates:
Is taking time to think before answering a question like this considered suspicious?
Was my approach fundamentally wrong, or was it just incomplete?
How would you approach the $10K-in-one-month US launch question?
If you suspected a candidate was using AI, how would you handle it?
Is this kind of question reasonable for a PM intern interview?
I’m open to criticism of my answer. I want to learn from this and do better in my next interview.
Frankly, I would have sputtered my coffee all over the monitor when I heard they expected more than $1 in the first month.
I do understand the issue. I’ve heard of candidates who simply type into AI and read the result. So sharing your screen is somewhat reasonable.
It’s also reasonable for you to ask for a few moments to think.
You know you weren’t using AI. That should be enough. No need to argue.
- Launch-case framework: For a PM interview scenario asking how to launch in the US and reach $10K within one month, clarify whether the target means revenue or profit, the pricing, and whether US users already exist; then work backward into customer volume, conversion rates, acquisition channels, and explicit assumptions. The candidate’s initial competitor-positioning and acquisition research was a reasonable starting point but did not connect tightly enough to the revenue target.
- Feasibility and cross-functional ownership: One respondent advised acknowledging the ambition of immediate profitability, asking about ramp-up time, defining the steps leading to launch, and agreeing on post-launch follow-up metrics. Another commenter challenged treating first-month revenue as solely a PM responsibility, distinguishing product discovery from marketing distribution, sales capture, and growth optimization.
- Interview dynamics: A respondent considered asking for a few moments to think reasonable and screen sharing understandable given AI misuse concerns. The candidate reported being accused of using AI despite explaining the pause, after which the interview ended without substantive feedback.