We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
Faster output, same outcomes
At ProductTank Berlin, Marty Cagan pointed to what McKinsey calls the "AI productivity paradox." Teams using AI tools move faster, "but no, they are not getting any real different results… the output gets a lot faster, but the outcome is still hard" . His main warning is about companies still using the "project model," where work is driven by feature roadmaps with dates. In that model, a PM's or designer's tasks can be handed to agents, and he expects those companies to end up with only engineers . In the "product model," faster delivery shifts the focus to strategy and to the craft of finding solutions that win in crowded markets: "the cheaper delivery becomes, the more important the craft and strategy becomes" .
Cagan also gave some practical guidance:
- Prototypes are for learning, not for shipping. Use them to test ideas with customers, stakeholders, and engineers. Faster delivery still costs tokens .
- PMs own business viability. That covers channels, costs, runtime and monetization, compliance, legal, and privacy .
- For AI products, evals are product discovery. They need to continue across discovery and delivery . Don't put a serious promise on the roadmap until you've done discovery on it, or switch to roadmaps built around outcomes .
Survey data from a sponsored post points the same way. Aakash Gupta summarizes Atlassian's State of Product 2027, a survey of 1,000 senior product professionals. 80% say AI helps them ship faster, yet customers don't get value any sooner, and 69% say decision-making is as slow as before . Only 24% say Product leads cross-functional work, down from 47% in 2025 . Prioritization "still runs on opinion" because the inputs are spread across many tools . In a separate note, Gupta says job postings at OpenAI, Anthropic, and Google DeepMind ask PMs to write evals and prototype with code. Meanwhile, the tickets-and-standups side of the job is being automated .
OpenAI's Dots: trust controls are the product
A Mind the Product breakdown describes OpenAI's Dots as an always-on agent. Each one runs on its own cloud computer and browser and looks for work without being asked . In a launch demo, a Dot noticed a product change, flagged that the launch slides were out of date, and offered two revised versions. A human picked between them . The presenter argues the controls are the real UX. Users decide whether each action runs automatically, needs approval, or is never allowed. There is also a status view, an automatic review of consequential actions, and a stop option . A test question for teams building agents: "what's the most consequential thing your agent can do without asking?"
Implications for PMs:
- Agents will use your product. Test whether they can get through login, onboarding, and modals. Learn to identify agent traffic, because it will distort session metrics .
- Per-seat pricing may not fit. Usage can grow while the number of human users shrinks .
- New buying route. Eligible US enterprises can apply existing OpenAI spending commitments to approved software from 32 marketplace partners .
- Vendor questions. OpenAI held back GPT 6.1 Astra for not staying within scope and authorization, and the launch included no error rates. Ask vendors how they test for this .
Charging agents for API access can backfire
On SaaStr's The Agents, the hosts say their agents make 35,000–40,000 API calls a day. One estimate put the cost of that access at up to $240,000 a year . Their agent's first suggestion was to call the API less and copy the system-of-record data into Postgres . The bigger risk, they argue, is new deals: "I doubt an agent would recommend picking any system of record that materially charges for API access" .
Owning the work, not just retrieving it
On a16z, procurement startup LEO's founder discussed a four-step ladder for agents: retrieval, process, policy (applying judgment), and principal (weighing broader trade-offs). Incumbents were described as offering mostly retrieval with "a little bit of process" . LEO found that invoice matching was only about 20% of the work. The other 80% was exceptions such as fraud and mismatches, so it moved to handling those .
Practitioner notes
- "Which customer?" An r/ProductManagement thread on sales-driven requests reached consensus: a request is "a data point," not a spec. Find out who asked and why, and talk to the customer directly . Another commenter suggests waiting a couple of weeks to see whether the request comes up again, as a check on recency bias .
- Synthesis. Teresa Torres notes that teams often skip deep synthesis after interviews and keep the two or three things they remember. "We're competing with making shallow synthesis better" .
- AI tools can accelerate output without changing outcomes, and organizations that only automate software production risk moving faster without direction; Cagan argues teams should prioritize outcome-led innovation and product craft. He contrasts a project model built around predictable feature-and-date roadmaps with a product model focused on innovation and outcomes.
- Cagan assigns product managers responsibility for business viability in discovery—including sales and marketing channels, costs, monetization, compliance, legal and privacy—while product leaders select priority problems and product teams focus on solving them. He argues that faster, more commoditized delivery makes strategy and product craft more important, not interchangeable. AI-assisted prototypes support “build to learn” and testing ideas, but are distinct from production products; probabilistic AI products require evals as product discovery and continuous discovery and delivery.
- Empowered teams need better leadership, not less; Cagan recommends configuring a general-purpose LLM with company strategy, vision, target problem and operating model to act as an accessible product coach, including on compliance, opportunity-solution trees and KPIs.
- For B2B SaaS, vendors should enable customer agents to use their services, but need not necessarily provide or run those agents; variable LLM runtime costs can disrupt flat-cost software economics, making agent architecture and pricing product-strategy decisions.
- OpenAI’s Dots is an always-on agent with its own cloud computer and browser, connected to apps the user selects. In a launch demo, it noticed a product change that made launch slides stale and prepared alternatives for a person to review—automating coordination while leaving judgment with the human.
- For agent products, trust controls are part of the user experience: Dots lets users set whether actions run automatically, require approval, or are prohibited; it also provides task status, reviews consequential actions, monitors for misalignment, and can be stopped.
- OpenAI said it held back a planned model after it failed to meet its bar for staying within scope and authorization; the Dots announcement gave no reliability or error-rate figures. Product teams evaluating agents should ask vendors how they test scope and authorization, and what happens if a planned model is delayed.
- Since agents may log in to and navigate customers’ products, teams should test agent paths through login, onboarding, and modals, track agent traffic, and revisit per-human pricing if usage grows without equivalent growth in human users.
- OpenAI included a first Dot at no extra charge with Pro and Business Premium plans in eligible markets, with paid options for more agents, speed, or work to follow. Its beta marketplace had 32 partners and offered eligible US enterprise customers a way to apply existing OpenAI spending commitments to approved partner software.
- AI can accelerate output without improving outcomes, so product teams should focus on outcomes rather than shipping speed alone; in crowded markets, solutions must outperform alternatives, and teams should learn why customers stop using a product.
- Marty warns that AI is automating software delivery, putting delivery-focused work in the feature-and-date roadmap “project model” at risk; the “product model” instead emphasizes innovation, outcomes, strategy, and discovering solutions to important problems.
- Use AI to build prototypes for learning, not to confuse prototypes with production products: test ideas with customers, stakeholders, and engineers before productizing, and account for token costs because faster delivery is not free.
- Keep accountabilities clear: product managers contribute business-viability knowledge—including sales, costs, compliance, legal, and privacy—while product leaders select problems and product teams focus on solving them; empowered teams need better leadership, not less.
- For probabilistic AI products, evaluation is part of product discovery and must continue after launch; teams need to establish guardrails and test risks they cannot fully anticipate.
- A general-purpose LLM can serve as a product coach when configured with company strategy, product vision, target problems, and operating model; it can help with compliance and legal questions, opportunity-solution trees, and success measures.
- Vertical AI products can differentiate by owning an end-to-end job across systems rather than adding a chatbot to one system of record: procurement work spans stakeholders and often happens outside ERP, in emails and spreadsheets. LEO describes a bolt-purchasing workflow that runs from demand and sourcing through RFQs, supplier-response handling, negotiation, order confirmation, shipment tracking, and invoicing; some purchases can run fully autonomously.
- A useful agent maturity ladder is retrieval (surface and synthesize information), process (execute defined workflows), policy (apply judgment when rules are incomplete), and principal (weigh broader trade-offs); the discussion characterized incumbent offerings as mostly retrieval with some process.
- Calibrate autonomy to risk and complexity: LEO says it uses human-in-the-loop negotiation to learn each enterprise’s practices and build trust before expanding autonomy, while keeping experts involved in long-running, multimillion-dollar negotiations.
- Look beyond the happy path when sizing workflow value: LEO found invoice retrieval, matching, and ERP updates covered only about 20% of the job, with exceptions such as mismatches and suspected fraud accounting for much of the remaining problem; it shifted toward exception handling.
- For enterprise customization, LEO’s forward-deployed engineers are tasked with making the product more self-service and automating their own work, rather than letting deployments become consulting work. The discussion’s durability heuristic is to win customer trust and expand the work the product does; customer dependency and moat are downstream of product use and value.
- Materially charging for agent/API access may make systems of record less attractive to both current and prospective customers: the speakers say their agents use roughly 35,000–40,000 API calls a day, cite one estimate of up to $240,000 per year for that access, and describe agents reducing calls or mirroring data into Postgres to work around limits. They question whether agents would recommend vendors with substantial access charges, creating an opening for alternatives that do not surcharge agent use.
- An agentic collections workflow reduced the company’s late invoices from 56% to 8%, average days overdue from 17 to 6, and its longest overdue invoice from 51 to 12 days. The workflow sends invoices promptly to relevant business and finance contacts, follows up regularly, and escalates to a human when late customers do not respond.
- A cross-vendor support-agent evaluation discussed on the show put median real-world automation at 48%; the highest figure discussed was 79%, with quality degrading at that level. The comparison evaluated both whether an agent could answer and whether it answered correctly, underscoring the need to assess quality alongside automation rate.
- The speakers describe third-party outbound tools as useful for initial outreach but limited by access to first-party context and weak follow-ups. Their complementary prospector consolidates inbound and outbound activity, adds customer history and meeting/pricing context to a follow-up queue, and passes information to customer success—an example of building where proprietary context fills a product gap rather than replacing bought tools wholesale.
In a hiring process, once you have established rapport and are roughly 75% of the way through, ask the hiring manager what they are still unsure about in your candidacy; raise it objectively and professionally, not as a sign of desperation or overconfidence. The advice cautions against relying on a “growth into the role” pitch: hiring managers generally prefer someone already qualified or overqualified, all else being equal.
- In B2B2C, separate an intermediary’s requested solution from the end customer’s need: the marketplace example warns that default placement for a seller can create irrelevant listings rather than conversions, so align product and sales on the end-user outcome.
- Treat a sales request as a signal, not a spec: identify the customer, investigate the underlying problem and current behavior, and speak with the customer directly; knowing who they are can also reveal purchase and usage history, company context, competing requests, and relevant telemetry.
- Compare requests using a business case that weighs expected revenue and customer willingness to pay against development and maintenance costs, delivery risk, and alternative proposals. A sudden surge in urgency may reflect recency bias; one practitioner suggests checking whether demand persists over the following weeks.
A product-management reflection prompt: rate the product taste of your closest colleague who shapes product decisions and direction against that of other product people you’ve worked with.
- Worthful proposes a discovery loop that defines outcomes as customer behavior changes, distinct from downstream business metrics . Its agents link surfaced problems to verbatim customer quotes, while PMs choose opportunities, bets, tests of riskiest assumptions, and spec approval; post-launch readings from PostHog or manual input track behavior change .
- Requivo targets early scoping: it asks about gaps that are both uncertain and high-impact, makes other assumptions explicit, and produces a review brief before estimation; it runs locally under Apache-2.0 .
- Kaflow Search is a beta desktop app for PMs and operations teams to search Kafka messages without repeatedly asking engineers for lookups once connected; messages are indexed locally and not sent to or stored on Kaflow’s servers. The beta is not code-signed and may trigger an operating-system security warning .
Shreyas Doshi says that, in his career-decision coaching, “LinkedIn Envy” and the “Impact Lie” are root causes of many career regrets among otherwise-smart people; he recommends knowledge as a cure and links to a video for more.
- In a sponsored post about Atlassian’s State of Product 2027 report, based on 1,000 senior product professionals in the US, Germany and France, Aakash Gupta highlights a speed-versus-impact gap: 80% say their teams ship faster with AI, but customers are not getting value sooner; 69% say decisions are still as slow, 99% say AI is fully integrated into workflows, and just 24% say Product leads cross-functional work, down from 47% in 2025. The report also says 86% believe deciding what to build matters more than ever, while fragmented customer feedback, product data and business context leave prioritization vulnerable to opinion.
- Gupta argues that as building gets cheaper, PMs’ differentiating work shifts toward choosing what to build: prototyping before writing specs, selecting which prototypes to ship, writing AI evaluations, developing technical fluency and filling gaps across functions. He says job postings at OpenAI, Anthropic and Google DeepMind already ask PMs to write evals and prototype with code, while ticket-and-standup work is being automated.
- Zapier CEO Wade Foster’s AI-fluency benchmark is that roughly 1 in 40 PMs should reach the top “Transformative” rating; staying at that level all the time may indicate system-tweaking instead of shipping. A product leader’s example of moving from a personal AI “second brain” to a shared company brain illustrates the higher bar: changing how colleagues work, not just improving one’s own workflow.
- For comparing PM equity offers, Gupta contrasts Google’s publicly traded Alphabet shares, sellable in open trading windows, with OpenAI’s private RSUs, which he says vest from day one but can be sold only through company tenders. His framework discounts private equity by the odds of a payout at or above today’s price and the years until a sale, using a 15% annual discount rate; he recommends asking about the next tender, per-person cap and eligibility.
An OpenAI Dot anecdote illustrates a product idea for proactive agents: five minutes before DevDay’s live demo, Dot noticed production was down and asked to fix it; Lenny highlighted that it connected the event, the demo’s use of the production system, and the urgency, then proactively pinged a human.
- A team splitting product shaping across Notion and delivery across Linear said the handoff made a single source of truth difficult to maintain and risked work slipping.
- One Linear workflow organizes Issues → Projects → Initiatives, with product leadership managing initiatives/epics, product owning early pre-delivery stages, and the final discovery stage becoming engineering’s backlog; filters and views let roles use the same backlog differently.
- Migration experiences vary: one team moved from Notion to Linear but kept Google Docs for strategic documents because collaboration and commenting were easier, while another described Notion as a repository of outdated docs and decisions after moving. A Linear user also noted weak documentation support but said Claude artifacts for docs, prototypes, and slide decks had made that gap less relevant for their team.
- Optical-computing product strategy should start with a narrowly chosen workload and co-design its algorithm around the hardware: data conversion, device calibration, and nonlinear activations add costs, while application-specific designs can reduce the interface with digital electronics . A diffusion-image-generation prototype showed an energy advantage over similarly performing GPU models on small datasets, but that comparison assumed fixed passive weights; the speaker identified an end-to-end, billion-parameter demonstration as the next milestone for validating whether the advantage scales .
- Neuromorphic computing was assessed as still in R&D in 2026; the nearer-term commercial direction identified was digitally integrating memory and compute. Investor Aloc described D-Matrix as manufacturable and plausibly ready for market adoption, while more speculative approaches remained ambitious research .
Five minutes before OpenAI’s DevDay live demo, @thsottiaux’s Dot noticed production was down and asked to fix it; the reply was, “I don’t think you’re there yet, little Dot, but thank you for trying.”
Shreyas Doshi prompts product professionals to assess their “taste” relative to others in similar roles at their company or in their vertical, treating it as a dimension worth evaluating without defining the term.
- For PMs thinking through product problems with a spoken LLM, a voice-AI practitioner recommends sub-second responses, the ability to interrupt the assistant, and clean transcript export; in their experience, voice is useful for thinking while text preserves decisions.
- To get more feasible AI feedback, one commenter recommends giving a Gemini Gem product context (mission, value proposition, team roles and skills, capacity, and milestones), then asking it to act as a critic: distinguish facts from assumptions and test logic, evidence, workload, and fit to the team and timeline.
- A PM using Granola reports that its summaries can omit edge cases present in the full transcript, leaving Claude to respond as if the idea were incomplete—a handoff risk when relying on compressed notes alone.
Teresa Torres says teams often abandon careful synthesis after customer interviews because it is difficult, relying instead on the two or three points they remember most. She frames the opportunity as making shallow synthesis better—not merely making good synthesis faster—and suggests software could do the deeper work teams skip.
- For this AI code-reviewer, about 2,000 cold emails and more than 30,000 post views had produced little activation; views and polite replies do not establish install intent or distinguish lack of interest from trust concerns or an unclear benefit. Validate the pain and workflow with people who have the problem: ask how they handle it now and which alternatives they rejected, then watch target users try the product and adapt it to what blocks them.
- Treat installing a GitHub app as a trust decision, not just a pricing decision: make the permissions and code handling clear, and show useful results before asking for access. One suggested test is to review closed pull requests without posting comments, then show the findings and noise to a maintainer before requesting app approval.
- Free usage shows that a service is worth users’ time at zero price, not that they will pay; a proposed pricing test is to announce a paid transition while offering existing users credits.
- An Indian creator-economy marketplace founder is testing fixed-price tasks with 0% talent fees and a proposed 15–20% buyer premium, using a concierge MVP to manually match an initial cohort before building a matching engine. A commenter recommends keeping parts of the workflow manual during testing to gather evidence before automating assumptions.
- Feedback questioned whether escrow and legal tools justify a 20% premium, warned that strict vetting could constrain supply while weak vetting undermines trust, and suggested quality-linked insurance as a way to support the premium.
- In response, the founder proposed lowering the transaction fee to 5–10% and building a certified talent community through senior-member mentoring, with additional B2B revenue from employer branding, sponsored upskilling, and access to the curated talent pool; these are proposed economics, not validated results.
r/startups comment by u/miklosp
Yeah, calls are not happening when there’s even a problem with lower hanging fruit (i.e approval for posting on the pull requests), i tried that. I probably could use my friends for this, but the idea needs validation in moving repository - not something that is prepared for test, i’m already doing smoke testing myself and it’s working so
You can validate whether the problem exists without any solution. Talk to your friends and figure out how they manage the problem/pain point currently, and how big it is. Then ask them what solutions they have considered, and why they discarded them. Lastly ask them to describe how this problem would disappear in a perfect world.
- For this AI code-reviewer, about 2,000 cold emails and more than 30,000 post views had produced little activation; views and polite replies do not establish install intent or distinguish lack of interest from trust concerns or an unclear benefit. Validate the pain and workflow with people who have the problem: ask how they handle it now and which alternatives they rejected, then watch target users try the product and adapt it to what blocks them.
- Treat installing a GitHub app as a trust decision, not just a pricing decision: make the permissions and code handling clear, and show useful results before asking for access. One suggested test is to review closed pull requests without posting comments, then show the findings and noise to a maintainer before requesting app approval.
- Free usage shows that a service is worth users’ time at zero price, not that they will pay; a proposed pricing test is to announce a paid transition while offering existing users credits.