ZeroNoise Logo zeronoise
Post
AI makes teams ship faster, but decisions and outcomes haven't caught up
•
4 min read
• 198 docs
Marty Cagan and Atlassian survey data both point to a gap between AI-driven output and real outcomes. OpenAI's Dots raises design questions for agent-facing products, and SaaStr warns that charging agents for API access can backfire.

Faster output, same outcomes

At ProductTank Berlin, Marty Cagan pointed to what McKinsey calls the "AI productivity paradox." Teams using AI tools move faster, "but no, they are not getting any real different results… the output gets a lot faster, but the outcome is still hard" . His main warning is about companies still using the "project model," where work is driven by feature roadmaps with dates. In that model, a PM's or designer's tasks can be handed to agents, and he expects those companies to end up with only engineers . In the "product model," faster delivery shifts the focus to strategy and to the craft of finding solutions that win in crowded markets: "the cheaper delivery becomes, the more important the craft and strategy becomes" .

Cagan also gave some practical guidance:

  • Prototypes are for learning, not for shipping. Use them to test ideas with customers, stakeholders, and engineers. Faster delivery still costs tokens .
  • PMs own business viability. That covers channels, costs, runtime and monetization, compliance, legal, and privacy .
  • For AI products, evals are product discovery. They need to continue across discovery and delivery . Don't put a serious promise on the roadmap until you've done discovery on it, or switch to roadmaps built around outcomes .

Survey data from a sponsored post points the same way. Aakash Gupta summarizes Atlassian's State of Product 2027, a survey of 1,000 senior product professionals. 80% say AI helps them ship faster, yet customers don't get value any sooner, and 69% say decision-making is as slow as before . Only 24% say Product leads cross-functional work, down from 47% in 2025 . Prioritization "still runs on opinion" because the inputs are spread across many tools . In a separate note, Gupta says job postings at OpenAI, Anthropic, and Google DeepMind ask PMs to write evals and prototype with code. Meanwhile, the tickets-and-standups side of the job is being automated .

OpenAI's Dots: trust controls are the product

A Mind the Product breakdown describes OpenAI's Dots as an always-on agent. Each one runs on its own cloud computer and browser and looks for work without being asked . In a launch demo, a Dot noticed a product change, flagged that the launch slides were out of date, and offered two revised versions. A human picked between them . The presenter argues the controls are the real UX. Users decide whether each action runs automatically, needs approval, or is never allowed. There is also a status view, an automatic review of consequential actions, and a stop option . A test question for teams building agents: "what's the most consequential thing your agent can do without asking?"

Implications for PMs:

  • Agents will use your product. Test whether they can get through login, onboarding, and modals. Learn to identify agent traffic, because it will distort session metrics .
  • Per-seat pricing may not fit. Usage can grow while the number of human users shrinks .
  • New buying route. Eligible US enterprises can apply existing OpenAI spending commitments to approved software from 32 marketplace partners .
  • Vendor questions. OpenAI held back GPT 6.1 Astra for not staying within scope and authorization, and the launch included no error rates. Ask vendors how they test for this .

Charging agents for API access can backfire

On SaaStr's The Agents, the hosts say their agents make 35,000–40,000 API calls a day. One estimate put the cost of that access at up to $240,000 a year . Their agent's first suggestion was to call the API less and copy the system-of-record data into Postgres . The bigger risk, they argue, is new deals: "I doubt an agent would recommend picking any system of record that materially charges for API access" .

Owning the work, not just retrieving it

On a16z, procurement startup LEO's founder discussed a four-step ladder for agents: retrieval, process, policy (applying judgment), and principal (weighing broader trade-offs). Incumbents were described as offering mostly retrieval with "a little bit of process" . LEO found that invoice matching was only about 20% of the work. The other 80% was exceptions such as fraud and mismatches, so it moved to handling those .

Practitioner notes

  • "Which customer?" An r/ProductManagement thread on sales-driven requests reached consensus: a request is "a data point," not a spec. Find out who asked and why, and talk to the customer directly . Another commenter suggests waiting a couple of weeks to see whether the request comes up again, as a check on recency bias .
  • Synthesis. Teresa Torres notes that teams often skip deep synthesis after interviews and keep the two or three things they remember. "We're competing with making shallow synthesis better" .
AI makes teams ship faster, but decisions and outcomes haven't caught up
Marty Cagan
Profile
  • AI tools can accelerate output without changing outcomes, and organizations that only automate software production risk moving faster without direction; Cagan argues teams should prioritize outcome-led innovation and product craft. He contrasts a project model built around predictable feature-and-date roadmaps with a product model focused on innovation and outcomes.
  • Cagan assigns product managers responsibility for business viability in discovery—including sales and marketing channels, costs, monetization, compliance, legal and privacy—while product leaders select priority problems and product teams focus on solving them. He argues that faster, more commoditized delivery makes strategy and product craft more important, not interchangeable. AI-assisted prototypes support “build to learn” and testing ideas, but are distinct from production products; probabilistic AI products require evals as product discovery and continuous discovery and delivery.
  • Empowered teams need better leadership, not less; Cagan recommends configuring a general-purpose LLM with company strategy, vision, target problem and operating model to act as an accessible product coach, including on compliance, opportunity-solution trees and KPIs.
  • For B2B SaaS, vendors should enable customer agents to use their services, but need not necessarily provide or run those agents; variable LLM runtime costs can disrupt flat-cost software economics, making agent architecture and pricing product-strategy decisions.
Fireside Chat with Marty Cagan and Elias Lieberich - ProductTank Berlin
Mind the Product
  • OpenAI’s Dots is an always-on agent with its own cloud computer and browser, connected to apps the user selects. In a launch demo, it noticed a product change that made launch slides stale and prepared alternatives for a person to review—automating coordination while leaving judgment with the human.
  • For agent products, trust controls are part of the user experience: Dots lets users set whether actions run automatically, require approval, or are prohibited; it also provides task status, reviews consequential actions, monitors for misalignment, and can be stopped.
  • OpenAI said it held back a planned model after it failed to meet its bar for staying within scope and authorization; the Dots announcement gave no reliability or error-rate figures. Product teams evaluating agents should ask vendors how they test scope and authorization, and what happens if a planned model is delayed.
  • Since agents may log in to and navigate customers’ products, teams should test agent paths through login, onboarding, and modals, track agent traffic, and revisit per-human pricing if usage grows without equivalent growth in human users.
  • OpenAI included a first Dot at no extra charge with Pro and Business Premium plans in eligible markets, with paid options for more agents, speed, or work to follow. Its beta marketplace had 32 partners and offered eligible US enterprise customers a way to apply existing OpenAI spending commitments to approved partner software.
Everything You Need To Know About OpenAI's Dots in 17 minutes
Mind the Product
  • AI can accelerate output without improving outcomes, so product teams should focus on outcomes rather than shipping speed alone; in crowded markets, solutions must outperform alternatives, and teams should learn why customers stop using a product.
  • Marty warns that AI is automating software delivery, putting delivery-focused work in the feature-and-date roadmap “project model” at risk; the “product model” instead emphasizes innovation, outcomes, strategy, and discovering solutions to important problems.
  • Use AI to build prototypes for learning, not to confuse prototypes with production products: test ideas with customers, stakeholders, and engineers before productizing, and account for token costs because faster delivery is not free.
  • Keep accountabilities clear: product managers contribute business-viability knowledge—including sales, costs, compliance, legal, and privacy—while product leaders select problems and product teams focus on solving them; empowered teams need better leadership, not less.
  • For probabilistic AI products, evaluation is part of product discovery and must continue after launch; teams need to establish guardrails and test risks they cannot fully anticipate.
  • A general-purpose LLM can serve as a product coach when configured with company strategy, product vision, target problems, and operating model; it can help with compliance and legal questions, opportunity-solution trees, and success measures.
Fireside Chat with Marty Cagan and Elias Lieberich - ProductTank Berlin
  • Vertical AI products can differentiate by owning an end-to-end job across systems rather than adding a chatbot to one system of record: procurement work spans stakeholders and often happens outside ERP, in emails and spreadsheets. LEO describes a bolt-purchasing workflow that runs from demand and sourcing through RFQs, supplier-response handling, negotiation, order confirmation, shipment tracking, and invoicing; some purchases can run fully autonomously.
  • A useful agent maturity ladder is retrieval (surface and synthesize information), process (execute defined workflows), policy (apply judgment when rules are incomplete), and principal (weigh broader trade-offs); the discussion characterized incumbent offerings as mostly retrieval with some process.
  • Calibrate autonomy to risk and complexity: LEO says it uses human-in-the-loop negotiation to learn each enterprise’s practices and build trust before expanding autonomy, while keeping experts involved in long-running, multimillion-dollar negotiations.
  • Look beyond the happy path when sizing workflow value: LEO found invoice retrieval, matching, and ERP updates covered only about 20% of the job, with exceptions such as mismatches and suspected fraud accounting for much of the remaining problem; it shifted toward exception handling.
  • For enterprise customization, LEO’s forward-deployed engineers are tasked with making the product more self-service and automating their own work, rather than letting deployments become consulting work. The discussion’s durability heuristic is to win customer trust and expand the work the product does; customer dependency and moat are downstream of product use and value.
Why AI Is Reinventing How Businesses Buy Everything
SaaStr AI
  • Materially charging for agent/API access may make systems of record less attractive to both current and prospective customers: the speakers say their agents use roughly 35,000–40,000 API calls a day, cite one estimate of up to $240,000 per year for that access, and describe agents reducing calls or mirroring data into Postgres to work around limits. They question whether agents would recommend vendors with substantial access charges, creating an opening for alternatives that do not surcharge agent use.
  • An agentic collections workflow reduced the company’s late invoices from 56% to 8%, average days overdue from 17 to 6, and its longest overdue invoice from 51 to 12 days. The workflow sends invoices promptly to relevant business and finance contacts, follows up regularly, and escalates to a human when late customers do not respond.
  • A cross-vendor support-agent evaluation discussed on the show put median real-world automation at 48%; the highest figure discussed was 79%, with quality degrading at that level. The comparison evaluated both whether an agent could answer and whether it answered correctly, underscoring the need to assess quality alongside automation rate.
  • The speakers describe third-party outbound tools as useful for initial outreach but limited by access to first-party context and weak follow-ups. Their complementary prospector consolidates inbound and outbound activity, adds customer history and meeting/pricing context to a follow-up queue, and passes information to customer success—an example of building where proprietary context fills a product gap rather than replacing bought tools wholesale.
Agent Pricing is Chaos. Here's What We're Seeing From the Buyer Side | The Agents #015
Shreyas Doshi
Profile

In a hiring process, once you have established rapport and are roughly 75% of the way through, ask the hiring manager what they are still unsure about in your candidacy; raise it objectively and professionally, not as a sign of desperation or overconfidence. The advice cautions against relying on a “growth into the role” pitch: hiring managers generally prefer someone already qualified or overqualified, all else being equal.

Ask What They're Still Unsure About
Product Management
  • In B2B2C, separate an intermediary’s requested solution from the end customer’s need: the marketplace example warns that default placement for a seller can create irrelevant listings rather than conversions, so align product and sales on the end-user outcome.
  • Treat a sales request as a signal, not a spec: identify the customer, investigate the underlying problem and current behavior, and speak with the customer directly; knowing who they are can also reveal purchase and usage history, company context, competing requests, and relevant telemetry.
  • Compare requests using a business case that weighs expected revenue and customer willingness to pay against development and maintenance costs, delivery risk, and alternative proposals. A sudden surge in urgency may reflect recency bias; one practitioner suggests checking whether demand persists over the following weeks.
When sales says “the customer wants it”, do you ask which customer? I would never, ever take a request from anyone as a spec. It's a data point. And yes, if a customer asks for something it's highly releva… Hell yes. Not only that, but I need to talk to the customer. It's not about credibility. You are a PM, they are a sales team. Your job is… I work up a business case with them. I say great idea. Let’s figure out exactly what they need. How much more revenue will this drive? Do… My biggest problem with feature requests coming from sales is recency bias. Something happened recently, so it suddenly feels extremely i…
Shreyas Doshi

A product-management reflection prompt: rate the product taste of your closest colleague who shapes product decisions and direction against that of other product people you’ve worked with.

If you work on products: How would you rate the Taste of your closest colleague who also shapes product decisions & direction? (compa…
Product Management
  • Worthful proposes a discovery loop that defines outcomes as customer behavior changes, distinct from downstream business metrics . Its agents link surfaced problems to verbatim customer quotes, while PMs choose opportunities, bets, tests of riskiest assumptions, and spec approval; post-launch readings from PostHog or manual input track behavior change .
  • Requivo targets early scoping: it asks about gaps that are both uncertain and high-impact, makes other assumptions explicit, and produces a review brief before estimation; it runs locally under Apache-2.0 .
  • Kaflow Search is a beta desktop app for PMs and operations teams to search Kafka messages without repeatedly asking engineers for lookups once connected; messages are indexed locally and not sent to or stored on Kaflow’s servers. The beta is not code-signed and may trigger an operating-system security warning .
Hi folks, I'm working on Worthful: an IDE for product teams (Integrated Discovery Environment) [https://worthful.io](https://worthful.io)… I'm building Requivo, an open-source tool for the start of scoping. Five years of B2B product work and the same pattern every time: a cli… **Search Kafka events yourself, without asking an engineer every time.** Finding a specific order or user event among millions of Kafka m…
Shreyas Doshi

Shreyas Doshi says that, in his career-decision coaching, “LinkedIn Envy” and the “Impact Lie” are root causes of many career regrets among otherwise-smart people; he recommends knowledge as a cure and links to a video for more.

From coaching 100s of people over the years on career decisions, LinkedIn Envy and the Impact Lie are the root cause of many career regre…
Aakash Gupta
  • In a sponsored post about Atlassian’s State of Product 2027 report, based on 1,000 senior product professionals in the US, Germany and France, Aakash Gupta highlights a speed-versus-impact gap: 80% say their teams ship faster with AI, but customers are not getting value sooner; 69% say decisions are still as slow, 99% say AI is fully integrated into workflows, and just 24% say Product leads cross-functional work, down from 47% in 2025. The report also says 86% believe deciding what to build matters more than ever, while fragmented customer feedback, product data and business context leave prioritization vulnerable to opinion.
  • Gupta argues that as building gets cheaper, PMs’ differentiating work shifts toward choosing what to build: prototyping before writing specs, selecting which prototypes to ship, writing AI evaluations, developing technical fluency and filling gaps across functions. He says job postings at OpenAI, Anthropic and Google DeepMind already ask PMs to write evals and prototype with code, while ticket-and-standup work is being automated.
  • Zapier CEO Wade Foster’s AI-fluency benchmark is that roughly 1 in 40 PMs should reach the top “Transformative” rating; staying at that level all the time may indicate system-tweaking instead of shipping. A product leader’s example of moving from a personal AI “second brain” to a shared company brain illustrates the higher bar: changing how colleagues work, not just improving one’s own workflow.
  • For comparing PM equity offers, Gupta contrasts Google’s publicly traded Alphabet shares, sellable in open trading windows, with OpenAI’s private RSUs, which he says vest from day one but can be sold only through company tenders. His framework discounts private equity by the odds of a payout at or above today’s price and the years until a sale, using a 15% annual discount rate; he recommends asking about the next tender, per-person cap and eligibility.
Atlassian's State of Product 2027 report is out, and one finding explains most of what I've been hearing from product leaders this year. … I'm biased (I coach PMs for a living) but Andrew Chen is right. And honestly I think he's underselling it. His whole argument is one line… Zapier's CEO thinks about 1 PM in 40 should hold the top AI fluency rating. Most people read the rubric and assume Transformative is the … OpenAI and Google can both offer you $400K of stock. Only one of them lets you sell it next quarter. Google pays in Alphabet shares. New …
Lenny Rachitsky

An OpenAI Dot anecdote illustrates a product idea for proactive agents: five minutes before DevDay’s live demo, Dot noticed production was down and asked to fix it; Lenny highlighted that it connected the event, the demo’s use of the production system, and the urgency, then proactively pinged a human.

Wild story: Five minutes before OpenAI's DevDay live demo, [@thsottiaux](https://x.com/thsottiaux)'s Dot noticed production was down and … "It understood there's a pretty important thing happening, called called DevDay. There's this production system. There's a live demo. It'…
Product Management
  • A team splitting product shaping across Notion and delivery across Linear said the handoff made a single source of truth difficult to maintain and risked work slipping.
  • One Linear workflow organizes Issues → Projects → Initiatives, with product leadership managing initiatives/epics, product owning early pre-delivery stages, and the final discovery stage becoming engineering’s backlog; filters and views let roles use the same backlog differently.
  • Migration experiences vary: one team moved from Notion to Linear but kept Google Docs for strategic documents because collaboration and commenting were easier, while another described Notion as a repository of outdated docs and decisions after moving. A Linear user also noted weak documentation support but said Claude artifacts for docs, prototypes, and slide decks had made that gap less relevant for their team.
Anyone using Linear for product work? No need to complicate things. Linear hierarchy goes Issues - Projects - Initiatives. In my company, product leadership manages Initiative… We moved off Notion to Linear a year ago and I would never go back. We started using Google Docs for high level strategic documents becau… We have a smaller org so it’s not as bad. Plus, attribution between docs and linear plus agent integrations make finding things pretty qu… We use Linear exclusively (switched from Jira) and while it lacks some things like easy documentation, the rapid improvement of claude ar…
Y Combinator
  • Optical-computing product strategy should start with a narrowly chosen workload and co-design its algorithm around the hardware: data conversion, device calibration, and nonlinear activations add costs, while application-specific designs can reduce the interface with digital electronics . A diffusion-image-generation prototype showed an energy advantage over similarly performing GPU models on small datasets, but that comparison assumed fixed passive weights; the speaker identified an end-to-end, billion-parameter demonstration as the next milestone for validating whether the advantage scales .
  • Neuromorphic computing was assessed as still in R&D in 2026; the nearer-term commercial direction identified was digitally integrating memory and compute. Investor Aloc described D-Matrix as manufacturable and plausibly ready for market adoption, while more speculative approaches remained ambitious research .
What If We Stopped Using GPUs? | YC Paper Club
Lenny Rachitsky

Five minutes before OpenAI’s DevDay live demo, @thsottiaux’s Dot noticed production was down and asked to fix it; the reply was, “I don’t think you’re there yet, little Dot, but thank you for trying.”

Wild story: Five minutes before OpenAI's DevDay live demo, [@thsottiaux](https://x.com/thsottiaux)'s Dot noticed production was down and …
Shreyas Doshi

Shreyas Doshi prompts product professionals to assess their “taste” relative to others in similar roles at their company or in their vertical, treating it as a dimension worth evaluating without defining the term.

If you work on products: How would you rate your Taste? (compared to others in similar roles in your company or your vertical)
Product Management
  • For PMs thinking through product problems with a spoken LLM, a voice-AI practitioner recommends sub-second responses, the ability to interrupt the assistant, and clean transcript export; in their experience, voice is useful for thinking while text preserves decisions.
  • To get more feasible AI feedback, one commenter recommends giving a Gemini Gem product context (mission, value proposition, team roles and skills, capacity, and milestones), then asking it to act as a critic: distinguish facts from assumptions and test logic, evidence, workload, and fit to the team and timeline.
  • A PM using Granola reports that its summaries can omit edge cases present in the full transcript, leaving Claude to respond as if the idea were incomplete—a handoff risk when relying on compressed notes alone.
I build voice AI for a living and I've tried most of the talk-to-an-LLM setups. honest review: latency is everything. if there's a 2-3 se… AI can easily produce a lot more options than you can use, or a super complex plan that your team cannot follow. It happens when the requ… I use granola for talking my ideas through and it does a solid job of sorting out my thoughts and then capturing the salient points. Prob…
Teresa Torres

Teresa Torres says teams often abandon careful synthesis after customer interviews because it is difficult, relying instead on the two or three points they remember most. She frames the opportunity as making shallow synthesis better—not merely making good synthesis faster—and suggests software could do the deeper work teams skip.

Teresa Torres wants to believe every product team does careful synthesis after each customer interview. The reality? It's hard, so teams …
The community for ventures designed to scale rapidly | Read our rules before posting ❤️
  • For this AI code-reviewer, about 2,000 cold emails and more than 30,000 post views had produced little activation; views and polite replies do not establish install intent or distinguish lack of interest from trust concerns or an unclear benefit. Validate the pain and workflow with people who have the problem: ask how they handle it now and which alternatives they rejected, then watch target users try the product and adapt it to what blocks them.
  • Treat installing a GitHub app as a trust decision, not just a pricing decision: make the permissions and code handling clear, and show useful results before asking for access. One suggested test is to review closed pull requests without posting comments, then show the findings and noise to a maintainer before requesting app approval.
  • Free usage shows that a service is worth users’ time at zero price, not that they will pay; a proposed pricing test is to announce a paid transition while offering existing users credits.
How do you figure out distribution and trust in a market with huge players? - i will not promote I wouldn't make everything free yet. You already offer all features free to open source projects, so price isn't the only possible barrie… You can validate whether the problem exists without any solution. Talk to your friends and figure out how they manage the problem/pain po… From how you describe it, I don't think the product is really validated yet. Nothing wrong with that, but ads and cold email are for spee… Your one installation won't tell you much if that repository rarely opens pull requests. Pick a smaller repo that merged several PRs this… If you've got users at the low low cost of free you've just validated that the service is worth their time when it's free. That's it, it'…
Product Management
  • An Indian creator-economy marketplace founder is testing fixed-price tasks with 0% talent fees and a proposed 15–20% buyer premium, using a concierge MVP to manually match an initial cohort before building a matching engine. A commenter recommends keeping parts of the workflow manual during testing to gather evidence before automating assumptions.
  • Feedback questioned whether escrow and legal tools justify a 20% premium, warned that strict vetting could constrain supply while weak vetting undermines trust, and suggested quality-linked insurance as a way to support the premium.
  • In response, the founder proposed lowering the transaction fee to 5–10% and building a certified talent community through senior-member mentoring, with additional B2B revenue from employer branding, sponsored upskilling, and access to the curated talent pool; these are proposed economics, not validated results.
Validating a two-sided marketplace model: How do I maximize B2B brand premiums so the talent side pays 0% fees? I don't have the answer but I like that you are doing doing it manually to test it out. Often you can automate more of one side and do mo… That's a hefty premium. The Escrow and IP & Legal Automation are ok, but I don't think they justify 20% markup. The pre-vetted talent is … You are completely right, and that feedback actually forced me to rethink the core unit economics. Charging 20% just for software feature…