We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
Big Ideas
AI has made execution cheap; shared understanding and judgment are becoming the constraint. One PM team reports that a one-sentence goal sent design, engineering, sales, CS, and marketing toward different interpretations, turning a simple dashboard into three months of rework. An AI-generated PRD helped because every function used it as an evolving, tracked alignment artifact—not because AI authored it. Replit describes the complementary workflow: PMs state requirements in natural language, iterate on interactive prototypes, often in under an hour, and hand off with code already started; its teams still write PRDs after prototyping. The risk is cognitive: people can accept agent-recommended choices they cannot later explain. Keep prototypes and PRDs as shared objects, but require explicit rationale and human challenge.
Trust is product architecture, not a model claim. Computer-use agents score 85% on OSWorld-Verified, but that still leaves 15 failures per 100; production buyers care about verification, escalation, error handling, security, ROI, and workflow context more than model identity. A Stripe CFO-copilot demo turns that into requirements: start with a high-stakes persona, ground answers in company policy, show the exact supporting section and confidence/source, and route below-70 responses to human review. Specify evidence, abstention, recovery, and auditability alongside the happy path.
Tactical Playbook
Run decision-first research at AI speed. When stakeholders want evidence in 48 hours, Typeform’s Research Flow combines quantitative and qualitative data; its AI moderator probes each response, with up to three follow-ups per question. Use this sequence: define the decision and stakeholder, screen the audience, collect a baseline, probe the reasons, then verify synthesis against raw responses or clips. In the team’s trust study, 25 sessions that would take 10–12 hours to field produced insights in hours or days; a 3.5/5 trust average became useful only after follow-ups surfaced accuracy and verifiability blockers, mentioned by 24 of 25 respondents 90 times. Keep human review: the researcher fact-checked every highlight for two months.
Discover admin products through verbs, not tables. A startup discussion separates commodity CRUD screens from the back-office operating system of approvals, refunds, permissions, audit logs, and manual fixes. Practitioners recommend mapping what support does by hand—resend, unlock, refund—delaying the panel until a product-specific process requires it, and letting ops change roles and views without a redeploy. This converts a vague feature request into workflow discovery and a maintainability requirement.
Case Studies & Lessons
Vertical SaaS: let constraints define the wedge. A martial-arts-school owner building Retention OS estimates that about 75% of white belts leave before blue belt, usually in the first 90 days; at 100 students paying $150/month, losing five monthly is $9,000 a year. Discovery revealed that the owner pays but parents decide whether a child stays, while a six-day-a-week instructor ignores anything taking more than a couple of seconds to log; existing billing tools do not solve that one job. Validate buyer, user, and operating constraint before expanding—small TAM is a wedge-versus-ceiling question, not a reason to go horizontal.
Career Corner
Promotions follow measurable outcomes. One PM’s example is a clean packet structure: measure the inherited problem, align on a fix, execute, then show SLA adherence rising from below 30% to above 90% while the denominator grew tenfold. The author says senior-PM framing was easy once a result leadership cared about existed; a manager adds trust with engineering, leadership, customers, and strong 360 feedback. Build your case around baseline, intervention, company-relevant metric, and stakeholder proof, with your manager as an ally.
Tools & Resources
AI pricing is a product-design decision. One AI-pricing framework argues that agentic products shift value from seats to consumption; price units can progress from tokens/compute to credits, work minutes, or outcomes. Enterprise packaging also needs entitlements, commitments, ramps, caps, and real-time metering; show usage and warn before limits so monetization does not break user momentum. Choose the invoice unit customers can understand and forecast, then test margin and usage before locking packaging.
- Pixar's story-reel process — validate before building the expensive thing. Films start as ~4,000 storyboard drawings assembled into a narrated/temp-dialogue "story reel" — a cheap, low-res 2D version of the movie — and teams iterate ~8 times on plot, script, and dialogue before committing to 3D production. Lasseter: "If it's working in our story reels, when we animate it and put color to it, it's going to work even better. If it's not working in story reels, the animation won't save it." Jobs: it "lets us beta test and iterate on our films before we actually make it," a reason the hit rate differs. The point: story and characters must work in the cheap 2D form first, surfacing concept problems before expensive production.
- "Singles and doubles" content portfolio strategy (Paramount → Disney under Eisner/Katzenberg): keep production costs low, avoid A-list stars/directors, and bet on script/story quality plus "high concept" (an idea whose originality can be conveyed briefly); source undervalued talent (comebacks, up-and-comers) — "a young Warren Buffett strategy of looking for cigar butts" with a guaranteed return. Eisner memo: "Not even the greatest screenwriter or actor or director can be counted on to save a film that lacks a strong underlying concept." Katzenberg: "celebrity can open a film, but celebrity can't carry a film." Result: 27 of the first 33 films were profitable.
- Disney Plus — a strategic product bet (rationale, trade-offs, outcomes). After ESPN's subscriber losses surfaced in 2015, Disney chose to build its own streamer rather than stay a licensing supplier: Netflix's algorithm would otherwise control whether families saw Disney content, and Disney had no direct consumer relationships or data (no emails, no demographics, no DTC revenue except parks). Launched Nov 2019 at $6.99/month with "incredibly laudable" clarity of product vision. Trade-offs taken: walked away from hundreds of millions in pure-profit Netflix licensing (~2% of company profits) and ordered every studio to ramp production for the service. Outcomes: 10M signups in 24 hours, 26M in the first quarter, 100M in 16 months (vs. 60–90M five-year target), but ~$13B cumulative losses before turning barely profitable (~$1B last year), and 132M subscribers remains subscale vs. Netflix's 325M.
- The content-treadmill vs. brand-scarcity trade-off. A scale streaming service's #1 job is retention — "you need to feed the beast" with a constant fire hose of new content — structurally the opposite of Disney's flywheel of scarce, premium content; every marginal sequel/spin-off risks devaluing the differentiated library, and "if you have a bunch of bombs in a row, I think less of the Disney brand." Streaming is a scale-economies business (Netflix: $45B revenue, $13.5B operating income vs. Disney's streaming at half that revenue and barely profitable), so the winning service is the "kitchen sink" — broad content for broad appeal. Disney's realistic path: clear #2, subsidized by other businesses, dialing back production so the flywheel heals.
- Portfolio risk signal for franchise-driven businesses. Since ~2016 (Moana, Zootopia) Disney has launched no new commercially successful franchise — "everything after that that was big box office... is harvesting existing IP," and new Pixar originals now gross only ~$200–400M while sequels set records. "Product market fit is an evolving thing... the market is changing," so owned distribution (even at lower profit) is now the way new IP reaches audiences.
- Subscription bundling as churn mitigation. ESPN's new DTC service (ESPN Unlimited, $30/month) launched at the end of 2025, aggressively bundled with Disney Plus and Hulu (both for an extra $6/month) — bundling captures casual subscribers and reduces churn, since bundle subscribers don't re-evaluate each app's monthly value.
- Career tactic: reframe a candidacy on the future. Running for Disney CEO as the #2 to an unpopular lame-duck incumbent, Bob Iger knew every board question would be "why you from the failed previous administration?" His only chance: "completely reframe the situation. This isn't about what happened in the past. This is only about the future" — he hired a political campaign consultant and ran on three concrete pillars (high-quality branded content, embrace technology, expand global reach), and won the role.
- Pricing power when supply cannot scale. When Eisner took over, Disney park ticket/parking prices had been flat since Walt's death (parking: $1) despite 1970s inflation — "5 to 10x" headroom meant raising prices added pure incremental profit with zero operational change. Today, attendance is down from its 157M pre-pandemic peak (145M/year) but per-visit spend has climbed ~5%/year for decades, making parks ~60% of operating income; with physical capacity capped, pricing is the scaling lever, backed by $60B announced park/cruise capex in 2023.
- Internal tooling as a product decision. Disney invested ~$10M in fixed-cost software (CAPS, built with Pixar) that replaced hand inking/painting and expensive multiplane cameras — "better, faster, and it keeps the budget down" (The Little Mermaid: 3 multiplane shots; The Lion King: hundreds). One fixed cost that raised quality while cutting marginal production cost.
Typeform's lead PM Sun and senior director of research Lee present Research Flow, an AI-moderated mixed-methods research tool, as a fix for a core PM bottleneck: PMs juggling roadmap, leadership ideas, competitor moves, and customer feedback need validated data fast, but researchers often can't deliver in the 48-hour window stakeholders demand, so teams historically moved forward without research . The trade-off they cite: surveys are fast but "just tell you what happened," while moderated interviews add nuance but require scheduling, live presence, note-taking, and transcript analysis .
- Core method: Research Flow runs quantitative and qualitative data in one study — the PM sets the structure and an AI moderator follows up with each respondent based on research goals . Follow-ups adapt to each answer and reference earlier statements, so the session feels like a real conversation rather than survey branching logic .
- Implementation steps: build study questions in a chat interface with an agent trained on research methods; it probes (e.g., "what decision will this research drive?") and suggests uses such as shaping product direction or prioritizing concerns, or upload an existing script . Lee notes that being crisp on what decision the research feeds and who the stakeholder is makes someone a better product builder . Studies start with a welcome screen and screeners, and each question can carry up to three AI-moderator follow-ups to capture nuance behind Likert-scale answers . The team steers non-researchers to this self-serve tool when the four-person research group can't service every request .
- Recruitment: from inside the product, source respondents from Prolific (general population/B2C) or User Interviews (experts such as buyers and doctors), or share a link with existing customers; studies can run in multiple languages with responses generated in that language, and via video, voice, or text depending on the audience's comfort .
- Live case study (trust in AI): a 5-question study (4 multiple choice, 1 open-ended ) with 25 respondents put average trust at ~3.5/5 . Synthesized insights: accuracy/verifiability are the core trust blockers; users want AI to do the heavy lifting but not decide alone; technical tasks (data analysis, code interpretation, Q&A) earn trust while emotional depth is off-limits; creative tasks spark enthusiasm but hesitation . 24 of 25 respondents raised accuracy/reliability, across 90 mentions .
- Impact and output: Lee estimates traditional fielding would have taken 10–12 hours, versus synthesized insights in "a few hours or just a couple of days or just a few minutes in some cases" ; during the webinar, completions rose from 25 to 37 in about nine minutes with topics, sentiment, quotes, and mentions updating in real time . Outputs include an executive summary, topic/sentiment breakdowns, per-respondent summaries, and highlight reels with verified, non-hallucinated quotes .
- Adoption and influence lessons: Lee initially rejected the tool — "hard no... making my life harder" — and after the team iterated on feedback from internal researchers and dozens of customers, her team adopted it in late January/early February, replacing competitor tools and internal workflows ; she fact-checked every synthesized highlight for two months before conceding the tool's synthesis was as good as or better than her own . An average like 3.5/5 is not actionable — stakeholders ask "why," which the follow-up questions answer — and one authentic sound bite (e.g., a sports fan recounting ChatGPT confidently giving wrong information) can be more influential than all the data when making a roadmap case .
Computer-use agents have crossed from demo to production: the best model scores 85% on OSWorld-Verified (was 42% a year ago; humans ~72%), and production deployments are built on raw APIs with custom harnesses, not consumer agent products .
- Buyers pick vendors on reliability, security, and ROI - the model is rarely the deciding factor (models are already good enough); the biggest differences come from the harness around the model: verification, escalation, and error handling . Failure modes matter more than benchmarks, so design for failure from the start .
- The moat is context, not model: execution is commoditized, while durable advantage is workflow-specific knowledge (runbooks, credentials, test cases, a recorded video of the job); buyers want proof of hours saved and junior-engineer operability .
- Best-fit tasks are protocol-following, standardized, repeatable work with clear verification (CRM updates, QA, portal logins, IT tickets, order/contract processing); agents fail where output can't be cross-checked or success isn't observable at runtime .
- Field cases: a CPG platform runs ~15-20M portal interactions/month with agents as self-healing fallback for scrapers, halving the scraper-maintenance engineering team; a systems integrator runs 27 workflows processing ~1,500-2,100 IT tickets/day, aiming to redeploy 20-25% of headcount; an agency automated recruiting end-to-end using a cheap non-frontier model .
- Economics: agent inference costs ~$6-8/hour (range $3-15 depending on harness), roughly break-even with offshore BPO (~$10/hour) and 70-80% gross margin vs US back-office labor (~$30-45/hour); agents are slower than people (2-3 min human task takes 8-10 min) but run 24/7 and scale without hiring. A common pattern: run once, cache as deterministic code, and call the model only on breakage to cut cost per run .
- Next frontier: multi-agent orchestration lacks standard frameworks; progress will come on accuracy, latency, and cost (e.g., accessibility-tree grounding instead of screenshots) .
Cursor's Head of Talent Adam Ward shared a hiring playbook via Lenny Rachitsky; Cursor was acquired by Elon Musk for $60B . Key takeaways:
- Forward-deployed engineers are the hottest role in tech right now—like the mobile-engineer boom, but compressed from two years to days/weeks .
- The traditional recruiting funnel ("funnel of doom") produces mediocre teams because initial responders aren't the top 20%; instead, treat every hire like an executive search: rigorously scope what "great" means, map the 50 best people in the world, and pursue them relentlessly .
- Role scoping is the most under-invested step; without it there's no sourcing strategy, pitch, or assessment; resist the "logo shortcut" of copying another company's definition .
- Work trials are the highest predictor of success; Cursor runs on-site work trials that generate technical, values, and collaboration signals; removing them dropped hiring confidence significantly .
- Post-offer is a danger zone (reneging, long gaps); Cursor runs it as a campaign—dinners with incoming hires, laptop shipped early, community before day one .
- Sourcing: avoid "Who's the best engineer?"—ask hyper-specific questions tied to scoping; triangulate names independently mentioned by multiple trusted people .
- For "bad timing" candidates, shrink the ask to "our next conversation," plant seeds over weeks/months, and use "home games" (shared meals, office visits) .
- Closing is a team sport: daily standups on candidate motivations/objections, dedicated Slack channels; "caring is free" is the biggest unfair advantage—candidate satisfaction hinges on feeling genuinely wanted; focus on "the one," not the 10 .
Lenny Rachitsky's podcast episode with Adam Ward, Head of Talent at Cursor, covers the playbook behind building high talent-density teams: why the traditional recruiting funnel—the 'funnel of doom'—guarantees mediocre hiring; Adam's three-step playbook for hiring the top 1%; the rise of the forward-deployed engineer; today's 'tale of two cities' talent market; the biggest mistake founders make with their first recruiting hire; and the worst question to ask when sourcing talent . Context: Cursor is described as having competed with every major AI lab and incumbent, continued to win, built one of the most talent-dense teams in history, and been bought by Elon Musk for $60 billion to help SpaceX win the AI race .
- Service design leader Vonnie Lang (UK public sector) argues the line between optional consumer products and services people rely on is blurring: when a product becomes a lifeline, teams inherit a duty of care whether they choose it or not, and the classic optimization/growth playbook "wasn't designed for any of this" — a new playbook is needed .
- "Average users" and "edge cases" don't exist . Case study: a nameless supermarket bank (grocery and banking arms serving very different demographics) revised its customer strategy and asked the team to "start with the people who have the least" — underserved families with two parents earning under £25,000 combined, often in persistent debt . Context: one in five UK residents (13.4M people, 4M children) live in poverty, below 60% of typical median household income . Research produced a composite persona of a young mom running a full envelope-budgeting system who knew every pound — one of the most sophisticated money managers the team met, disproving the assumption that struggling users are bad with money .
- Flip the design order: build first for the people who have the hardest time staying in your product or service, and the benefits lift everyone — plain English on gov.uk, for instance, helps second-language speakers, people under heavy cognitive load, and users in a hurry .
- "Eat your greens": separate needs from wants — wants get users in, but long-term needs keep them there, and needs shift across life chapters, so design for the user's future self . Follow design decisions forward and ask who the user is 6 months, 3 years, or 5 years later; buy now pay later was made frictionless by design so people committed without weighing consequences — teams should add "speed bumps" (guardrails) as in school zones .
- "Leave the building": dashboards and completion rates can look great while missing the real story — two users' identical click patterns in a banking app hid one person casually checking discretionary spending versus another anxiously doing mental arithmetic about direct debits and overdraft charges . Research doesn't need to be expensive: two people, train fare, a box of chocolates, and a day excused from meetings yielded ~200 post-it notes of insights; a quick, short, sharp, regular "sniff test" beats a six-month ethnography . Watching a user at a bus stop with a child in their arm writes the design brief — one-handed, low-data, on-the-move, interruptible — and failing any of those means "your product fails in the real world" .
- Closing call: protect the user care that drew you into the craft even as deadlines, dashboards, targets, KPIs, and shareholders push it out — you can still move fast without breaking the people on the other end .
- Hiten Shah (@hnshah) argues agents are the new website: the late-1990s website cycle — where a service market formed around builders who reduced the "cost of misunderstanding" — is repeating at much higher speed. A proposal "can tell you more about a market than a demo" because it prices the distance between seeing an agent work and trusting it with real work .
- The most revealing line in an agent proposal is monthly maintenance: an example proposal (email, invoices, call-memory, privacy/control, keep-local) included discovery, phased implementation, and a monthly maintenance relationship, and the durable business will belong to whoever remains responsible for an agent's performance long after the initial build .
- Agent maintenance becomes trust preservation because agents degrade silently — acting on stale context, missing exceptions, or producing plausible-but-wrong answers, with skills drift — unlike a broken website that announces its failure. Systems need evaluation, monitoring, correction, permission review, and periodic redesign .
- Recurring demand is ordinary business work (email triage, call memory, CRM hygiene, invoices, internal search, lead routing, reporting); agents become valuable when they hold context across people, tabs, inboxes, and memory and reliably carry work forward. The build work is "operating design": deciding what the agent can see, which tools it may use, when it may act, how output is checked, and what happens under uncertainty, plus memory, permissions, escalation paths, and learning from failure .
- Market signals: Fiverr reported an 18,347% increase in searches for AI agent freelancers over six months in 2025 (from a small base), and Microsoft's 2026 Work Trend Index found organizational factors (culture, manager support, talent practices) accounted for more than twice the reported AI impact of individual effort — individual experimentation is outpacing company redesign, and that gap is becoming a market .
- Platforms will compress the middle: Shopify turns natural-language requests into Flow automations and generates custom admin apps via Sidekick; Salesforce says Agentforce reached $800M ARR and 29,000 deals by early 2026. Horizontal "we make AI agents" offers will age quickly; durability rises with the cost of error and specificity of context — strongest in vertical workflows, cross-system integration, private/local environments, and evaluation-heavy systems, because platforms can standardize actions but accountability takes longer to standardize .
- The framing post: every business owner eventually asked "Who can make this work for me?" about websites — a question that created an enormous industry — and they are about to ask it again about agents .
- Shreyas Doshi argues there is no single template for a great PM: highly effective PMs differ widely, succeeding via analytical strength, ability to drive teams, design talent, or deep domain intuition, despite weaknesses elsewhere; the idealized "always humble, great communicator, instantly responsive" PM doesn't reflect reality.
- There is similarly no magic trait: Doshi has seen equally effective PMs with fixed and with intense growth mindsets.
- The one mindset he finds highly correlated with PM effectiveness is not shying away from hard, ambiguous problems. This is a mindset, not a skill — skills are coachable, mindsets are learnable but not teachable; it typically changes through experiences, often failures, only when someone is ready.
- Because all product work involves problems, Doshi says the advanced skill is judgment about which problems to solve and which to leave alone. Strong problem solvers without this discernment try to fix everything and end up on a treadmill: solving one problem creates downstream problems, and chasing those keeps teams on unimportant work; leave some nails alone.
Priam Bagani, head of product at an AI billing infrastructure company, presents a framework for pricing and monetizing AI products in a Product School talk. Key points:
- AI product pricing is fundamentally different from SaaS: per-inference/agent-run margins, user-level unit economics, and consumption (not seats) as the value driver .
- Pricing unit options form a ladder: raw tokens/compute → credits (e.g., Lovable) → units of work (e.g., CodeRabbit prices its Slack agent per minute) → outcome-based (e.g., Finn $0.99 per resolved ticket; HubSpot $50 per resolved conversation); choose based on product and customer .
- Cursor's packaging: Pro $20 includes $20 of frontier-model spend; Pro Plus $60 includes $70 of frontier spend plus generous own-model usage; economics work via auto-routing to own models, spend caps, and 60-70% of revenue from enterprise deals (seat fees, pooled usage, governance controls) .
- Lovable's credit model: subscription + credits with rollover/top-ups and daily 5 free credits that expire daily to drive habitual usage; credits abstract across design, hosting, and AI features, and users see per-action minutes/credits consumed .
- Enterprise AI contracts require four primitives: entitlements (pooled usage, departmental splits, overage rates, not-to-exceed caps), commitments, ramps (e.g., $3M deal spread as $700K/$1M/$1.3M over years), and price/settlement — every element is negotiable .
- Monetization infrastructure must meter/rate usage events in real time and keep an auditable credit ledger; design principles: preserve user momentum (hold-and-authorize, thresholds/alerts at 80%/95%), provide usage transparency, and treat limit-hit moments as smooth monetization opportunities (upgrade/top-up) .
- Manage pricing changes deliberately: Figma Make ran free from launch (Dec 2025) until they understood usage before enforcing pricing; Cursor publicly apologized for a poorly communicated pricing change — consider grandfathering, cushioning, and clear communication .
- Design questions: What unit of value will customers recognize on the invoice? Is usage bounded (sell predictability) or volatile (pass through/dual fixed+consumption)? Who decides and how fast (PLG vs enterprise)? What happens at usage walls? How will pricing evolve?
- A PM at a 60-person B2B SaaS for regulated orgs (bank statements, government notices, utility bills) deploying its first status page on Atlassian Statuspage faces CEO fears that larger competitors (none with a status page) will use it to attack uptime; enterprise contracts already require publishing uptime stats and SLAs, and the page is live but not connected to health APIs yet .
- Visibility should not be all-or-nothing: Statuspage can restrict which services, components, and metrics each authenticated customer group sees — e.g., SLA-based transparency per customer — or just expose basic health with a simple "watermelon" display .
- A login-gated status page is a viable middle ground: only logged-in customers see it (and content can be tailored if a subset of the base is affected); if login itself is down, swap in a public-facing page .
- Launch quietly and validate: publish the page without an announcement, gather client-stakeholder feedback on its value, then use that feedback to prioritize changes and argue for/against keeping it .
- One PM's launched public cloud status page — with region- and feature-level categories so customers subscribe only to relevant events — deflected "hundreds of thousands of support calls" over years and was expanded beyond Sev1 incidents .
- Outages become public whether or not you have a page; your own page lets you control the story, and in regulated enterprise sales a public page signals operational maturity while its absence makes security/procurement assume poor reliability .
- If a public page would expose poor reliability, customers already know; use it as leverage to push reliability improvements if the PM has enough clout to make it a CEO goal — and if uptime numbers are weak, consider tiered pricing: a standard tier tolerant of brief outages plus a higher-priced "High SLA Guarantee" tier on a VIP instance that receives updates only after the general population proves reliability .
- Before committing, ask what problem the status page solves and whether customers actually care about availability .
Kavak's agent-per-customer bet (case study). Kavak — a vertically integrated used-car marketplace in Latin America, headed by AI lead Ali Masa — bet the company on agents, asking 'how would we build Kavak in 2035 with GPT-10-level intelligence?' . The answer: every customer gets a dedicated agent with its own virtual machine, memory of years of interaction history, and a hard long-term goal (maximize customer lifetime value), rather than workflow automation . The bet: agents outperform the best hired human on conversion, LTV, and customer experience . Results: agents handle 96% of interactions and 95% of transactions, with 100,000–200,000 agents instantiated daily ; NPS and customer satisfaction tripled; agent conversion beat the human team by 50% at launch and now 2.1x ; car loans that normally take 2+ months in Mexico are approved in under 3 minutes ; warranty rates fell ~20–26% after mechanics received the agent sidekick 'El Mike' . An agent CEO running a carved-out city in Mexico raised profits 50% in its first month against a 2x goal .
Eval-driven development. 'Evals over agent demos': spend roughly equal engineer time, tokens, and money on evals as on the agents themselves, and evaluate business outcomes — did the customer convert and re-engage — rather than superficial KPIs like call count or minutes . Kavak's measurement shifted from transactional (cars bought/sold) to relational (LTV per customer) .
Architecture trade-offs. Kavak ran tens of thousands of agents on a multi-agent workflow-graph harness in production, then concluded with Opus 4.5 that the graph paradigm constrained the intelligence; it destroyed two years of working systems and rebuilt on a VM-per-agent harness with memory, evals, and CLI access to every company API — the 'self-improving organization' .
Scaling playbook. To go AI-native: (1) redesign the company and its APIs around agents rather than layering ChatGPT onto the existing org; (2) generate data and feedback loops by putting agents live in front of customers; (3) shift success metrics to relationships and LTV . The transformation must be top-down with a clear 3–5 year plan, not bottom-up hackathon use cases .
Token-tier ROI framework. Tier 3 = agent tokens whose per-token ROI is directly measurable; tier 2 = indirectly measurable (e.g., developers in the codebase); tier 1 = generic copilot/chat usage with unknowable value — where most companies sit. Track that each token maps to business benefit .
Org and career implications. The org becomes flat, senior, cross-functional teams that build agents, work for agents, or work in the physical world . Kavak retrains everyone from CEO to mechanics through a six-week Jedi Academy that ends with shipping production agents; staying relevant means upgrading skills every month or two . Masa's advice to builders: tools are democratized, so build deep for an AI-native future — superficial AI adoption yields only ~6–10% improvement while redesigning the company around AI yields the large gains .
Lenny Rachitsky shared a video with Adam Ward (Head of Talent at Cursor) addressing hiring mistakes, covering the “funnel of doom” and how to fix your hiring process ; the full conversation is available on YouTube .
Shreyas Doshi published a new video on the problem-solving mindset, covering whether this mindset is immediately coachable, why problem solvers must also develop great judgment, and the downsides of solving all problems .
Lenny Rachitsky posted a video titled "The worst question to ask when sourcing great talent" , and in a reply pointed to his full conversation with @wardadamp on YouTube (https://www.youtube.com/watch?v=zegYJ6dhIg4) .
Atlassian PM evangelist Axel Suria shared AI-assisted GTM workflows for product teams:
- As AI lowers the cost of building, discoverability and market differentiation (GTM) become costlier; distribution is increasingly the hard problem — quoting Box CEO Aaron Levie .
- Differentiates skills (bundled instructions telling an agent how to act) from agents (orchestrate multiple skills via tool calling and loops until an outcome) .
- Dia's Morning Brief surfaces a daily, context-rich to-do recap (Slack, Jira, calendar via MCP) with embedded skills like drafting briefing talking points .
- Built a Rovo Studio automation that triggers when a Jira Product Discovery work item moves to 'production': it drafts a landing page brief and passes it to a Replit agent that generates/publishes a page mock-up, then comments the link back on the work item — cutting hours of back-and-forth with the creative team to minutes .
- Jira Product Discovery tips: custom fields with a RICE formula for stack-ranking; program/tactics/deliverables hierarchy in board view; matrix, timeline, and opportunity-solution tree views linking bets to deliverables and OKRs; shareable live roadmap URL so stakeholders can self-serve status .
- Atlassian's Teamwork Graph is now exposed via MCP/CLI, allowing agents to pull Atlassian context into Claude, Cursor, or Codex .
- For measuring webinar impact, cross-referenced Databricks, Salesforce, and goals via MCPs to build a dashboard in ~30 minutes (vs weeks of tickets); turned the workflow into a reusable global skill and made it available to his team through a Slack on-demand command .
- Recommended pattern: start from your heaviest day-to-day problems and then find AI solutions; be kind to yourself amid pressure to deliver more with less .
Hiten Shah shares cost-strategy guidance (via Shishir Mehrotra): at enough volume, every AI workload becomes a decision about what to rent and what to own, and open models give companies a path to turn repeated work into infrastructure they control .
Grammarly's approach: many companies hit cost problems by routing every call to a frontier model ; instead, because Grammarly was building its own models before GPT-2, most of its 100B+ weekly LLM calls go to models it built itself, often fine-tuned from open source (e.g., Llama or Gemma) and served on its own infrastructure. The instinct to reach for the biggest model for everything is misguided — most workloads that run at scale don't actually need it .
Hiten Shah (@hnshah) notes AI now makes it possible to "procrastinate at 100 commits a day" — high commit volume no longer signals real progress, a caution for PMs tracking team productivity via activity metrics.
- Hiten Shah describes four AI marketing systems he and Sam Asante built around marketing jobs they already do, including two Hermes agents in Slack and a media system that saves hours and may be turned into a product .
- The systems gain advantage from specificity: they encode the builders' context and decisions, turning each marketer's private operating system into executable software; tools that look too specific are where the edge comes from .
- One system deliberately uses multiple sub-agents and a real-time interface, complexity the author normally avoids but accepted because the job required it; a model becomes an agent when its harness lets it see, do, remember, and check .
- Samuel Spitz, AI product PM at Replit, argues the traditional PRD is dying: it's annoying and slow to write, translating the product in your head to text and back is lossy, requirements constantly change, and implementation reveals hidden constraints . At Replit, PRDs are still written but almost always after a prototyping phase, which speeds up product development .
- New AI prototyping workflow for PMs: express precise requirements in natural language, let a vibe-coding tool like Replit generate a first branded version, iterate based on team feedback, then hand off to engineers; PMs without design or coding skills can do this, often in under an hour, cutting weeks off the implementation cycle, and sometimes much of the code is already written by the agent before engineers start .
- AI makes internal tooling accessible to PMs: low-code tools were still semi-technical, but now PMs without engineering backgrounds can build dashboards, automations, and workflow tools; people closest to bottlenecks can fix inefficiencies, and five to ten such tools can transform an organization .
- AI compresses timelines and levels expertise: a week-long internal tool build can become a one-prompt task, and PMs can push back on "impossible" estimates by pointing to a working prototype they built in one prompt .
- PMs also use AI coding agents for long-tail tasks such as one-off data requests, emails, go-to-market lead sourcing, and recruiting research .
- Replit builds slide decks as websites because AI models excel at web design but produce templated slide outputs; this is used for pitch decks, sales decks, and internal presentations, with decks typically taking 4-6 minutes .
- OpenClaw founder validated product-market fit by building a WhatsApp relay for personal annoyance and observing strong emotional reactions from friends, including nontechnical friends who were upset when told the tool wasn't ready for them.
- Launching the agent in a public Discord server triggered overnight virality (800 messages while he slept), and within 8 months the project attracted 18,000+ issue/PR contributors and 111,000+ total issues/PRs.
- Feature creep: every new feature shipped with a config option to avoid breaking setups, accumulating ~9,500 configuration options, making comprehensive testing impossible; "It is infinitely harder to evolve software that has users."
- Dependency risk: over-optimizing for one model provider backfired when the provider gave 24 hours' notice before disabling subscriptions; "Your dependencies business model is your business model."
- Security hardening trade-off: extensive sandboxing and permissions made updates slower and broke user workflows, while most users didn't value the security abstractions.
- Personal brand as a moat: "Everything you can build can be forked or cloned, but your name cannot," making personal brand more important than any single product.
- Enjoyment drives output: "Fun is velocity. The weeks I enjoyed building, the product got visibly better."
- First-user strategy: the founder should be user number one, with friends as the next ~20 users; if the builder isn't excited by the product, it likely won't succeed.
- Distribution is the hardest problem in an AI era where building is easy: "eyeballs are kind of the the most expensive currency."
- For new startups, he advises building something you personally want to use and choosing "hard and boring" categories because they attract users who appreciate the solution.
Can Agents Use a Computer Yet? We've Got the Data
Can Agents Use a Computer Yet? We’ve Got the Data

It seems obvious to say, but if you leave Silicon Valley and go out into the rest of the world and tell them, “there are these things called agents, which are pretty smart, and can do tasks with you, and automate some of the repetitive parts of your work”, chances are the first question you’ll get back is, “Can they use a computer?”
This is a good question! Can they, really? The long horizon of productivity potential, out in the real economy we’re going to go unlock over decades, runs through pretty everyday work: can an agent sit (metaphorically) at a desk 24/7, and be trusted to use a web browser, fill out forms, click the right buttons, and not make mistakes? This is the domain of Business Process Outsourcing (BPO), which historically meant, “can this work be outsourced?” but now has a new agentic frontier. We wrote about this last year (opens in new tab), when the computer-use landscape was still mostly a bunch of demos. A lot has happened since then.
The models have improved faster than almost anyone expected. Computer-using agents are beginning to hold up in production at scale and on narrow, repeatable workflows: updating systems of record, moving data through portals, processing tickets, checking records, and handling the long tail of software where no clean API exists. With the right infrastructure, computer-use capabilities can now be deployed to tackle end-to-end tasks at scale, which before required either human supervision or direct human completion.
Today, workflows leveraging computer-use are far from perfect: agents are brittle when work drifts off the runbook, and for certain use-cases where caching is intractable (more below) they are expensive enough that the math does not work everywhere. But we’re seeing production deployments for standardized back-office work, especially where labor would otherwise be clicking through legacy systems by hand; the cost curve is starting to look compelling, considering that workflows leveraging computer-use offer structural advantages such as 24/7 availability and - most importantly - scalability to meet demand.
The first wave of computer-use infrastructure was about making agents capable: seeing, clicking, typing, recovering from mistakes. The next wave is about making them useful inside actual companies. As raw UI navigation becomes a model-layer commodity, the model is no longer the main bottleneck and the durable advantage moves up the stack: context, permissions, process knowledge, validation, escalation, error handling, caching, and the hard-earned understanding of how work actually gets done inside one specific customer’s organization to map a workflow end-to-end. In other words, the frontier is shifting from “can the agent use a computer?” to “can it reliably do this job?”
From Humans Watching Every Step to Real Autonomous Workflows

A year ago the best computer-using model scored 42% on OSWorld-Verified; today’s best scores 85%, above the ~72% humans manage on the same tasks (this means they successfully completed 85 of 100 tasks). In production these general frontier models run much like they do in the benchmark: the labs expose computer use as an API - the model gets a screenshot, returns clicks and keystrokes, with OpenAI’s CUA also layering in accessibility-tree or DOM data where available - and builders wrap that loop in their own harness: a sandboxed VM or browser, plus the orchestration, verification, and retry logic around it. Notably, almost nobody deploys consumer products (Claude, ChatGPT agent mode) for this - founders and enterprises build on the raw APIs, or buy from vendors who package them. And the capability jump is what made those setups viable - “the models weren’t good enough to use in production on their own until Opus 4.6 in February 2026,” as one founder building in the space put it. Somewhere in the last eighteen months, computer use capabilities crossed from demo to being deployable in the field.
Of course, benchmarks aren’t always the best proxy for the viability of a real-world deployment. OSWorld counts completed tasks, so 85% still means 15 of 100 failed, and a business process only finishes if every step does. Back-office work doesn’t grade on a curve: if a person reviews every output, no labor was saved. (It’s analogous to what’s happening in coding right now: the scarce resource is no longer writing the code, it’s vouching for it.)
We found that the best way to think about what matters is to go beyond the benchmark and focus on the core question: can a business process be reliably automated with computer-use capabilities? Under this lens, what makes the biggest difference is everything around the model - that is: verification, escalation, error handling when a retailer portal changes its layout overnight.
Perhaps the clearest tell is that one operator we spoke with, who runs millions of automated tasks a month, couldn’t tell us which model executes them; he hadn’t needed to find out. His vendor swaps models underneath him the way a cloud provider swaps hardware. But he did trust the computer-using agent to run these tasks. Bottom line: when your heaviest users stop checking the leaderboard, the leaderboard has stopped being the story.
Thus, the chart above explains why production deployments exist in 2026 and didn’t in 2024. From there on, what determines whether they work is everything else - and that’s the rest of this piece.
Agents are Protocol Following
We had various conversations with teams running workflows leveraging computer-use capabilities in production, and we learned from their experiences that protocol following tasks work best. Unsurprisingly, computer-using agents break on more complex workflows where accuracy is harder to verify. The overall takeaway is that computer-using agents are strongest on standardized, repeatable tasks with a clear, well-defined path. The real unlock is the long tail of software where no clean API exists and a person would otherwise be clicking through a UI by hand. In practice, the work looks like updating records in a CRM, QA, logging into government and insurance portals, pulling data off databases and regulatory pages, retail order processing, contract processing, or IT tickets in ServiceNow.
We believe the voice of the user here tells the story much better than any theory. Some examples: a CPG data platform walked us through how they run ~15-20M automated portal interactions a month, using agents as a self-healing fallback for hand-coded scrapers - when a retailer portal changes its UI, the agent diagnoses the break, fixes the automation, and keeps data flowing before an engineer ever sees the error. Once implemented, they told us, they cut the engineering team dedicated to scraper maintenance in half and re-allocated staff capacity to other workflows. In another case, from a global systems integrator, we learned they have 27 live workflows leveraging computer use agents that process ~1,500-2,100 IT tickets a day, with the ultimate goal of redeploying 20-25% of headcount on low-margin managed-services contracts. And finally an agency walked us through how they automated a recruiting workflow end-to-end to populate data in an applicant tracking platform as soon as a candidate interview was over. To do so they run a cheap non-frontier model because it “does everything we need and does it well.”
The clearest pattern is workflows where in theory a computer-use agent could operate and solve the task, but there is either no clear answer for “what good looks like” (i.e. they are hard to evaluate) or no reliable way of determining whether a task succeeded. Usually, issues arise fast when: (1) you can’t cross-check the output - think of an agent extracting payment terms from contracts into an ERP: if it reads “net 60” as “net 30,” the record looks perfectly plausible, passes every visual check, and nobody catches it until an invoice goes out wrong; and (2) in some cases there is no signal to verify success at the time the task runs - think of an agent submitting a claim on an insurance portal: the submission goes through, the screen says “received,” task done. Except two days later an adjuster calls the office because a policy number needs confirming before the claim can be processed. A human who filed that claim picks up the phone and sorts it out in thirty seconds; the agent has no idea the call ever happened, and the claim quietly stalls. Net-net is that a smarter model doesn’t fix a process whose ground truth shows up as a phone call to somebody’s desk a week later, unless the harness is designed to handle the edge case from the get-go.
Buyers Care About Infrastructure
For the buyers we spoke with, the model itself is rarely the deciding factor, as “the models today are already good enough.” In practice, they evaluate and pay for everything around the model: the infrastructure to run reliably at scale, pass security review, and prove ROI. Users don’t care whether the solution uses a given frontier model; rather they focus on whether it can actually get the task done at scale and reliably. Period.
As a consequence, failure modes matter more than any benchmark, and design for failure needs to be a first-class concern from the start because a solution that does not handle failures well will never be adopted in production. One example of what this looks like in practice and a pattern we ran into more than once: the agent runs the workflow once, the system caches it as deterministic code, runs execute as cheap repeatable code from then on, and the model comes back only when something breaks - to diagnose, fix, and re-cache. With this approach, cost per run falls over a workflow’s lifetime, and cheaper models just lower the bill. What’s interesting about it is how it handles uncertainty. Where before deterministic code simply failed, or a human had to review and fix every break, here the agent absorbs that uncertainty on its own. It’s one way of designing for failure, and it shows what buyers are actually rewarding and using at scale.
We did not encounter more sophisticated use cases among the users we spoke with, which tells us the market is still chipping away at the low-hanging fruit. That said, there is a long list of workflows that can be automated this way before anyone needs the harder tasks.
The Model Is Not the Differentiator, Context Is
For founders, the more important shift is what is becoming commoditized. Building a computer-using agent used to mean wrestling with Selenium or Playwright, or more recently Stagehand, and stitching together DOM or video recordings to capture a workflow. That whole execution layer is getting abstracted away, the same way Claude Code abstracted the scaffolding around coding agents. If clicking the right button is no longer the hard part, it is no longer the moat.
Unsurprisingly, the context and knowledge of the workflow are durable. The hard part is not whether an agent can navigate an SAP screen, it is whether it understands how a particular company actually gets work done: the tribal knowledge, the internal terminology, the preferred formats, who to escalate to and when, how to handle failures, how to verify output reliably. In practice that context lives in runbooks, in access and credentials, in test cases and guardrails for when a workflow goes off-script, and increasingly in a single recorded video of someone doing the job once. None of it is general. All of it is specific to a company, and often to one team. That said, it is exactly the kind of specific, unglamorous problem that focused startups tend to solve better than model providers, which is why we think the next generation of agentic coworkers gets built at the application and context layer, not the model layer.
The buyers we spoke with squarely confirmed this. They picked vendors on whether the product reported hours saved without extra work, and whether a junior engineer could run it. For now, the moat is not the frontier capability—rather, it’s being the vendor an enterprise is allowed, and able, to use at scale in production.
The Real Inflection is Economic, Not Just Technical

The cost data is also encouraging. Take the numbers above as an order of magnitude, not a precise quote. Running an agent costs roughly $6-8 per hour of inference today, but in practice anywhere between $3 and $15 depending on how the harness is built - how often it screenshots, how much context it carries, how much of the work it can hand off to deterministic code. These figures describe the agent operating the UI screenshot by screenshot with a frontier model - the most expensive mode there is. Well-built harnesses reserve that mode for what actually needs it, and let cheap deterministic code handle the repeatable parts - not every workflow can be optimized this way, but where it can, blended cost drops fast. So read the comparison as the worst case, and even then an agent is roughly break-even against offshore BPO at ~$10/hour fully loaded, and pencils out to a 70-80% gross margin against US back-office labor at ~$30-45/hour. In production, the harness drives real-world cost as much as the model does.
Same caveat on speed. In agentic mode, agents are still slower than people, and it is not close - a task someone finishes in two to three minutes can take an agent eight to ten, and academic benchmarks put the gap even wider. Deterministic runs flip this: code executes faster than any human - but for the agentic work the argument is not speed. It is that an agent runs around the clock, costs a fraction of US labor, and scales without hiring.
This comparison works for BPO buyers and ops teams, but the unit economics look different if you are the one selling agent-hours, because costs are less predictable in the wild. COGS are inference plus retries (i.e. failed runs still burn tokens) and margins compress when context grows or screenshot frequency goes up. Vendors manage this by pricing per task, per hour, or per outcome, each with a different risk profile depending on workflow variance. There are also the practical realities of things like monitoring, maintenance and human escalation, which are being priced in, just as they would be with human workforce. There is no one-size-fits-all business model here yet, and the answer varies by vertical.
And the math only gets better - inference keeps getting cheaper, and open-source models are getting good enough for a growing share of these workflows. For any task an agent can reliably solve, embedding computer-use will likely be way more convenient than human labor. So the real question is no longer whether the economics work - it’s how far the set of tasks that can be solved reliably extends, which is where things are heading next.
Where Do We Go From Here?
Over the past year, labs and a wave of startups have poured hundreds of millions into computer-use RL environments - the sandboxes where a model practices real tasks and gets rewarded for finishing them - with companies like Mechanize, Habitat, Fleet, Chakra, Deeptune, Matrices, and Originator building the training and eval substrate underneath the frontier models. That spend is what shows up as better reasoning, better state tracking, and more tolerance for apps that misbehave. The models still need careful harnessing to hold up in production - run-caching being the clearest example - but the raw capability was bought, deliberately, through this training infrastructure and will only get better over time.
Architecture, though, is a different story. Most computer-use systems deployed in production today are single-agent: one model, one task, one session. As workflows get more complex and latency becomes a constraint, multi-agent architectures start to matter. For example, a planner decomposes the workflow, executor agents handle subtasks in parallel and long-running agents bring their own problems: memory, trust, and failure rates that compound over time. The teams doing interesting work here are all building bespoke orchestration, because no standard framework exists yet. The Claude Code analogy is instructive: when coding agents matured, a scaffolding layer emerged to abstract the orchestration away. The same is likely to happen for workflows leveraging computer-use capabilities, and that abstraction layer is one of the more interesting unsolved infrastructure problems in the space.
From here, future developments run along three lines: accuracy, latency, and cost. Accuracy is most important, and, as we explained above, represents the difference between a cool demo and actually solving the problem - catching anomalies, checking its own work, escalating only when it actually needs to. Latency is the one most likely to surprise people: some teams already cut it today by grounding on the accessibility tree instead of screenshots. Standard Intelligence’s general computer action model is trained on an 11-million-hour video dataset, running at 30 FPS, and is an early signal that the step-by-step screenshot loop slowing today’s agents is a solvable problem, not a permanent tax. Cost keeps falling as inference gets cheaper, and smaller non-frontier models take over the routine clicks. These three vectors in addition to security and governance (e.g. credentials, audit logs, data retention, prompt injection, and accountability, permissioning).
Enterprises can and are benefiting from computer-using agents for narrow workflows that have high volume, repetitive steps, stable business rules with legacy interfaces or missing APIs. For now, they are best suited for tasks with immediate, machine-observable evidence of success, tolerable failure consequences, and clear escalation routes. But with the above developments, improvements are real and rapid, making computer-using agents more viable for more types of work.
Let’s just say - the future for computer-use capabilities is bright!
Computer-use agents have crossed from demo to production: the best model scores 85% on OSWorld-Verified (was 42% a year ago; humans ~72%), and production deployments are built on raw APIs with custom harnesses, not consumer agent products .
- Buyers pick vendors on reliability, security, and ROI - the model is rarely the deciding factor (models are already good enough); the biggest differences come from the harness around the model: verification, escalation, and error handling . Failure modes matter more than benchmarks, so design for failure from the start .
- The moat is context, not model: execution is commoditized, while durable advantage is workflow-specific knowledge (runbooks, credentials, test cases, a recorded video of the job); buyers want proof of hours saved and junior-engineer operability .
- Best-fit tasks are protocol-following, standardized, repeatable work with clear verification (CRM updates, QA, portal logins, IT tickets, order/contract processing); agents fail where output can't be cross-checked or success isn't observable at runtime .
- Field cases: a CPG platform runs ~15-20M portal interactions/month with agents as self-healing fallback for scrapers, halving the scraper-maintenance engineering team; a systems integrator runs 27 workflows processing ~1,500-2,100 IT tickets/day, aiming to redeploy 20-25% of headcount; an agency automated recruiting end-to-end using a cheap non-frontier model .
- Economics: agent inference costs ~$6-8/hour (range $3-15 depending on harness), roughly break-even with offshore BPO (~$10/hour) and 70-80% gross margin vs US back-office labor (~$30-45/hour); agents are slower than people (2-3 min human task takes 8-10 min) but run 24/7 and scale without hiring. A common pattern: run once, cache as deterministic code, and call the model only on breakage to cut cost per run .
- Next frontier: multi-agent orchestration lacks standard frameworks; progress will come on accuracy, latency, and cost (e.g., accessibility-tree grounding instead of screenshots) .