ZeroNoise Logo zeronoise
Post
AI Product Work Is Moving From Model Demos to Workflow Proof
3 min read
267 docs
This brief tracks the shift from AI demos to workflow proof: private evaluation, revealed user demand, focused ICPs, production handoff, and evidence-based career positioning.

Big Ideas

Evaluate AI against the work, not the leaderboard. The current frontier-evaluation debate shows why public scores are insufficient: the interview cites Llama 4 underperforming on held-out private benchmarks while appearing highly capable on public ones. For PMs, the implication is to build an evaluation set from real enterprise workflows, with rubrics that account for capability, cost, latency, and flexibility. Longer-running agents also require stable, retryable evaluation infrastructure and fewer tasks assessed against richer criteria.

The next personalization layer is workflow ownership. A current product essay argues that software should become malleable rather than expose a fixed settings menu: let users rearrange the interface, extend preferences beyond predefined toggles, turn feedback into product changes, and treat the seams between tools as part of the experience. The author’s half-hour rebuild of an eight-agent workspace illustrates the wedge: remove repeated, personal friction before adding broad feature breadth.

Tactical Playbook

Test pull with friction before reducing friction. A fast AI prototype proves feasibility for its builder, not viability for a market; the Visory interview explicitly separates “pain” from demand and recommends testing whether people have actually acted on the problem.

  1. Ask for evidence tied to a real job. Before building, Visory requested two sensitive board packs plus an audio recording of the customer reviewing one; ten people willing to clear that hurdle became the threshold to start.
  2. After launch, compare stated and revealed preference. Cross-pack search won enthusiastic reactions and demos, but time-poor users went straight to a key-signals view. Ask for three recent instances of the problem and what the user did instead.
  3. Model time to trust, not just time to value. Visory users checked the product manually for two cycles and only relied on it around the third; infrequent-use products may need a simulated second-cycle experience rather than a standard 30-day trial.

Case Studies & Lessons

Bolt’s pivot paired focus with commercial adaptation. Its team describes a cloud-IDE market with strong hype but weak willingness to pay; after several failed 2024 directions, Bolt was the last attempt before a planned shutdown, and reported ARR rose from $0.5M to $5.5M in 30 days. The company then narrowed toward professional product builders—especially PMs, designers, and engineers—and says B2B revenue grew 10× year over year. The PM lesson is to go deep on an ICP and its workflows: detached prototyping becomes more useful when it uses production components and has a clean developer handoff. When customers exhausted a $9 subscription in under a day, the team shipped usage-based pricing within 72 hours.

Career Corner

Optimize for operating language, not tool names. A sample of 421 PM postings across 62 company job boards found “roadmap” in 85.5% of listings, while cross-functional work, prioritization, and stakeholder management each exceeded 55%. Jira appeared in 1.7%, Figma in 2.4%, and Productboard in none; “experimentation” appeared in 24.2% versus 2.1% for “A/B testing.” Use the terminology of the target role when translating experience, but treat the study as directional: it covers tech companies on one ATS and excludes agencies and regulated employers.

Tools & Resources

Workflow-grounded evaluation: The interview describes ValSmith, which turns a company’s GitHub codebase into an internal benchmark for comparing coding agents on performance and ROI. It is a useful model for PMs evaluating AI vendors: the best general model may not be best for a specific repository, and token-efficient choices can be unintuitive.

Want personalized briefs on the topics you care about?

Create your own agent and get daily or weekly cited briefs on the topics you care about, shaped by the people, newsletters, podcasts, blogs, and channels you choose.