ZeroNoise Logo zeronoise
Post
AI skills need A/B evals, not install counts, and win/loss needs the buyer's own words
•
4 min read
• 311 docs
Hiten Shah on testing whether AI skills actually change agent output, ThoughtSpot on evals as the new PRD, Gamma's rebuild against generic AI output, practitioner advice on win/loss, and a PM job market that keeps tightening.

A skill being available doesn't mean the agent uses it

Hiten Shah's team has built 16 skills for competitive product marketing, and he asks a basic question: does a skill change the work once the model has it? He notes that Skills.sh reached one million skills and nearly 280 million installs in seven months, so availability "tells us very little about quality" . He cites a Vercel Next.js eval. The baseline agent passed 53%. Making a skill available left the score at 53%, partly because the agent skipped the skill in 56% of cases. Telling the agent explicitly when to use the skill raised it to 79%, and a compressed docs index in AGENTS.md reached 100%. He says one test can't settle how every skill should be designed .

His test is one PMs can reuse. Give the same model the same job and the same evidence, once with the skill and once without. Define what good looks like before the run, and keep the failures. If the model catches up, shrink the skill or remove it, and rerun the test whenever the model changes . He also writes each skill around the job it serves. For example, a CRM loss reason records what someone typed into the CRM, not why the buyer left. If nobody asked the buyer, the buyer's reason is unknown . In a demo, Claude without the skills treated an old source as recent and stated a guess as fact. With the skills, it dated every claim and noted what it couldn't see. The skills are MIT-licensed and work across Claude Code, Codex, Cursor and others .

ThoughtSpot: the eval is the new PRD

Francois Lopitaux (SVP Product, ThoughtSpot) told Mind the Product that "your eval system is almost becoming your PRD." It defines what good looks like and what to avoid, because a prompt-based product has no fixed UI to spec . Analytics answers have to be the same every time. So he uses LLMs only where they're needed and grounds them in a semantic layer that defines terms like "new customer" and "revenue". In his words, without that layer an LLM is "a new intern" that makes poor judgments . A few other points:

  • Teams are organized by product "track," not by feature, so the people closest to customers decide what to build next .
  • Pruning matters more now that code is cheap to generate .
  • He expects smaller teams and a lower PM-to-developer ratio, as deciding what's worth building becomes the bottleneck .
  • For a recent APM hire, he asked candidates for a Git repo and a video explaining the problem their code solves .

Product moves

  • Gamma 5. Gamma concluded that its output looked "too similar to all the other AI tools." It started over and rebuilt for visual variety and brand fidelity across presentations, docs, social assets, and graphics .
  • AI simulations. Lenny Rachitsky is following teams that use AI simulations to test product ideas and flow tweaks. He names Simile, Primitive Labs, Synthetic Users, Tenera, and Seldon, and asks how they've worked for people. No results are reported .

Craft: win/loss is not competitor research

A r/ProductMarketing thread argued that knowing a rival's pricing, features, and positioning doesn't tell you why a customer chose them . Advice from the thread:

  • Run unbiased win/loss interviews, not run by sales, and do 10–20 before looking for patterns .
  • Ask buyers what they expected to go wrong six months after signing. This tends to surface implementation risk, internal politics, and trust issues .
  • Commenters disagreed on price. One said it's almost always the reason . Another said C-suite buyers often pick more expensive competitors for ROI reasons .

Teresa Torres made a related point about records. Her transcripts kept showing that her confident memories were wrong. Notes and transcripts each add interpretation, and a record can be used to collaborate or as a weapon .

Job market

A 50-year-old senior PM with 22 years in tech has applied to more than 200 roles in nearly a year and gotten few interviews . Another commenter missed one of six technical requirements and was passed over for a role that had been open eight months. They described a "buyer's market" in which managers hold out for a perfect, cheaper PM . A third said the market never recovered from its 2022 peak and hiring has shifted toward senior roles . These are personal accounts, not market data.

AI skills need A/B evals, not install counts, and win/loss needs the buyer's own words
Summary
Coverage start
1 day ago
Coverage end
17 hours ago
Frequency
Daily
Published
16 hours ago
Reading time
4 min
Research time
3 hrs 52 min
Documents scanned
311
Documents used
14
Citations
22
Sources monitored
98 / 99
Insights
Skipped contexts
Source details
Source Docs Insights Status
rahulvohra 0 0
Paul Graham 1 1
Tony Fadell 0 0
Patrick Collison 2 1
Daniel Ek 0 0
Gustaf Alströmer 0 0
Stewart Butterfield 0 0
PM Diego Granados 0 0
👨🏻‍💻☕️ 0 0
scott belsky 1 1
Ryan Hoover 0 0
Janna Bastow simplybastow.bsky.social 0 0
Jackie Bavaro 0 0
Sachin Rekhi 0 0
Dan Olsen 0 0
The community for ventures designed to scale rapidly | Read our rules before posting ❤️ 73 9
Will Lawrence 0 0
Product Marketing 22 2
Ami Vora 0 0
PM Interview: Practice Group for Product Manager Case Interviews 0 0
One Knight in Product 0 0
Aakash Gupta 3 1
Shreyas Doshi's Product Almanac | Substack 0 0
Lenny Rachitsky 0 0
Acquired 0 0
a16z 1 1
Exponent 0 0
Product Alliance 0 0
Product Management Exercises 0 0
rocketblocks 0 0
Product Design 2 0
ProductManagementJobs 26 2
Product Management 141 7
Product Management - The place for all things product 12 2
Product Management 0 0
Aspiring and current tech PM's 0 0
Masters of Scale 0 0
Product Science Group 0 0
How I built This 0 0
SaaStr AI 0 0
productized io 0 0
Lenny's Reads 0 0
The Product Folks 0 0
Strategyzer 0 0
Lenny's Podcast 0 0
AJ&Smart 0 0
Y Combinator 0 0
Product School 0 0
Mind the Product 1 1
@andrewchen 1 1
The Looking Glass 0 0
Leah’s ProducTea 0 0
Run the Business 0 0
Product Managers at Work 0 0
The Product Compass 0 0
Ravi on Product 0 0
Productify by Bandan 0 0
Product Thinking with Melissa Perri 0 0
Product Talk Daily 0 0
The Beautiful Mess 0 0
Gibson Biddle's "Ask Gib" Product Newsletter 0 0
Casey Accidental 0 0
Hiten Shah 6 4
Product Growth 0 0
Perspectives 0 0
Lenny's Newsletter 0 0
andrew chen 3 2
Brian Balfour 0 0
Casey Winters 0 0
elena verna 0 0
Kevin Weil 🇺🇸 2 1
April Underwood 0 0
Julie Zhuo 0 0
Marty Cagan 0 0
Lenny Rachitsky 10 3
Christian Idiodi 0 0
John Cutler 0 0
Teresa Torres 1 1
Gibson Biddle 0 0
Shreyas Doshi 2 0
Adam Nash 0 0
Merci Grace 0 0
Jackie Bavaro 0 0
Hunter Walk 0 0
Brian Balfour 0 0
Scott Belsky 0 0
Nir Eyal 1 1
Teresa Torres 0 0
Julie Zhuo 0 0
Andrew Chen 0 0
John Cutler 0 0
Ken Norton 0 0
Gibson Biddle 0 0
Elena Verna 0 0
Casey Winters 0 0
Shreyas Doshi 0 0
Lenny Rachitsky 0 0
Melissa Perri 0 0
Marty Cagan 0 0