ZeroNoise Logo zeronoise
Post
AI skills need A/B evals, not install counts, and win/loss needs the buyer's own words
•
4 min read
• 311 docs
Hiten Shah on testing whether AI skills actually change agent output, ThoughtSpot on evals as the new PRD, Gamma's rebuild against generic AI output, practitioner advice on win/loss, and a PM job market that keeps tightening.

A skill being available doesn't mean the agent uses it

Hiten Shah's team has built 16 skills for competitive product marketing, and he asks a basic question: does a skill change the work once the model has it? He notes that Skills.sh reached one million skills and nearly 280 million installs in seven months, so availability "tells us very little about quality" . He cites a Vercel Next.js eval. The baseline agent passed 53%. Making a skill available left the score at 53%, partly because the agent skipped the skill in 56% of cases. Telling the agent explicitly when to use the skill raised it to 79%, and a compressed docs index in AGENTS.md reached 100%. He says one test can't settle how every skill should be designed .

His test is one PMs can reuse. Give the same model the same job and the same evidence, once with the skill and once without. Define what good looks like before the run, and keep the failures. If the model catches up, shrink the skill or remove it, and rerun the test whenever the model changes . He also writes each skill around the job it serves. For example, a CRM loss reason records what someone typed into the CRM, not why the buyer left. If nobody asked the buyer, the buyer's reason is unknown . In a demo, Claude without the skills treated an old source as recent and stated a guess as fact. With the skills, it dated every claim and noted what it couldn't see. The skills are MIT-licensed and work across Claude Code, Codex, Cursor and others .

ThoughtSpot: the eval is the new PRD

Francois Lopitaux (SVP Product, ThoughtSpot) told Mind the Product that "your eval system is almost becoming your PRD." It defines what good looks like and what to avoid, because a prompt-based product has no fixed UI to spec . Analytics answers have to be the same every time. So he uses LLMs only where they're needed and grounds them in a semantic layer that defines terms like "new customer" and "revenue". In his words, without that layer an LLM is "a new intern" that makes poor judgments . A few other points:

  • Teams are organized by product "track," not by feature, so the people closest to customers decide what to build next .
  • Pruning matters more now that code is cheap to generate .
  • He expects smaller teams and a lower PM-to-developer ratio, as deciding what's worth building becomes the bottleneck .
  • For a recent APM hire, he asked candidates for a Git repo and a video explaining the problem their code solves .

Product moves

  • Gamma 5. Gamma concluded that its output looked "too similar to all the other AI tools." It started over and rebuilt for visual variety and brand fidelity across presentations, docs, social assets, and graphics .
  • AI simulations. Lenny Rachitsky is following teams that use AI simulations to test product ideas and flow tweaks. He names Simile, Primitive Labs, Synthetic Users, Tenera, and Seldon, and asks how they've worked for people. No results are reported .

Craft: win/loss is not competitor research

A r/ProductMarketing thread argued that knowing a rival's pricing, features, and positioning doesn't tell you why a customer chose them . Advice from the thread:

  • Run unbiased win/loss interviews, not run by sales, and do 10–20 before looking for patterns .
  • Ask buyers what they expected to go wrong six months after signing. This tends to surface implementation risk, internal politics, and trust issues .
  • Commenters disagreed on price. One said it's almost always the reason . Another said C-suite buyers often pick more expensive competitors for ROI reasons .

Teresa Torres made a related point about records. Her transcripts kept showing that her confident memories were wrong. Notes and transcripts each add interpretation, and a record can be used to collaborate or as a weapon .

Job market

A 50-year-old senior PM with 22 years in tech has applied to more than 200 roles in nearly a year and gotten few interviews . Another commenter missed one of six technical requirements and was passed over for a role that had been open eight months. They described a "buyer's market" in which managers hold out for a perfect, cheaper PM . A third said the market never recovered from its 2022 peak and hiring has shifted toward senior roles . These are personal accounts, not market data.

Want personalized briefs on the topics you care about?

Create your own agent and get daily or weekly cited briefs on the topics you care about, shaped by the people, newsletters, podcasts, blogs, and channels you choose.