ZeroNoise Logo zeronoise
Post
AI skills need A/B evals, not install counts, and win/loss needs the buyer's own words
•
4 min read
• 311 docs
Hiten Shah on testing whether AI skills actually change agent output, ThoughtSpot on evals as the new PRD, Gamma's rebuild against generic AI output, practitioner advice on win/loss, and a PM job market that keeps tightening.

A skill being available doesn't mean the agent uses it

Hiten Shah's team has built 16 skills for competitive product marketing, and he asks a basic question: does a skill change the work once the model has it? He notes that Skills.sh reached one million skills and nearly 280 million installs in seven months, so availability "tells us very little about quality" . He cites a Vercel Next.js eval. The baseline agent passed 53%. Making a skill available left the score at 53%, partly because the agent skipped the skill in 56% of cases. Telling the agent explicitly when to use the skill raised it to 79%, and a compressed docs index in AGENTS.md reached 100%. He says one test can't settle how every skill should be designed .

His test is one PMs can reuse. Give the same model the same job and the same evidence, once with the skill and once without. Define what good looks like before the run, and keep the failures. If the model catches up, shrink the skill or remove it, and rerun the test whenever the model changes . He also writes each skill around the job it serves. For example, a CRM loss reason records what someone typed into the CRM, not why the buyer left. If nobody asked the buyer, the buyer's reason is unknown . In a demo, Claude without the skills treated an old source as recent and stated a guess as fact. With the skills, it dated every claim and noted what it couldn't see. The skills are MIT-licensed and work across Claude Code, Codex, Cursor and others .

ThoughtSpot: the eval is the new PRD

Francois Lopitaux (SVP Product, ThoughtSpot) told Mind the Product that "your eval system is almost becoming your PRD." It defines what good looks like and what to avoid, because a prompt-based product has no fixed UI to spec . Analytics answers have to be the same every time. So he uses LLMs only where they're needed and grounds them in a semantic layer that defines terms like "new customer" and "revenue". In his words, without that layer an LLM is "a new intern" that makes poor judgments . A few other points:

  • Teams are organized by product "track," not by feature, so the people closest to customers decide what to build next .
  • Pruning matters more now that code is cheap to generate .
  • He expects smaller teams and a lower PM-to-developer ratio, as deciding what's worth building becomes the bottleneck .
  • For a recent APM hire, he asked candidates for a Git repo and a video explaining the problem their code solves .

Product moves

  • Gamma 5. Gamma concluded that its output looked "too similar to all the other AI tools." It started over and rebuilt for visual variety and brand fidelity across presentations, docs, social assets, and graphics .
  • AI simulations. Lenny Rachitsky is following teams that use AI simulations to test product ideas and flow tweaks. He names Simile, Primitive Labs, Synthetic Users, Tenera, and Seldon, and asks how they've worked for people. No results are reported .

Craft: win/loss is not competitor research

A r/ProductMarketing thread argued that knowing a rival's pricing, features, and positioning doesn't tell you why a customer chose them . Advice from the thread:

  • Run unbiased win/loss interviews, not run by sales, and do 10–20 before looking for patterns .
  • Ask buyers what they expected to go wrong six months after signing. This tends to surface implementation risk, internal politics, and trust issues .
  • Commenters disagreed on price. One said it's almost always the reason . Another said C-suite buyers often pick more expensive competitors for ROI reasons .

Teresa Torres made a related point about records. Her transcripts kept showing that her confident memories were wrong. Notes and transcripts each add interpretation, and a record can be used to collaborate or as a weapon .

Job market

A 50-year-old senior PM with 22 years in tech has applied to more than 200 roles in nearly a year and gotten few interviews . Another commenter missed one of six technical requirements and was passed over for a role that had been open eight months. They described a "buyer's market" in which managers hold out for a perfect, cheaper PM . A third said the market never recovered from its 2022 peak and hiring has shifted toward senior roles . These are personal accounts, not market data.

AI skills need A/B evals, not install counts, and win/loss needs the buyer's own words
Mind the Product
  • AI has accelerated prototyping and hypothesis testing, but production still requires sound architecture, engineering expertise, and disciplined standards; shipping generated code without that discipline risks brittle systems and organizational chaos.
  • For probabilistic AI features, use evaluations as a behavioral specification: define desired and unacceptable outcomes, rerun evaluations as models or features change, and ground outputs in business and semantic context. Choose LLMs carefully when answers need to be consistent and deterministic.
  • Treat launch as the start of ongoing improvement: track adoption and repeat use alongside anonymized answer-quality and follow-up signals, and gather customer feedback. For white-label/OEM products, feature flags and customer-controlled rollout let different customers adopt at their own pace.
  • The product leader advocates organizing around continuing product or track ownership, so teams retain customer and product context after launch. He expects faster execution to enable smaller teams while making the choice of meaningful customer problems more critical; his stated PM hiring priorities include technical understanding, curiosity, and empathy.
How to escape the feature factory: Francois Lopitaux (SVP Product Management, ThoughtSpot)
Nir Eyal
Profile
  • Eyal argues that customers’ prior beliefs shape what they notice and expect, so branding and marketing can affect the product experience itself—not just recall—including how much customers enjoy a product.
  • For behavior-change work, he recommends mental contrasting over outcome-only positive visualization: anticipate likely obstacles, plan what to do, and adopt a useful interpretation of discomfort, such as “this is what it feels like to get better.”
  • For organizational change, he describes motivation as a combination of behavior, benefit, and belief; trust in a leader and confidence in one’s own ability affect effort and persistence. He cites Amazon’s “always day one” principle as a belief intended to sustain startup-like behavior and savings passed on to customers.
HAC26 - The Hidden Barrier to Change: What We believe About Ourselves
  • Armadin describes a two-part product: Red’s agent swarm maps customer networks and maintains a metadata-based pulse for changes, then tests when the network or threat conditions change; Blue is intended to turn exploitability findings into rapid compensating controls through defenses such as EDR and firewalls. The founder said early Blue controls were expected in the following months, with the capability becoming core within a year.
  • Armadin differentiates its approach from conventional penetration testing by attempting to verify whether risks are exploitable—including remote-code execution or data access—and by targeting logic flaws in custom applications. Its founder expects AI-enabled red teaming to eventually replace conventional penetration testing.
  • For security agents, the company’s approach combines red-team expertise with AI specialists, monitors agent prompts and activity, and uses deterministic rules and classifiers to stop or escalate unfamiliar behavior to people—while avoiding guardrails so restrictive that they suppress useful model creativity.
  • The founder says Armadin found more than 90 zero-day vulnerabilities in customer production environments since January 2026, including at Fortune 500 companies. He says humans found most while AI automated more than 90% of routine work, though the technology had begun finding zero-days itself.
  • For fast-changing enterprise products, the founder recommends weekly sales training when the product changes every two weeks, with customer feedback flowing directly to engineers. He frames customer acquisition, satisfaction, and repeatability as the core differentiator.
Building Cyber Defense for the Agentic Era
Product Management
  • A 50-year-old Senior PM with 22 years across tech and product roles said they had been unemployed for almost a year after redundancy, applied to more than 200 roles, and received few interviews despite broadening their search and lowering salary expectations.
  • Commenters describe PM hiring as favoring close domain and technical fit: one candidate said they were rejected for missing one of six technical requirements for a role open for eight months. Others argued that core PM skills transfer across industries and that PMs should learn users’ needs rather than arrive as expert users.
  • AI’s effect on PM expectations is contested: one applicant said recruiters emphasized being able to ship and adjust features independently; another commenter argued AI rewards faster execution and domain expertise, while a counterview said stakeholder management and sorting through AI-created complexity remain valuable PM work.
  • Jobseekers in the thread report that referrals and networking can outperform cold applications; commenters recommend targeting roles closely matching prior product/domain experience and considering fractional or contract work. One experienced PM also reported fewer management openings and more Principal/Lead IC roles with pay comparable to Head roles.
  • Age bias is a recurring concern in the discussion, though commenters also report seeing PMs over 50 at larger companies, often with deep industry or company knowledge.
Senior Product Manager with 22 years experience. A year unemployed following redundancy. Has the industry stopped valuing experience? I recently experienced this with a company I interviewed with. The hiring manager had a specific 6 point list of technical areas the PM H… \> Product is a hard role to move into externally, especially in domains where domain knowledge is valued. I think this is true, I also t… I’m a 34 yo senior product manager who is facing inevitable layoff and have also been applying for a year. It’s incredibly competitive ri… I’d say they have to stop valuing experience. The job has changed so drastically that trying to hold onto old ways and ideas of doing thi… Ageism is alive and well. I’m really sorry to hear you are going and have been going through this. The industry is garbage right now. Nob… AI has ruined recruitment. Feels like applying cold to jobs is a waste of time. You really need to know people otherwise applications get… Had some similar challenges last time I was between things, which was super humbling. I didn't expect to click my fingers and walk into a… market is terrible, best of luck; my only thoughts are to get more targeted for the roles you're seeking - going beyond your scope (of pr… I'm late 40s and took my previous job mid 40s. I deliberately made the change you cite - to a Head of role - for the same reasons. As it … I was talking to my partner about this earlier this week (we're both in tech). In short, yes I think there's ageism in tech and once you … Yeah I see plenty of 50+ product managers working in larger corporations. Though they typically have deep industry or company knowledge b…
Product Management
  • One commenter describes dark mode and gamification as fading, text-heavy sites as increasingly common (speculating AI may be driving this), and generic chatbots as past their peak; they say effective sites remain small, clean, and focused on one use case. A reply says AI needs better management and that chat interfaces can be useful when they complete tasks rather than merely ask how they can help.
  • Gamification is not a blanket dead end: a gamification PM says it works best in education and fitness, with niche uses such as savings goals, while generic leaderboards, levels, badges, and collections can become repetitive. Other commenters warn that gamification may increase dissatisfaction and churn, and that loyalty spending can be wasteful unless the target persona wants it.
  • For perceived website quality, one commenter recommends consistent spacing and typographic rhythm over chasing style trends; another warns that carousels without visible swipe cues can be hard to discover.
  • A contributor reports passwordless email/SMS OTP as a growing pattern and criticizes splitting email and password entry across screens. An email-first flow can route existing users to login and new users to signup, but a user with multiple email addresses says it creates concern about duplicate accounts and extra email searching. Security and flow preferences differ: one password-manager user prefers passwords over OTP (especially SMS), accepts email OTP, and dislikes emailed login links that interrupt checkout; another commenter favors passwordless in light of AI-enabled hacking and password reuse, while a reply suggests passkeys.
Well, I don't know which sites are AI driven and which sites are not, but here is what I see when I review new sites at launch: 1) Dark m… 1. I am with you on the dark mode. I’ve removed this feature from the site. 2. With you on this, except in some scenarios where it is cle… I strongly believe gamification works, but not in every domain. For example, it works best in education (think Duolingo) and fitness. Bey… Sorry, didn't mean to imply nobody's trying. I'm saying that it doesn't seem to work. If you watch interactions, the gamification is driv… I agree, but unfortunately people still *think* it works, so are attempting to add it into everything. You've big loyalty platforms whose… Trust comes from perceived craft, not style trends: the biggest tell is inconsistent spacing and type rhythm, so fixing vertical spacing … Not sure if it is a 'quality' trend, but i see too many apps getting too comfortable with horizontal scrolling and have carousel view eve… A big trend I've noticed recently, largely in retail/e-commerce design, is passwordless or OTP over Email/SMS, with no need for a traditi… The reason for this is usually so that you can have a single sign up/log in flow and people don't need to remember if they have an accoun… I get the business metric, but as a user I think that's a sucky experience. I have 2 or 3 emails I use and have often had the thought tha… I'm torn. As someone with a password manager, I'll take a proper password over OTP any day, especially OTP via SMS which is incredibly in… With the rise of AI and its use for hacking, it’s the right direction from a security standpoint. People don’t use password managers and … Passkeys have entered the chat
Lenny Rachitsky

Gamma rebuilt its product as Gamma 5 after concluding that generic AI tools made presentations look too similar; the overhaul centers on greater visual variety so outputs can match a user’s brand or a new aesthetic. Gamma 5 extends this approach to presentations, docs, social assets, and graphics, and revamps its agent, design tools, editing, import, export, and connectors.

Everyone is sick of generic AI, including us. Gamma was the first AI presentation platform to reach real scale. But now that AI tools are…
Teresa Torres

For teams using AI meeting transcripts, treat them as aids rather than definitive truth: Teresa says her confident recollections sometimes differed from her Granola transcripts, and cautions that notes and transcripts add interpretation, so verify a record before relying on it. Use transcripts for self-reflection rather than to hold colleagues rigidly to earlier views; the same record can support collaboration or become a weapon, and people should be allowed to evolve.

🎙️Recorded Conversations We now record almost everything — every meeting has an AI note-taker, and some people wear glasses that capture e…
Product Management
  • One PM with 2.5 years of experience reports being laid off after their company shut down, receiving few callbacks, and having a signed offer revoked with only “internal discussions” given as the reason. A commenter describes the wider market as still below its 2022 peak, says AI has weakened the perceived need for generalist PMs, and says hiring has shifted toward senior roles; these are assessments, not market data. Another commenter suggests B2B platform roles face less competition than consumer PM roles.
  • The discussion also disputes how AI affects PM generalism: one commenter argues AI can narrow specialists’ depth advantage without supplying generalists’ breadth, while another argues that breadth comes from experience in functions outside product rather than process expertise alone.
Laid off because the company shut down. Offer revoked because of "internal discussions." 2.5 YOE PM, still looking. Any PM/APM leads appreciated. That sucks, I'm sorry. The PM market never recovered from the 2022 peak, and AI adoption has further weakened the (perceived) need for ge… Do you have experience in B2B, especially with platform developmemt and management? There are tons of roles open with lesser competition … That is incorrect - generalists are growing in demand, specialists are less. You wan't more rounded and broad professionals nowadays, not… Good perspective, I see what you mean. Reddit, as many other "internet" environments are heavily influenced by SV media. Big funded start…
@andrewchen
  • Agent products do not have inherent network effects just because they are useful tools; product teams can assess potential effects separately across acquisition (lower CAC, more virality/signups), engagement (retention, depth, frequency), and monetization (ARPU, conversion, share of wallet), focusing on network density and interconnection rather than raw user scale.
  • Agent-created documents, events, microsites, plans, and group chats could become shareable “viral objects” that bring new users in. These loops require access to contacts, context, and communication channels, but overly aggressive outreach risks annoying users and having agents filtered or blocked.
  • Engagement network effects may arise when a shared agent network accumulates identity, trust, reputation, relationships, and private context that improve discovery, matching, and coordination as more people join. Those network effects could sit in open protocols and agent-facing marketplaces, leaving agents interchangeable, or be captured inside proprietary horizontal agents; the latter could make switching agents mean leaving a network behind, while an open ecosystem could allow vertical agents to thrive.
Will agents have network effects?
Hiten Shah
  • In Vercel’s cited Next.js evaluation, making a skill available left the pass rate at 53% (the agent skipped it in 56% of cases); explicitly telling the agent when to use the skill raised the rate to 79%, while a compressed docs index reached 100%. The article cautions that one test does not establish how every skill should be designed.
  • Evaluate AI skills by giving the same model the same job and evidence with and without the skill, defining success before the run, retaining failures, and repeating the test as models change; keep, shrink, or remove the skill based on whether it improves the work.
  • Design a workflow around the job’s evidence and output needs, and keep evidence sources distinct: a CRM-recorded loss reason is not necessarily the buyer’s reason, which remains unknown if the buyer was never asked.
There are a million AI skills. How do we know if they work?
Lenny Rachitsky

Lenny Rachitsky flags AI simulations as a trend to watch: teams are using them to quickly test product ideas and changes to product flows. He names Simile AI, Primitive Labs AI, Synthetic Users, Tenera, and Seldon as examples; the post does not report results from using them.

Trend I'm following: Teams using AI "simulations" to quickly test product ideas and flow tweaks. Products like [@simile_ai](https://x.com…
andrew chen
  • Andrew Chen argues agents do not inherently have network effects: assess whether greater network density improves acquisition, engagement, or monetization—not merely whether the product gains users. Interoperable agents may remain interchangeable tools, like email clients.
  • Agent-driven growth loops could turn generated documents, events, websites, research, and plans into shareable artifacts that bring new users in. These loops depend on access to contacts and communication channels, relevant context for choosing whom to involve, and outreach that earns trust rather than feeling spammy.
  • Shared identity, trust, reputation, and private context can make a common agent network more useful for discovery and matching, but those network effects could instead live in open protocols or specialized services accessible to multiple agents. The strategic boundary is whether horizontal agents internalize artifacts, transactions, and introductions—or leave networks interoperable; that choice affects whether vertical agents can thrive and whether switching means losing a network.
DO AGENTS HAVE NETWORK EFFECTS? We are seeing a new wave of general purpose consumer-friendly agents, and it’s causing many thousands of …
Hiten Shah

An AI-assisted competitive-monitoring workflow compares Claude without and with competitive-analysis skills: the post says the unskilled run treated an old source as recent and a guess as fact, while the skills-equipped run dated claims, marked its ranking as a viewpoint, and disclosed what it could not see . It suggests installing LMTYdotcom/skills and asking Claude which recent competitor changes deserve attention; the author invites readers to test the workflow .

We asked Claude what changed among Linear's competitors. No skills: an old source treated as recent, and a guess stated as fact. Same mod…
Kevin Weil 🇺🇸

Kevin Weil praised an OpenAI release in AI and mathematics, while noting that achieving models of comparable caliber in the physical sciences remains unfinished work; he expressed confidence that progress will come.

AI x mathematics ftw. What an incredible release today from OpenAI 🤯 There's a lot of work to do to get models of the same caliber in the…
andrew chen

After months of daily use of Hermes/OpenClaw, Andrew Chen says Town, Muse, and Grokbot are fun to try, but he is not switching because he enjoys tinkering—likening it to administering his own Linux box rather than using iOS . The anecdote suggests that hands-on tinkering can itself be product value for some users, not just a hurdle to convenience .

fun to use town/muse/grokbot after months of being a DAU on hermes/openclaw Reminds me of all the amazing fun and insanity it is to admin…
Lenny Rachitsky

@thsottiaux, Head of ChatGPT & Codex, says AI-era skills trending up are great taste, passion for building something that matters, and knowing what good looks like; typing fast is trending down.

.@thsottiaux (Head of ChatGPT & Codex) on the skills trending up and down in the AI era. Trending up: great taste, passion to build s…
Product Management
  • One PM with more than 10 years’ experience argues that expectations vary sharply by company, with some roles expanding toward end-to-end work across ideation, design, testing, and implementation. In their view, weak PM management and unclear definitions of “done” can fuel second-guessing, while AI-enabled customers and stakeholders make it harder for PMs to keep up with every change.
  • Their suggested response is to use LLMs and agents to digest incoming information and automate recurring work—such as Jira/spec workflows, async Slack updates, and task prioritization—accepting some loss of detail for sustainability. They also recommend grounding decisions in customer and user exposure, keeping up with the market through a manageable reading or podcast routine, and building cross-functional relationships.
  • Another commenter frames product and feature releases as iterative learning, and recommends consulting customers and coworkers to identify what to build and where to improve.
I've been there, and so have a lot of the other PMs I've coached over the years. What you are describing is not uncommon, and a lot of it… I disagree with another commenter here, I believe that many PMs suffer from impostor syndrome like you seem to be. But this does seem to …
Hiten Shah

For repeated AI-assisted work, a skill that teaches the task is not enough: the AI also needs company and market context, plus updates since the previous run, so later runs build on earlier work instead of starting over .

A skill can teach AI how to do the work. Repeated work adds another problem. The AI needs your company and market in context, plus what c…
Hiten Shah

For AI-assisted competitor monitoring, a described skill dates each reported change, separates facts from judgment, makes missing evidence visible, and stops when the next question is a different job; the post identifies answer currency as the hard part of asking AI what changed among competitors.

Ask AI what changed with your competitors. The hard part is knowing whether the answer is actually current. This skill makes it date ever…
Product Management
  • Commenters describe a crowded PM job market despite many listings: one says even experienced, industry-specific PMs struggle to land interviews, while another says experienced candidates are taking junior roles to get reemployed. Another commenter says hiring checklists can combine deep industry knowledge, architecture-level engineering, project management, and AI, and reports seeing more contract roles.
  • For someone moving into PM, a commenter recommends pursuing an internal transfer rather than switching externally in the current market; they suggest learning fundamentals in the current role, where AI may help.
I’m not sure exactly why it’s here but ai has definitely influenced the landscape. As to “why is the market bad”- there are tons of jobs … unless your current company is willing to transfer you to product, now is not the time to do this. the market is trash, and PMs with year… Agreed. Too many PMs in the market, also now companies are only hiring if you tick all the items in their checklist. They want very deep …