We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
Big Ideas
AI leverage compounds when teams productize their context. Sachin Rekhi’s “Compounding OS” shifts AI from individual productivity to team-wide productivity that improves with use. His playbook is to standardize an agentic platform, build a shared skills library or marketplace, and connect AI to repositories, the design system, data layer, and code so it can contribute to a growing company brain. Apply this by taking one recurring PM workflow and turning its prompt, context, outputs, and review criteria into a reusable skill; improve it from failures instead of leaving expertise in one person’s chat history.
Reliability must be specified, not inferred from demos. Hiten Shah’s warning is practical: one great run proves only that a task could work, while a system that succeeds 90% of the time has roughly a 59% chance of succeeding on all five independent runs. Build evals from normal, messy, and known-trouble examples; define what “good” means; rerun the same set after changing prompts, models, context, or tools; and save important failures as regression cases. For agents, evaluate state changes, rule-following, recovery, and cost—not only the final answer—because the full workflow, not a public benchmark, owns the behavior. Retry, refusal, escalation, and help-seeking thresholds are therefore part of the product specification.
Tactical Playbook
Make stakeholder input decision-grade. Shreyas Doshi’s framework rejects seniority as the shortcut: input quality depends on who you ask, how you frame the problem, whether you know the real goal, whether you listen without filters, and whether you understand the reasoning behind the input. Before asking for opinions, write the decision and desired outcome; select people with relevant context; ask for reasoning and constraints rather than a vote; then replay what you heard and identify what evidence would change the choice.
For cross-sell, start with a trigger—not a campaign. Validate the adoption journey with current adopters and comparable non-adopters, then define the buyer, problem, trigger, value proposition, and positioning; confirm whether the buyer is even the same role across products. The first test should target one observable trigger and one adjacent use case. Assign the hypothesis and message to product marketing, the conversation to sales or CS, and the activation event to product; use a phased rollout or holdout, tracking exposure, activation, time-to-expansion, and support burden.
Case Studies & Lessons
Skydio’s lean PM model ties platform investment to customer outcomes. Its roughly 20-PM organization combines a core platform team with vertically focused PMs; product leaders must get field signal, inspect data in “product data day,” and pair with engineering. Its DFR business reports 55,000 911 responses per month and 25 million Americans within two miles of a dock. The outcomes dashboard tracks whether the drone arrived first, response time, useful information, and calls where the drone prevented an unnecessary officer dispatch, with customers reporting the data. After years of R&D without delivered value, Skydio made DFR the number-one goal; only after X10, the dock, and remote-operations software created customer value did it expand investment elsewhere. Make the wedge’s outcome metric the gate for adjacent bets—not internal enthusiasm or platform completeness.
Career Corner
Nontraditional paths can compound into product leadership. Skydio’s product leader has a history degree, military experience, self-taught software and drone experience, and entered the company by building customer success from zero before moving into product. The actionable signal for career changers: seek roles where you can own a customer problem and show operational and technical learning, not just acquire a PM title.
Tools & Resources
- Evals 101: Hiten Shah’s free live session is Friday, August 14 at 10 AM PT.
- Compounding OS webinar: Sachin Rekhi’s free product-leader session is August 20 at 10 AM PT and covers the three-step playbook.
Blaine (OpenAI designer) on the ChatGPT–Codex merge and the new product design plugin: built to make ChatGPT a daily driver for designers and non-designers who need design; the team leaned into model strengths — computer use, image generation for interface design, and code — rather than forcing Figma mastery, so it focused first on faster code prototyping and later on AI-assisted ideation.
- Product/design work shifts from pixels to defining good output: deciding which evals grade success, how to measure, and heuristic evaluation of output format and structure.
-
LLM-assisted ideation behaves like
crazy eightsat volume: an LLM can produce 80 ideas in 8 minutes vs 8 designs in 8 minutes for humans; volume beats preciousness, and the sticky-note→lo-fi→hi-fi pipeline can be compressed (e.g., come back from lunch to ~80 mocked-up ideas). Method: poke with what-if changes and let the model run structured exercises, then harvest one or two inspiring outputs. - Treat evals as synthetic user research: posit what people will do, then judge whether the resulting experience is good or bad. The hard part is envisaging the prompt space — after hill-climbing on internal examples, an internal user immediately exposed a marketing-site use case that failed badly. Keep adding throwaway eval sets beyond a golden regression set, define both success and failure for the current work, and share early to check assumptions.
-
At very large scale, do not try to cover all user behaviors: design general mechanics and learn how they fail and succeed at the extremes; spend as much effort defining failure (e.g., know when the product
declares bankruptcy) as defining success, sample widely, and cushion bad outcomes. - Without traditional user access, the team ran experience-and-vibes research: internal design-team interviews, Twitter/Reddit community forums, internal tests, eval cases built from observed failures (e.g., icons), a couple of companies testing days before launch, and post-launch review of prompt classifications rather than raw prompts. Any information is good information.
- Expect the second 80%: prototypes arrive in a day or a weekend, then production work is grueling. OpenAI does little roadmap planning; when the ChatGPT desktop/Codex harness merged into the web product, the team had roughly three weeks to rewrite and re-test everything in a new environment, and priorities shift daily or weekly.
- Alignment at speed comes from a company-wide rallying cry rather than OKRs or roadmaps: anyone should be able to state the one or two most important things the company is doing, and everyone contributes to that from wherever they sit. This avoids cross-team competing priorities; the interviewer notes it works best in exponential growth and gets harder once growth plateaus.
- With AI agents, individuals are far more empowered (e.g., submitting bigger PRs directly), but teams risk operating in their own little worlds; keep sharing/show-and-tell rituals — the best group creative thinking is working alone and then reflecting together, which avoids groupthink.
- Career advice for AI-era product teams: titles no longer make sense, so focus on who you are working with, what you are doing, and where you fit on the team; like a jazz band, everyone has a specialty but must know the tunes and share responsibility for the output. Experience helps but is not a prerequisite — junior people without baggage can produce some of the best work; build taste by making many things (a thousand pots).
-
Demonstrated in the ChatGPT desktop app, the product design plugin saves preferences and known sources, browses the web (opening ~10 tabs), builds and organizes FigJam mood boards, produces 10 distinct visual styles, mocks key screens and flows, accepts markup/removal corrections, and generates working React prototypes with shareable links. The goal is many throwaway artifacts (
something every 10 or 15 minutes) so teams stay less precious and iterate faster.
Career advice from YC partner Gary: don't chase "what's hot"; the right question is what you're interested in and know uniquely. He regrets abandoning web programming (where he had ~5 years of experience on everyone) to follow the crowd to Windows Mobile, and declining a $70K-check co-founder invitation from his fraternity brothers — the group that became Palantir — for a Microsoft promotion, calling it "working backwards from the map instead of looking down at the territory" and a "$2 billion to $4 billion mistake at this point" . Earnestness means trusting your own direct experience over counter-narratives , and agency/taste is a practice built by doing, not inborn — he calls himself a late bloomer whose mistakes, introspection, and changed choices show it can be developed .
AI-era product work: the coordination layer "a lot of product managers historically did" — writing specs, marshaling engineers, QA — can now be absorbed by agents and "should not be human"; code is "no longer precious" . His learning tactic for the moment: just use every model and build trivial things as a "chassis" for building intuition .
Skillification methodology: do a business process once, "skillify" it into a reusable markdown file + code + tests on a cron job; any future failure is just a permanent bug fix — "a markdown file is an employee" that does the job perfectly every time, applicable to sales, marketing, support, and every business process. He claims companies now go from zero to ~$15M ARR in ~4 months with two or three people and a few hundred skill files .
His own PM automation practice (GStack): building agentic systems means finding bottlenecks and instructing an agent to create software or a markdown file that removes them; he "automated how [he] thought about PM," adding an engineering "edge manager" skill enforcing unit and end-to-end test coverage plus a QA loop, until his remaining job was blackbox verification .
Case study — cross-team dependency hell at Microsoft: as a Windows Mobile PM (~2003), he needed integration from the Windows team, which ignored emails and bugs (wouldn't even mark "won't fix"); he and his PM mentor went over "with a baseball bat" — seven levels of bureaucracy lay between them and the owner, and four hours went into a P3 bug blocking ~1,000 users. His lesson: a cron-run skill doing competitive research and surfacing blockers/dependencies could compress ~six-month releases to a day or two — "an org like Microsoft can't. But a startup can. And every startup must" .
Agentic management: the most interesting loops are "business loops" — loops whose output changes how the business operates — and YC is organizing around them . Example: Pedro (of Brex) open-sourced "crab trap," an agent that watches all OpenClaw network traffic, to safely use agents in a regulated fintech; agents read meeting transcripts from direct reports (two levels down) to surface what's broken and who's in conflict, letting him join any meeting with perfect context. Gary frames this as solving the failure mode where a business outgrows one person's head — agents + memory + retrieval give ground-truth visibility — and as productive conflict, "the search for truth," unbundled from human emotion .
Nir Eyal's Indistractable framework: the opposite of distraction is traction — any action done with intent; anything not what you planned, including 'productive' pseudo-work like checking email before a big task, counts as distraction . Most distraction originates from internal triggers — uncomfortable emotional states (loneliness → Facebook, uncertainty → Google) — rather than external pings; because all behavior is driven by a desire to escape discomfort, 'time management is pain management' .
Four implementation steps: (1) master internal triggers — e.g., reimagine a disliked task by focusing on it more intensely and adding variability, per Ian Bogost's Play Anything ; (2) make time for traction with a timeboxed calendar — turn values into time, schedule leisure and relationships explicitly, and measure success as working on a task as long as planned without distraction rather than finishing; Eyal says this beats to-do lists, which reinforce a failure identity (free seven-day schedule template at nirandfar.com ); (3) hack back external triggers — device settings and meetings, with colleagues the #1 workplace distraction (80% of survey respondents) ; (4) prevent distraction with pre-commitment pacts — effort, price, and identity .
Don't rely on willpower at the moment of temptation: 'the antidote to impulsiveness is forethought' — plan systems in advance . Eyal rejects digital detox as a temporary-diet fix; distraction's root cause is internal discomfort, and blaming tech companies creates learned helplessness . For product context, Eyal is a behavioral designer applying habit-formation across medical, fitness, and education products (e.g., FitBod, Kahoot, NYT) and wrote Hooked (building habit-forming products) and Indistractable (breaking distracting habits) ; he argues overuse is a problem for distracting products, not enterprise software .
- A bootstrapped BI product spent three years with positive feedback from workers, CISOs, and peers but zero contracts; feature requests were built, the goalposts shifted, and the founder was running out of money .
- The pattern of 'everyone likes it, nobody signs, endless feature requests before purchase' indicates users, not buyers, in the room: a user enjoys the demo; a buyer has a budget line and a problem their boss is asking about this quarter. Feature requests are a polite 'not now.' The useful questions are 'who else has to say yes?' and 'what breaks for them personally if they do nothing this quarter?' — no answer to the second means no deal .
- To draw brutal feedback instead of politeness, ask 'What would it take for you to buy right now?' and 'What am I missing? If you knew you couldn't offend me, what feedback would you give me?'
- If users aren't willing to stake relationship capital to propose the product, it's probably not good enough; internal advocacy needs an urgent problem, budget, and the right buyer . For BI, a CISO is a blocker, not a buyer; signers are whoever owns an embarrassing number — VP revenue ops, CFO, head of supply chain .
- Don't build pre-contract feature requests: when told 'if it had X we'd buy,' say 'great, I'll build X — here's a contract that starts when it ships'; most will go quiet, showing there was no deal . A small paid pilot turns polite interest into a real decision ; charge for it — a few thousand dollars for a scoped six-week pilot with a named success measure — because free pilots are demos with extra steps .
- Zero contracts is likely a positioning/buyer-urgency problem, not product quality: sell one painful outcome to one specific buyer rather than the whole BI product . Pick a recurring decision with a deadline and visible owner (regulatory report, audit exception) and ask what a late/wrong answer costs; for smaller buyers, sell a paid diagnostic on one dataset and disqualify accounts that can't name a consequence worth fixing . Frame value by segment: cost reduction for existing BI users, revenue growth for non-users .
- Enterprise prospects wanted mature software, not a product that 'may need an iteration or a feature added' ; risk management was often treated as a compliance checkbox rather than a financial benefit .
- Adoption barriers: users avoid learning another tool when Excel is 'good enough,' won't pay out of pocket, and corporate license procurement takes months — a bottom-up vs. top-down adoption consideration .
- For BI, the brand's emotional job is often 'protection from being the person blamed for a number,' not insight; ask the warmest contact what they'd tell their boss to get this signed, then use their words as the pitch .
- Bootstrapped and out of runway, don't try to build a repeatable sales motion; sell the outcome as a done-for-you service using the product to get one paid customer — it buys runway and produces honest feedback .
- Resource: 'Founding Sales' by Peter Kazanjy was recommended as useful reading for founders learning to sell .
Alden Jones, VP of Product at Skydio, runs a lean ~20-PM function across hardware, software, and autonomy. The team pairs a core platform org (hardware, core software, autonomy, comms, product ops) with vertically integrated PMs focused on four use-case segments (public safety DFR, ISR/defense, security, inspection) . PMs have no travel budget — 'if you need signal, go get it' — and must see robots in the field; the company runs 'product data day' to review customer data . PMs pair with engineering leaders; design is a core partner because autonomous products must earn user trust . For deep features, platform engineers 'strike' — embedding with a vertical team for ~6 months, e.g., to build Pathfinder for DFR (similar to forward deployed engineers) . Software ships fast; hardware programs run ~2 years (R10 indoor drone ~18 months), with revisions in between . Example deep cross-team bet: NightSense, active IR illumination so autonomy works at night — '12 hours of robots not doing what they need to do' otherwise .
DFR case study: Skydio replaces $3,000/hour helicopter support and one-line radio dispatches for 911 response; pre-sales, they map customers' 911 calls to decide where to place docks . Success is tracked in an outcomes dashboard: did the drone get there first, response time, did it provide useful information, and how often an officer was not dispatched because the drone showed a nuisance call — with customers self-reporting these metrics . Current scale: 55,000 911 calls responded per month; 25M Americans within 2 miles of a dock (up from 12M ~8 months earlier); in physical security 90–95% of alarms are false, and a drone can be on site in ≤50 seconds .
Prioritization playbook: after years of R&D without delivered value, Skydio ruthlessly made DFR the number one goal; once X10 + dock + remote-ops software created product-market fit, it 'earned the right' to invest in defense, security, and inspection . A $3.5B five-year investment commitment is revenue-funded, not debt, targeting US/allied supply chains .
Autonomy and regulatory strategy: Skydio deliberately earned FAA trust via autonomy, aviation-safety experts, and daily collaboration, obtaining groundbreaking waivers first; that trust enables ops like one pilot flying four drones and autonomous response, and full fleets are 'very near term' — 'not 10 years out' .
Career path: Jones is VP Product without an engineering degree — history degree, military truck officer in Iraq, Fortune 500 roles, self-taught software/drones at a telco, then Skydio's first customer success leader (built the function from zero) before moving into product; Skydio competes with Google and AI labs for autonomy and mechanical/electrical engineering talent .
Lenny Rachitsky published a newsletter titled "How to make people care about your startup" https://www.lennysnewsletter.com/p/how-to-make-people-care-about-your.
Lenny Rachitsky's podcast with Cursor Head of Talent Adam Ward covers the playbook for building high talent-density teams: why the traditional recruiting funnel (the 'funnel of doom') guarantees mediocre hiring, a three-step playbook for hiring the top 1%, the rise of the forward-deployed engineer, the current 'tale of two cities' talent market, the biggest mistake founders make with their first recruiting hire, and the worst sourcing question to ask . The key prerequisite he offers: 'If you can't objectively describe what great looks like for a role, you're not ready to hire for it' .
On whether you benefit from outside input, it's not about your level — it depends on who you ask for input, how you frame the problem, whether you truly know your real goals, whether you're actually listening to their input without filters, and whether you understand the logic behind their input . A world-class coach or advisor can help you work through all of these factors with a higher likelihood of finding what works for your specific situation .
- Hiten Shah argues that judging AI work one output at a time is the wrong standard: a great first result proves a task is possible, not that the system reliably works — "it could work" became "it works" . He calls "it worked" possibly the most dangerous sentence in AI right now .
- Eval methodology for AI features: (1) start from the job — collect examples of real work, including normal, messy, and already-troublesome cases; (2) write down what "good" means, including judgment-based quality bars (research that misses the source that mattered, a support agent answering without solving the customer's problem, a coding agent making tests green while leaving the codebase worse); (3) rerun the same examples after changing prompt, model, context, or tools; (4) when something important breaks, save the case so the next version must prove it handles it — the eval set becomes a memory of what the system has learned not to screw up .
- Reliability math: a system that succeeds 90% of the time, with that rate holding independently across five runs, has only a ~59% chance that all five succeed — a 90% system can still fail often enough to feel broken .
- Public benchmarks are not your product: once a model is wrapped in your system (prompt, retrieval, tools, memory, permissions, routing, retries, orchestration), you own the behavior; no public benchmark knows your support policy, customers, codebase, or research standards, so the eval must come from the job itself .
- For agents, output quality alone is insufficient — evaluate the work: the final answer can look correct while the system took unnecessary actions, changed the wrong state, ignored a rule, recovered from its own mistake, or spent ten times expectations getting there .
- AI lets teams build before specifying what good behavior means; evals force that missing specification into the open, and "knowing when the work is actually good is becoming part of the product" .
- Hiten Shah is teaching the full process in a free live session, "Evals 101: How to Know If Your AI Actually Works," Friday, August 14, 10 AM PT .
- Anxiety is common in PM and not a disqualifier: many PMs are introverts or anxious; what matters is clear communication and not melting in conflict, and keeping anxiety internal.
- Anxiety can be managed with preparation and repetition: one PM went from needing two hours of prep for a small customer call to holding three calls without worry.
- Practical product experience (e.g., 3 shipped products) can outweigh formal credentials in some eyes, but breaking into PM without a degree or prior team experience is a grind; startups and tiny companies may take a chance if you show real impact.
- PM is not an entry-level role; it requires prior experience in something. Side projects alone are nice but insufficient—candidates must articulate strategy and business approach. Degree requirements create an education barrier that can block candidates before reaching hiring managers.
Shreyas Doshi: the quality of input you get depends not on the other person's seniority but on 1) who you're asking for input, 2) how you're framing the problem, 3) whether you truly know your real goals, 4) whether you're actually listening to their input without filters, and 5) whether you understand the logic behind their input .
Scott Belsky argues the more people you consult for advice, the more likely you’ll make a bad or average decision: (1) people are, on average, risk-averse and, in quantity, blunt your boldness; (2) people may hate or envy opportunities they lack; (3) people tend to mix cynicism (bad) with criticism (good) . Shreyas Doshi adds that most people — even those perceived as higher status — are primarily trying to come across as intelligent or competent, and this need to impress isn’t usually compatible with effective advice .
Hiten Shah announced a Friday 10 AM PT breakdown of how to test AI work ("Evals 101" at https://www.hiten.com/evals-101). The core framing: changing a prompt and seeing a better output does not prove improvement — evals are what let you tell what actually got better vs. a better run .
Physical Intelligence's founder — who founded the company two years before the talk — shared strategy and career guidance from building general-purpose robot models :
- Teams should start from an open-source generalist policy (e.g., the PI0/PI5 models) and fine-tune it rather than scaling per-site specialist models; the exception is heavily constrained environments like surgical robots in operating rooms with no internet and weak GPUs, though local inference on a workstation remains possible .
- On career paths: a PhD is most valuable for learning to handle uncertainty and pick problems; it is not required for the engineering-heavy work (robot software stack, hardware, ML and data infrastructure), and the founder — who chose industry first — said he would still choose a PhD in hindsight .
- Concrete route into robotics from another background: buy a cheap robot, fine-tune an open-source model on it, share the result, and cold-email; the founder cites a former algorithmic-trading/legal hire who joined Physical Intelligence this way .
- A 'ChatGPT moment' for robotics is unlikely to spread as fast because distribution requires physical devices, but reaching ChatGPT-like capability is "very much on the horizon in the next few years" .
- Market signals: Waymo passed 250,000 weekly autonomous rides about a year before the talk, cited as evidence that ML systems can operate autonomously in the physical world ; two YC companies (Ultra, Weave) have post-trained Physical Intelligence models for real deployments in laundry folding and warehouse packaging .
- A single good AI output proves capability, not reliability: "it could work" gets mistaken for "it works," making overtrust easy .
- Even a 90% system success rate means only ~59% chance all five independent runs succeed — a system can score 90% and still feel broken, so measure the work, not just the output .
- Build evals from the job: collect examples including normal, messy, and known-trouble cases, define "what good means" as the quality bar, then rerun the same set after every change (prompt, model, context, tools, workflow) .
- Save failures into the eval set as regression cases; the eval set becomes a memory of what the system has learned not to screw up .
- Public benchmarks can't judge your product because your system includes model, prompt, retrieval, tools, memory, permissions, routing, retries, and orchestration — evaluate the full work, since agents can look correct while taking wrong actions, changing wrong state, or overspending .
- AI lets teams build before specifying requirements; evals force the missing specification of good behavior, making "knowing when the work is good" part of the product itself .
ElevenLabs (11 Labs), an AI voice synthesis and cloning startup (synthetic voiceovers, consent-based voice cloning, near-real-time translated dubbing preserving an actor's performance), is valued at $11B after its Feb 2026 Series D, more than triple the ~$3.3B valuation Disney had been working with a year earlier; Disney, spending ~$500M/year on voice acting, dubbing, and localization (growing 12% annually), weighs acquiring it as its largest AI bet to date . The evaluation framework: assess strategic fit (does it solve a real need; how rivals and Disney's competitors use similar tech), business case (Disney cost savings and whether ElevenLabs' standalone revenue supports $11B), and risks (talent retention, quality, integration time, regulation/IP) . Quantifying with 60% of spend automatable at 90% cost reduction yields $270M in year-one savings and roughly a 40-year payback on $11B; a pure cost lens is the worst-case view because it omits ElevenLabs' non-Disney revenue, 12%+ spend growth, IP/ownership value, and possible revenue/quality differentiation . The three options are acquire (most expensive; full IP/team ownership), license API (cheapest; no ownership, vendor-dependent), and build in-house (full ownership; high execution risk); the candidate recommends licensing the API for now, given technology, regulatory, and team unknowns . Interview-tactic takeaways from the coach: clarify how an unfamiliar company earns money, anchor frameworks with case-specific data points, state a clear hypothesis and next steps, define the quant problem before computing, contextualize intermediate numbers, and structure brainstorm responses before answering .
Andrew Chen flags the phrase "I asked my friends and they said X" as a red flag in decision-making, arguing it treats friends' opinions as a vote rather than reasoning from first principles . Scott Belsky explains why consulting many people makes bad/average decisions more likely: people are typically risk-averse and blunt boldness, may hate/envy opportunities they lack, and tend to mix cynicism (bad) with criticism (good) . These points are useful for product managers weighing stakeholder or user input against product conviction.
Hiten Shah observes that much AI improvement still relies on informal evaluation — change something, inspect a few outputs, decide it got better — and challenges practitioners to articulate how they actually know a change to a prompt, skill, workflow, or agent improved things .
One good AI result proves possibility, not reliability — a useful caveat for PMs assessing whether an AI feature is dependable enough for real product use.
In an a16z interview, YC's Gary Tan argues that the organizational coordination layer historically filled by product managers should be automated with AI agents — mid-level bureaucracy becomes agents, while executives set direction and individuals own execution . He illustrates the problem with a PM case from his Microsoft Windows Mobile days: a P3 bug blocking ~1,000 users went unfixed because the Windows team ignored emails and bug reports, forcing him and his PM mentor to escalate physically with a baseball bat — evidence of how cross-team friction and human cognitive limits (7±2) stall product delivery . For building AI-native product ops, Tan recommends converting any business process into a skill file (markdown + code + tests) run on a cron; a 'markdown file is an employee' that executes flawlessly every time, improving via bug-fix iterations, applicable across sales, marketing, and support . He advises PMs to build trivial things with every AI model to build intuition and to automate their own PM workflow (as he did with GStack), since agency and taste now outweigh traditional spec-and-QA execution .
𝕏 post by @hnshah
“It worked” might be the most dangerous sentence in AI right now.
We’re producing way more work with AI and still judging most of it one output at a time. https://x.com/i/article/2087420669026021376 (opens in new tab)
- Hiten Shah argues that judging AI work one output at a time is the wrong standard: a great first result proves a task is possible, not that the system reliably works — "it could work" became "it works" . He calls "it worked" possibly the most dangerous sentence in AI right now .
- Eval methodology for AI features: (1) start from the job — collect examples of real work, including normal, messy, and already-troublesome cases; (2) write down what "good" means, including judgment-based quality bars (research that misses the source that mattered, a support agent answering without solving the customer's problem, a coding agent making tests green while leaving the codebase worse); (3) rerun the same examples after changing prompt, model, context, or tools; (4) when something important breaks, save the case so the next version must prove it handles it — the eval set becomes a memory of what the system has learned not to screw up .
- Reliability math: a system that succeeds 90% of the time, with that rate holding independently across five runs, has only a ~59% chance that all five succeed — a 90% system can still fail often enough to feel broken .
- Public benchmarks are not your product: once a model is wrapped in your system (prompt, retrieval, tools, memory, permissions, routing, retries, orchestration), you own the behavior; no public benchmark knows your support policy, customers, codebase, or research standards, so the eval must come from the job itself .
- For agents, output quality alone is insufficient — evaluate the work: the final answer can look correct while the system took unnecessary actions, changed the wrong state, ignored a rule, recovered from its own mistake, or spent ten times expectations getting there .
- AI lets teams build before specifying what good behavior means; evals force that missing specification into the open, and "knowing when the work is actually good is becoming part of the product" .
- Hiten Shah is teaching the full process in a free live session, "Evals 101: How to Know If Your AI Actually Works," Friday, August 14, 10 AM PT .