We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
Big Ideas
AI adoption is becoming an operating-model problem, not a software-procurement problem. Sachin Rekhi reports a widening gap: average teams are seeing a 20–30% increase in engineering velocity, while the strongest teams are seeing 2–3×. His four-step playbook is to build a shared Compounding OS, set an AI-fluency floor, redesign the product process so agents work alongside product, design, and engineering, and use behavior-change principles rather than training and mandates alone. The useful leadership audit is whether the team’s system, skills, workflow, and habits are changing—not merely whether people have access to AI tools.
“AI PM” is not one job. Jonathan Evens separates modeling PMs, who define model behavior, capabilities, and acceptable error rates, from AI feature/product PMs, who apply AI to a domain workflow and design the surrounding experience—the last 20% that makes a general model work for a use case. Using AI to improve a PM’s own knowledge work is simply PM productivity, not a separate role. For AI features, establish product principles and existing north-star outcomes first, then add side-by-side output comparisons and behavioral signals to the evaluation system.
Measure successful work, not hours saved. A controlled study compared 108 people using agents with 110 doing the same context-heavy tasks without them. Task success rose from 59% to 93% while completion time fell 38%, or roughly 2.5× more successful outcomes per unit of time; the report cautions that it does not yet show how those gains translate into business outcomes.
Tactical Playbook
Debug AI quality at the source. Teresa Torres describes a customer finding a flat branch in an AI-generated opportunity tree; the team spent three weeks building four evaluation metrics and testing 16 variations rather than applying a quick prompt fix. The operating rule is to combine evals, guardrails, and orchestration; calibrate judges against production data; trace upstream fixes for downstream regressions; and consider an agent that audits its own work. A practitioner’s more operational checklist includes cost per token, latency, payload size, structured-output errors, first-pass rate, real-user failures marked mandatory to detect, and broad regression runs—while remembering that a better eval score can still produce worse UX if it makes the product materially slower.
Case Studies & Lessons
Lightfield used customer pull to escape a successful but unsatisfying product. Its presentation product reached about 2 million users per month, but the founders stopped it because they could not see discerning professionals finding it indispensable; they also judged that missing context about the presenter, audience, and relationship—not general reasoning—was the core limitation. They found 12 B2B pilots among sales and marketing users, followed requests from decks into research, lead qualification, and account expansion, then discovered that conflicting CRM, call-recorder, and warehouse data was the deeper problem. A first go-to-market assistant had daily users but no pricing power because it did not own the underlying data; after restarting around a CRM, 10 startups used the rough product daily and sent feedback roughly every two hours. The lesson is to distinguish reach and usage from indispensability, pricing power, and ownership of the workflow’s core data.
Hinge demonstrates an outcome-led consumer strategy. Its North Star is great dates and getting people off the platform, not maximizing engagement; fewer than 15% of users pay, and monetization is reserved for constraints or accelerators while the free experience remains “sacred.” Hinge segments daters, identifies their problems, forms hypotheses, and does not ship a change unless it increases the chance of two people meeting in real life.
Career Corner
For senior interviews with founders or executives, use the final five minutes to ask questions that expose the real role: What will you continue owning after hiring the head of product? What is the one superpower you want this person to have? What are the one or two things that truly matter enough to require an 11/10? The framework is aimed at senior roles; it is not designed for a PM3 reporting to a GPM.
Tools & Resources
All Things PM is a current community-built resource derived from 604 PM job postings across 95 companies and 137 career boards, with an updating AI-PM curriculum, a 205-concept knowledge graph, and interview/resume tools. Treat it as a map of hiring language rather than a definition of the job: its creator acknowledges that postings can miss stakeholder management, judgment, and politics, and a commenter argues that JDs often diverge from actual workflows.
- Treat product assumptions as revisable models. Nir Eyal defines beliefs as convictions open to revision based on evidence and as tools rather than truths. For a team stuck on a customer, problem, or stakeholder interpretation, his inquiry method is to identify where the difficulty is, state the assumed cause, ask whether it is true and absolutely true, examine the effect of holding the belief, imagine operating without it, and generate alternative interpretations. The goal is a portfolio of perspectives rather than a forced replacement belief, giving the team more freedom to choose a useful response.
- Design sustained motivation around behavior, benefit, and belief. Eyal argues that persistence is critical to reaching goals and that long-term motivation requires knowing what to do, why it matters or what benefit it provides, and believing in one’s ability or in the leadership behind the effort. For product adoption and internal execution, make all three explicit instead of relying on incentives or information alone.
- Treat expectation-setting as part of the product experience, but validate behavioral outcomes. Eyal describes identical wine being judged differently when framed as expensive versus cheap, and golfers performing better when told a putter was used by a famous golfer. He attributes the improvement to behavior changes such as greater relaxation and focus—not to positive thinking or “manifesting” alone. Product teams should therefore test how pricing, positioning, and launch framing affect perceived value and usage, while separating expectation effects from actual product capability.
- For B2B SaaS, repeated use of the same core workflow is a stronger intent signal than login frequency; one practitioner cites using that workflow three times in a week, while daily logins may reflect curiosity. Repeated feature depth, organic usage-limit hits, and teammate invitations that lead to actual teammate activity are also prioritized signals.
- A lightweight operating model is to score users from 0–3 on these signals and manually review users scoring 2+ instead of building a full lead-scoring model. Interpret limit events in trial context: reaching a cap on day 2 of a 14-day trial may indicate faster workflow adoption than reaching it on day 13.
- In a small beta, combine behavioral signals with direct conversations rather than relying only on product data. For short B2C trials, onboarding completion may be the only practical intent signal; one three-day-trial example improved trial starts from about 2.5% to 5% and weekly MRR from $60 to $300 by showing the paywall after onboarding and removing the “continue for free” option.
Use case: In interviews for senior product leadership roles with founders or senior executives, use the limited end-of-interview question period to ask sharper questions; this framework is not intended for a PM3 reporting to a GPM and is less applicable when the interviewer is a GPM at a FAANG company.
- Ownership: Ask, “What aspects of the product makes sense for you to continue owning after you hire the head of product?” This probes founder or executive ownership boundaries, which can reveal a source of future conflict, and signals that the candidate is optimizing for the company rather than personal ownership.
- Role-defining strength: Ask, “What is the one superpower you want this person to have?” instead of asking for an ideal-candidate profile; the sharper prompt is intended to make the interviewer think and provide more useful information for the next conversation.
- Highest bar: Ask, “What are the one or two things that truly matter that this person needs to score an 11 out of 10 on?” rather than the generic “what does success look like?” question.
- Outcome-based North Star: Hinge measures success through great dates and getting users off the platform rather than maximizing engagement; it does not ship work unless it believes the change increases the chance that two people meet in real life. A transferable implementation pattern is to define the customer outcome beyond the product, research the problems blocking it, form hypotheses, and filter prioritization through that outcome.
- Segment-led discovery and constraint-based monetization: Hinge’s culture and consumer-insights team studies different dater situations, identifies their problems, and develops hypotheses before selecting solutions. Its free experience is intended to let everyone succeed; it monetizes features that break constraints or accelerate success, especially where scarcity would lose its value if offered universally, while fewer than 15% of users pay.
- Strategy-driven product organization: Hinge treats org structure as a reflection of the company’s problem and strategy, reviewing the organization annually as strategy is updated. It deliberately brings engineering, AI, data, product, design, and research leaders into prioritization debates; a lack of disagreement over product-versus-technology investment is treated as a warning sign, while infrastructure preparation must be balanced against user-facing improvements.
- AI as a product-learning accelerator: Hinge’s biggest current AI benefit is a faster feedback loop: PMs can prototype near-final experiences connected to APIs and back-end systems, put them in users’ hands, create visualizations instead of relying on long feature specifications, expand user research, and begin translating native iOS work to Android.
- Guardrails for AI-assisted building: Hinge lets PMs and designers use AI for prototypes, faster learning, and communicating ideas, but engineers still own production code, architecture, and quality because scalability and maintainability have not yet been proven for non-engineer-built systems. Ben’s forecast is that broader non-engineer production work requires highly complete product specifications, reusable design systems, and documented edge cases; he estimates this is at least about two years away and prefers not to trade quality for speed prematurely.
- Trust and safety as core product investment: Trust and safety sits under Hinge’s Chief Product and Technology Officer and is described as roughly a third of the organization; the company uses AI, including traditional machine-learning models, to identify fake users and bad actors.
- Distinguish AI PM roles: A modeling PM defines what the model should do, how core capabilities such as factuality, long-context handling, and reasoning are measured, and which error rates are acceptable; in frontier labs, these responsibilities may be split across capability-specific PMs. An AI feature/product PM treats AI as a capability in service of a product, relying on deep domain and workflow expertise and designing the surrounding experience—the last mile that adapts a general model to a specific use case. Merely using an LLM for knowledge-work productivity does not constitute a separate AI-PM role, although rapid prototyping, vibe coding, and AI-assisted data analysis are useful PM skills.
- Use a two-layer measurement system: Keep established north-star measures such as retention, fulfilled user needs, satisfaction, information quality, and trust, but add AI-specific proxy metrics including continuous side-by-side win rates against a prior model, competitor, or internal baseline. Supplement those evaluations with behavioral signals—copying answers, following recommendations, clicking links, giving feedback, or consuming generated media to completion—to inform reinforcement-learning and human-feedback loops. Define product principles and the intended experience before building evaluations, then translate those principles into measurable tests that can teach or improve the model.
- Make trust an explicit product-design requirement: For AI features added to trusted products, set an acceptable error rate for the intended release scale, launch first with power users where appropriate, and study real failure cases such as conflicting source information. Improve user trust through factuality and grounding in trusted sources, clearly placed citations, visual cues that distinguish higher- from lower-confidence information, and redirection or qualification when questions are politically sensitive or genuinely ambiguous.
- Use synthetic users selectively: Synthetic users are especially useful for cold-start products or features, privacy-constrained research, automated regression testing, and checking edge cases after a change. They can accelerate iteration when the PM performs initial human QA and the team then validates with trusted testers before broader real-user exposure.
- Match team shape to problem complexity: A senior developer may take some AI problems a long way alone, but complex or unsolved products also require product principles, AI research, and repeated iteration; AI work is likely to involve more prototyping and greater overlap between engineering and product roles. Engineers can own technical quality checks, while PMs apply product sense to define principles and prioritize the capabilities and use cases that matter to customers. The practical starting point for an AI feature PM is a deep user problem or pain point—not a gimmick—followed by friction reduction and principles such as purpose over possibility and human accountability over full automation.
- Pivot and discovery: Lightfield’s team stopped its presentation product despite reaching about 2 million users per month because the founders did not like the product and could not see a path to making it indispensable for discerning professional users. They rejected simply waiting for better models because the core gap was missing context about the presenter, audience, and relationship—not general reasoning ability. They then identified sales and marketing users in their existing base, ran 12 free pilots, and followed customer requests from presentation decks into research, lead qualification, and account expansion. Connecting CRM, call-recording, and warehouse data exposed incomplete and conflicting records, leading them to reframe the problem as organizing business reality for humans and machines. An early go-to-market assistant had daily users but no pricing power because it did not own the underlying data and faced ten competitors; after restarting around a CRM, ten startups used the barely finished product daily and sent feedback roughly every two hours. The founder’s pivot heuristic was to find real pain, build a product that solves it, focus obsessively on customers, and ignore surrounding noise.
- Product design principles: Lightfield made a chronological activity log—the full history of interactions, documents, product usage, and payments—the canonical primitive from which conventional CRM fields and stages are updated. Because a fully unstructured approach made queries too slow, the team adopted a semi-structured model that stores large amounts of unstructured activity data while using it to infer causality. To avoid locking customers into a flawed initial data model, it used schemaless onboarding: connect email and other systems, assemble relationships automatically, and fill or revise fields later. For enterprise adoption, the team kept dashboards and table views while also supporting natural-language workflows; sales sequences that once required explicit conditions could instead be generated as agent-written recipes.
- Monetization and execution: Lightfield tested seat pricing, where the heaviest user consumed about 10,000 times more than the lightest, then pure consumption pricing, which led signups to avoid using the product. It settled on a platform fee plus seats for core CRM work and consumption pricing for pipeline generation, workflow automation, intelligence, and forecasting. It avoided outcome-based pricing because sales outcomes depend heavily on each customer’s product-market fit, so it charges for the work performed. Internally, a 40-person team removed fixed swim lanes: everyone joined the same daily standup, ranked cross-functional problems, and whoever was available took the next problem; planning was continuous, with a low bar to start projects but a high bar to ship them, reinforced by company-wide bug bashes. For prioritization, the team evaluates an account’s expansion potential over a three-year horizon and leans toward building for the fastest-growing customers rather than the average customer.
- A customer identified a flat, unstructured branch in an AI-generated opportunity solution tree; the team spent three weeks tracing the issue to its source instead of applying a quick fix, creating four new AI evaluation metrics and testing 16 experiment variations.
- Reliable AI product quality requires evals, guardrails, and orchestration—not prompt engineering alone. Judges need calibration against production data, upstream fixes can worsen downstream errors, and a self-auditing agent may be preferable to endless prompt tuning.
- Opportunity solution trees are only useful when their details are accurate enough to show teams exactly what to work on next.
- Evaluate AI agents on recurring work, not one-off outputs. For cross-system workflows, test whether an agent can identify meaningful changes, carry context between systems, reopen the loop later, and return when a human decision is required—not merely generate another draft.
- Use the second run as the core product test. A useful agent should preserve prior state, avoid duplicate work and repeated alerts, recognize when nothing meaningful changed, and make yesterday’s work reduce today’s effort. A practical quality bar is: “Would I give the Bot the same job again next week?”
- Design around jobs and continuity rather than organizational categories or isolated artifacts. Agents are most valuable when they remove coordination work—remembering to check systems, transferring context, and rebuilding briefs—while stopping at explicit human-approval boundaries.
- Evaluate AI agents as durable handoffs, not first-run demos. Model the job as a loop: notice a signal, decide whether it matters, make a change, observe the result, and retain what was learned rather than stopping at a single artifact. Test the second run for retained context and baselines, suppressed duplicate work, and the ability to stay quiet when nothing meaningful changed. Use “Would I give the Bot the same job again next week?” as the quality bar; success means responsibility has moved out of the user’s head.
- Design for cross-system traversal with explicit autonomy boundaries. In fragmented workflows, an agent can move across systems of record while carrying working context instead of relying only on predetermined system-to-system paths; this requires clean data, reliable systems, and stable access. Let the agent proceed through routine steps but stop for human approval before consequential actions such as changing budgets, sending messages, or publishing pages.
- Use repeatability and failure modes as product acceptance criteria. Shah cataloged more than 900 public Grok Bots and is testing which jobs merit continued handoff; failure signals include weak sources, lost state, repeated alerts, constant corrections, and broken logins.
- For evaluating AI-agent productivity, measure successful work per hour, not hours saved alone: task success increased from 59% to 93% while completion time fell 38%, yielding roughly 2.5× more successful outcomes per unit of time.
- A controlled study compared 108 people using agents with 110 completing the same context-heavy knowledge-work tasks without them; agent users completed more work faster and had higher confidence in their work. The report cautions that it does not yet show how reclaimed time or quality improvements translate into business outcomes, so PMs should pair productivity metrics with downstream ROI measures.
- Evaluate AI agents as ongoing jobs, not one-off demos. Test whether the agent preserves state on the second run, avoids duplicate work and false alarms, stays quiet when nothing meaningful changed, and knows when to stop for human approval at consequential steps. The practical quality bar is: “Would I give the Bot the same job again next week?”
- Prioritize loop continuity over artifact generation. A useful agent should reopen the relevant source with prior context, determine whether anything materially changed, carry context across systems, and return when a person needs to decide; the value is reducing the need to remember checks, transfer context, and rebuild briefs—not maximizing autonomy.
- AI-driven technological change is creating career uncertainty for tech workers, including questions about whether product management remains a viable long-term career and whether starting a company is the only path forward.
- Shreyas Doshi frames career decision-making around alignment with one’s authentic identity and life goals, explicitly aiming to help people avoid FOMO, envy, and career anxiety.
- Lenny Rachitsky reports that stage fright affected his talks, large meetings, and time in the spotlight, though it has improved in recent years.
- His practical approaches are low-stakes public-speaking practice through Ultraspeaking, considering beta-blockers only after personal research and consultation with a doctor, and repeating reframing mantras such as “Look what I get to do,” “Don’t perform. Be yourself,” and “Find the joy in it.”
- Tobi Lutke recommends deliberate belief rehearsal: he wrote “I like public speaking” for 10 minutes daily for a week while afraid of speaking, and says he now loves it; he adds that writing something about yourself 100 times can help the brain reconcile with that belief.
- Claude is merging Cowork and chat into a single Claude experience: users can ask quick questions or hand off reports, while Claude continues working after the laptop is closed, asks for clarification when needed, and preserves the user’s final say.
- The unified experience is rolling out to Pro and Max subscribers over the following few weeks.
- Instrumentation and metrics: The workflow uses PostHog for analytics and traces and Homeric for skills management and analytics. Beyond usage, it tracks cost per token, generation time, input/output payload sizes, structured-output errors, and— for coding agents—first-pass rate: how often a change passes reviews/tests and is deployable on the first iteration.
- User-grounded evals: Real user failures are added to the evaluation dataset, with some cases marked “mandatory to detect.” Those cases are run repeatedly, and the eval must return a failure in 100% of them.
- Model-change regression testing: New models trigger a broad eval rerun—described as 100 repetitions across 100 use cases rather than a single run over 10 cases—followed by dashboard review. Eval results can remain difficult to interpret, and better eval performance may not improve perceived UX if it causes a significant speed reduction.
- Prompt and skill validation: More detailed workflow steps have sometimes reduced consistency and eval scores, while reordering modular skill components materially changed generative quality; evals were used to detect these counterintuitive regressions.
Casey Winters describes an unexpected long-tail consequence of growth work at a large consumer company: when the company changes its privacy policy ten years later, “they will personally email you a thousand times.”
- AI monitoring products should distinguish meaningful movement from routine updates and interrupt users only when the change matters; otherwise, users still have to perform the filtering the AI was meant to handle.
Grok Bot testing highlights two product requirements for delegated AI workflows: bots should resume jobs autonomously with context intact, and their outputs should be usable without requiring users to inspect how the work was done. A Friday 10 AM PT demonstration will show marketing tasks the author would hand off to Grok Bot.
- A mid-sized B2B SaaS product-discovery team’s AI interview tool contaminated live customer insights with snarky placeholder labels from its testing dataset, including claims that customers found the feature “useless” or pricing a “cash grab.”
- The tool’s automated highlight reel removed hedging and stitched together dramatic fragments, making nuanced “this is fine but” trade-offs sound like broad customer anger. The incident exposed the need to treat AI-generated research summaries as drafts: verify training/test-data separation, inspect raw context, and sanity-check outputs before executive or customer readouts.
- AI coding agents are shifting the PM bottleneck from implementation to alignment: one PM reports that carefully crafted requirements and technical alignment now lag developers using Claude agents, while parallel iterations and refactors create ongoing work-allocation and coordination chaos.
- AI-assisted specifications, research, prototypes, and slides produced “pretty decent” results but remained insufficient for discovery and alignment; the PM also tried supplying more business context so developers could generate ideas independently.
- AI acceleration can weaken release controls: a commenter reports that code bypassed QA and reached production because AI “cleared it,” with bugs then discovered in production rather than before release.
Why Marketing Is a Good Test for AI Agents
The marketing software landscape has 15,505 products. A surprising amount of the work between them still lives in people’s heads.
In 2011, the marketing technology landscape had about 150 products. We spent fifteen years giving almost every part of marketing its own software.
Analytics measures behavior while CRM remembers relationships and support hears complaints. Ad platforms buy distribution. Social networks show what people are saying. The CMS holds what customers see.
As those systems specialized, people absorbed the coordination work between them.
A customer says something on X. Their product activity lives somewhere else. The CRM knows part of the relationship. Support already knows what went wrong. Analytics shows what they did. The website may still be making a promise that no longer fits.
Someone has to recognize that those facts belong to the same story and decide what happens next.
For a long time, that someone has been the marketer.
The marketer became the API.

That is why marketing is such a good place to find out whether AI agents are actually useful.
I’m testing this in public on Friday at 10 AM PT. I went through more than 900 public Grok Bots looking for marketing work worth handing off. I’m putting the strongest examples through one test. Would I give the Bot the same job again next week?
The work lives between the tools
We call it a martech stack because the word stack makes the pieces sound connected. In practice, much of the harder work happens across the gaps.
The problem is rarely opening the CRM. The work begins when the CRM says one thing, product usage says another, and the person on X sounds like they are about to leave.
Someone has to notice the mismatch and decide which source matters. They still have to carry the context into the next system and remember what already happened.
Thousands of individual software problems got solved without removing the glue work people still had to do.
The artifact got cheap
The first wave of AI for marketing went after the most visible part of the job, the thing you make. Blog posts, ads, emails, images, landing pages, and summaries all got cheaper quickly.

Content marketing also had the largest product outflow in the latest martech landscape as basic generation moved into the models themselves and the software marketers already use.
The first draft is becoming abundant. That makes the surrounding work easier to see.
Someone still has to know what deserves attention and which signals matter. They still have to explain why performance moved and decide what should happen next.
Marketers are already using AI at high rates while still reporting generic campaigns and slow response times because the context they need is spread across systems.
As the bottleneck moves toward context, agents become more interesting.
Marketing is a loop
A lot of marketing software is organized around artifacts like campaigns, emails, posts, landing pages, ads, and reports. The work itself behaves more like a loop.

You notice something and decide whether it matters. You make a change, watch what happens, and remember what you learned. Then the world changes and the loop starts again.
Posts keep producing information after publishing. A campaign changes as people respond. Competitor reports age the moment they are written. A lead can become more interesting overnight. Web pages can become wrong because the product changed somewhere else.
The artifact is one moment in a longer job.
Generative AI made that moment much easier to produce. Agents can take on more of what happens around it.
A useful agent can reopen the source tomorrow with yesterday’s context intact. It can decide whether anything changed enough to matter, carry that context into the next tool, and come back when a person needs to make a decision.
That is a much bigger handoff than asking for another draft.
A surprising amount of knowledge work is remembering to reopen the loop.
The missing layer may be traversal
Software integrations usually assume the path is known ahead of time. System A sends something to System B through an API, script, or workflow designed in advance.
That model works beautifully when the process is stable and predictable.
Marketing often asks for a path that depends on what you find.

You may need to open the CRM and find the account, then check product usage and support history. From there you may need to inspect what the company is saying publicly or what the current page promises before deciding what deserves attention.
Agents give us another way to handle some of that work because they can move through the systems themselves.
Grok Bot makes the idea unusually literal. Each Bot uses a shared computer. It can use connectors where they exist and work through websites where they do not. It can use files, keep context over time, and run again later.
Systems of record store specialized truths. Agents can traverse those truths while carrying working context.
This still depends on clean data, reliable systems, and stable ways to access them.
Some work may no longer require every path to be designed before the work can begin. A fragmented marketing stack can start behaving more like one system because the worker can move through it.
The second run tells you if the handoff held
Agent demos usually focus on the first run, which is also the easiest one. Everything is fresh. There is no prior state to preserve, no duplicate work to suppress, and no history to remember.
The real job starts on the second run.

A competitor watcher should know which changes it already reported, and an ad monitor should know its baseline. Lead and social Bots should know which accounts or replies they already handled.
Yesterday should make today easier.

That gives memory a much better test than asking whether the product has a memory feature.
If the agent remembers everything and I still have to remember what to tell it, I still own the job.
The second run also exposes a quality demos rarely celebrate, the ability to stay quiet.
A competitor watcher should be quiet when nothing meaningful changed. A social listening Bot should avoid resurfacing the same conversation. An ad monitor should ignore harmless noise.
A useful agent has to know when nothing happened.

A bad monitoring agent creates another feed to manage. Every false alarm puts a little more work back on the person who was trying to hand the job off.
Trust also depends on where the agent stops. Watching ad spend can run on its own while changing the budget may need approval. The same applies to finding a prospect and sending the message, or drafting a page and publishing it.
The quality bar is the handoff.
Can the agent keep going when the next step is routine and stop when the next step belongs to a person?
That matters more to me than maximum autonomy.
The jobs are escaping the org chart
I wanted to see what people were actually trying to turn into agent jobs, so I started cataloging public Grok Bots.
The catalog crossed 900 live public bot shares this week.
At that scale, patterns start to show. One of the clearest is where the work ends up.
Grok Bot’s official marketplace has a Marketing section with social listening, search briefs, ad monitoring, video production, and event operations.
Product contains Bots for current market research and competitor monitoring. Sales includes Bots that turn product activity into prospecting opportunities or help make a pitch deck. A marketer could use all of them.
The marketplace shelves still follow the org chart. The same jobs already spill across them.

Across the catalog, the same jobs repeat. Some gather what people are saying or watch for change. Others turn evidence into a brief, make the asset, or check the claim before it goes out.
Departments organize people. Jobs organize work. Agents make that distinction much easier to see.
The next generation of marketing software may organize itself around the jobs people want to stop carrying in their heads.
Which marketing jobs should still live in my head?
Imagine a Bot that watches customer discussions and competitor changes. It remembers which signals it already showed you. When something matters, it checks the CRM and product context, prepares the next action, and stops at the point where a person should decide.
Another Bot watches what happens after the change.
Each step is ordinary on its own. The value comes from the continuity between them.
The job keeps moving while the context stays attached.
For years, one of the big marketing questions was which software the team should buy.
There is another question worth asking now.
Which marketing jobs should still live in my head?
That question is much more demanding than asking whether an agent can complete a task.
A Bot can complete the task and still create more work than it removes. Weak sources, lost state, repeated alerts, constant corrections, or a broken login can all strand the workflow.
The architecture can look right and the handoff can still fail.
That is why I am testing the jobs instead of reviewing the Bot descriptions.
My quality bar is simple.
Would I give the Bot the same job again next week?

A yes means the responsibility moved. Anything else means the job is still mine.
Marketing makes agents prove it
The useful question is whether the responsibility leaves your head.
Marketing gives us a clean way to see that happen because the work changes every day and prior context matters. Much of the result can be checked, and the consequences make the boundaries visible.
A useful agent has to survive the second run, the quiet day when nothing changed, stale context, and a source that disappears. It also has to know when the next step belongs to a person.
If it can keep the job through those conditions, something meaningful has moved.
When that works, you no longer have to remember to check the thing, carry context from one system to another, or rebuild the same brief. The job keeps moving without needing to live in your head.
That is the promise I find interesting.
Generative AI made marketing artifacts much cheaper to produce. Agents may now lower the cost of the coordination around them.
That coordination is where a surprising amount of marketing lives. It’s also where the difference between a great demo and a useful worker becomes obvious.
I went through more than 900 public Grok Bots looking for marketing work worth handing off. Now I’m testing the strongest jobs against that bar.
Friday at 10 AM PT, I’ll show you the ones that earned another week.
- Evaluate AI agents on recurring work, not one-off outputs. For cross-system workflows, test whether an agent can identify meaningful changes, carry context between systems, reopen the loop later, and return when a human decision is required—not merely generate another draft.
- Use the second run as the core product test. A useful agent should preserve prior state, avoid duplicate work and repeated alerts, recognize when nothing meaningful changed, and make yesterday’s work reduce today’s effort. A practical quality bar is: “Would I give the Bot the same job again next week?”
- Design around jobs and continuity rather than organizational categories or isolated artifacts. Agents are most valuable when they remove coordination work—remembering to check systems, transferring context, and rebuilding briefs—while stopping at explicit human-approval boundaries.
- Evaluate AI agents as durable handoffs, not first-run demos. Model the job as a loop: notice a signal, decide whether it matters, make a change, observe the result, and retain what was learned rather than stopping at a single artifact. Test the second run for retained context and baselines, suppressed duplicate work, and the ability to stay quiet when nothing meaningful changed. Use “Would I give the Bot the same job again next week?” as the quality bar; success means responsibility has moved out of the user’s head.
- Design for cross-system traversal with explicit autonomy boundaries. In fragmented workflows, an agent can move across systems of record while carrying working context instead of relying only on predetermined system-to-system paths; this requires clean data, reliable systems, and stable access. Let the agent proceed through routine steps but stop for human approval before consequential actions such as changing budgets, sending messages, or publishing pages.
- Use repeatability and failure modes as product acceptance criteria. Shah cataloged more than 900 public Grok Bots and is testing which jobs merit continued handoff; failure signals include weak sources, lost state, repeated alerts, constant corrections, and broken logins.
- Evaluate AI agents as ongoing jobs, not one-off demos. Test whether the agent preserves state on the second run, avoids duplicate work and false alarms, stays quiet when nothing meaningful changed, and knows when to stop for human approval at consequential steps. The practical quality bar is: “Would I give the Bot the same job again next week?”
- Prioritize loop continuity over artifact generation. A useful agent should reopen the relevant source with prior context, determine whether anything materially changed, carry context across systems, and return when a person needs to decide; the value is reducing the need to remember checks, transfer context, and rebuild briefs—not maximizing autonomy.