ZeroNoise Logo zeronoise
Post
Agent Products Are Entering the Permission-and-Trust Test
3 min read
286 docs
Meta’s Muse launch puts agent access, approval, and failure handling under the microscope; Supermemory’s rapid shutdown adds a counterexample on retention and strategic fit. The brief pairs those signals with operating, stakeholder, and career tactics for PMs.

Coverage is incomplete: some monitored sources or documents could not be processed. This brief covers the available verified material.

Big Ideas

Agent products now need a permission architecture. Meta’s Muse is designed for task handoffs rather than chat: each user gets a dedicated cloud VM, Sentinel controls connected apps and accounts, and Muse asks before spending or sending messages. Internal testing nevertheless found private iCloud photos exposed, unreliable ticket and inventory monitoring, and mid-task logouts; the launch had already been delayed to address security. Meta’s “rule of two” avoids combining untrusted content, sensitive access, and external action, but still requires human approval gates. For any agent feature, specify what it can touch, who approves, when approval occurs, and what happens when it fails.

Standalone tools face a bundled-and-AI benchmark. A ProductManagement thread reports Miro acquired by Bending Spoons for $1.355B versus a last $17.5B valuation. In the same discussion, users cite moves to FigJam and AI-built prototypes, while others defend Miro’s differentiated interaction quality and report 250+ Miro users versus 10+ Figma users. The PM question is whether a product owns a workflow strongly enough to survive bundling—not whether its basic features can be copied.

Tactical Playbook

Productize repeatable AI work instead of automating judgment. Ramp’s creative queue saw more than 375 requests in six months. In a four-week overhaul, artists defined layouts, narrative patterns, and quality standards while builders encoded them; a Slack agent now returns editable decks, moving human work upstream to edge cases and new patterns. Apply the pattern to PM operations: choose a recurring artifact; encode its context, constraints, examples, and evaluation criteria; assign owners and feedback loops; route unfamiliar or defining work to experts. Judge success by whether output is grounded, editable, reviewable, and visible—not merely rendered. Ramp attributes a separate 48-hour launch to shared context and authority close to the work, not lower standards.

Use skip-levels for signal, not status. Bring what is working, friction, unclear requirements/priorities/ownership, risks leadership may not see, one improvement, and a concrete request for a decision or help; skip the PowerPoint. The payoff is longer-horizon context and less-filtered ground truth, not another status channel.

Case Studies & Lessons

Supermemory chose unshipping over launch reach. About one month after launching Company Brain, the team reported 500k+ X impressions, 2M+ across platforms, and hundreds of companies using it; it discontinued Company Brain and Nova and refunded charged users. Most of its first 100 signups churned. The founder cited worsening product clarity, doubt that Supermemory was best positioned to build the broader company-brain future, and conflict with infrastructure customers, then refocused on the memory API. Use retention, strategic fit, and channel conflict as continue/kill criteria; launch reach is not product proof.

Career Corner

Keep the reasoning path in the job. Shreyas Doshi warns that AI can increase answer and prototype volume while teams bypass the thought process and atrophy core product skills. Make the PM own the problem frame, evidence trail, and decision rationale even when AI supplies drafts. Aakash Gupta similarly argues influence is less commoditized than SQL, impact sizing, or technical fluency; his practical models are reciprocity, matching media richness to stakes, careful delivery, warm introductions, and attention to small credibility signals.

Tools & Resources

Agent-native launch pattern: cfo.ai. Users reply with a business, the agent does work in public, and the output becomes the next demo; the team also rebuilt the spreadsheet layer around agents rather than placing an agent atop legacy software. For agent products, design usage to generate inspectable proof and revisit the underlying workflow, not only the interface.

Agent Products Are Entering the Permission-and-Trust Test
Julie Zhuo
  • When to pursue hyperpersonalization: Zhuo argues it is most valuable for high-frequency workflows where small frictions compound, especially across seams between tools, and for products where users want software to express their individual preferences. She places this shift on a spectrum from fixed use-case software to platforms and all-purpose builders, with increasing user freedom and imagination.
  • Product principles: Make the interface malleable; let preferences go beyond predefined toggles through descriptive intent; let user feedback change the product; and treat cross-tool seams as part of the product experience.
  • Dogfooding case: Zhuo found her remote workflow for managing eight AI-agent terminal windows cumbersome and six clicks deep, so she built a mostly working alternative in about 30 minutes and then tailored it with a custom grid, agent-status indicators, and support for both Claude and Codex. Because she was the target user, she could identify and remove friction and maintain a highly efficient feedback loop.
  • Personalized learning case: For her children’s math and reading app, Zhuo used stories tied to each child’s interests, friends, hobbies, and recent trips, added lesson-level difficulty and content feedback, and introduced rewards based on their requests. She cites a study of 145 ninth-graders in which personalized algebra problems were solved faster and more accurately, with the largest gains among struggling students and benefits persisting after personalization was removed.
I Never Want to Use Third-Party Software Again
Product Management
  • Valuation and growth signal: The thread reports Bending Spoons acquiring Miro for $1.355B versus a last valuation of $17.5B; the original poster frames the reset as a warning about heavily funded products pursuing hypergrowth that they cannot sustain.
  • Standalone-tool economics are under pressure: Commenters report moving from Miro to Jira Whiteboard, FigJam, Confluence Whiteboards, or AI-built prototypes because alternatives are cheaper, “good enough,” or bundled with existing software; one commenter nevertheless reports 250+ Miro users versus 10+ Figma users at their company.
  • AI is changing prototyping and collaboration workflows: A designer says their organization is using Figma less because of AI; another says Figma Make credits can be exhausted in a day, Claude is cheaper for similar work, and Make does not use existing design-system components without extra effort. Other comments say Claude-generated whiteboards, Claude Code, and MCPs are putting some visual-collaboration use cases at risk.
  • Differentiation lesson for PMs: One Miro user cites fluid object movement and connector prediction as meaningful UX advantages, while another says basic whiteboarding is easier to copy; collaboration products therefore need to prove that distinctive interaction quality and workflow value justify standalone pricing over bundled or AI alternatives.
Miro is acquired by Bending Spoons at $1.355B compared to last valuation of $17.5B My org is moving to Jira whiteboard. It’s a bit clunky but it’s not worth the added expense to keep miro around. Figjam is good enough and cheaper, and Atlassian put out confluence whiteboard. Hard to justify the Miro licensing cost. (Personally I th… I don’t think you need alternative if Miro works. Their product will continue. Eventbrite and Meetup continue to operate while being part… Depends. Miro can be used across the company from many departments, Figma is clearly aimed at designers. Just as an example, we have 250+… I’m a designer now at a growth SaaS, previously FAANG, who’s used Figma the entirety of its existence and we are using it less-and-less t… Figma Make looked promising until they started charging for it. Most of their user base is tech or tech-adjacent and is likely using Clau… Uff. Wouldn’t dare to in this moment. Claude is producing whiteboard as artefact in chat. My guess is, they will come for whiteboard tool… I think the valuation plunge is more a signal of how the market perceives visual collaboration moving forward. A lot of use cases those t… Compared to competitors, the way objects move when you drag them on the board feels much more fluid and intuitive. Their connector arrows…
Hiten Shah
  • Brand as software as an operating model: Hiten Shah argues that a guideline only records what a team has learned, while software can apply that knowledge to the next deck, landing page, or launch so work starts from accumulated company context. The proposed “institutional memory with hands” connects customer calls, product data, analytics, Slack, and Notion, carrying audience context, product truths, promises, approved examples, and escalation rules into the work.
  • Implementation pattern: Treat the knowledge layer as a product with owners, builders, evaluation criteria, maintenance, feedback loops, and a roadmap; version, test, and continuously improve it. Automate proven, repeatable work, but route unfamiliar or defining decisions to human experts and keep review mandatory where reliability is uncertain; a partial output that forces a rebuild can be worse than no tool.
  • Ramp case study: Ramp’s creative queue received more than 375 requests in its first six months, prompting a four-week overhaul focused on decks, whitepapers, and one-pagers. Artists defined layouts, narrative patterns, and quality standards while builders encoded them into a system. A Slack agent can turn an attached Markdown file, Google Doc, or Notion page into editable HTML, PDF, and PowerPoint outputs; the approach expanded to whitepapers, one-pagers, and landing pages, moving creative effort upstream toward new patterns and edge cases.
  • Cross-functional launch lesson: A Ramp team spanning brand, product, engineering, communications, and partnerships moved from an opportunity to a named, designed, developed site with customer proof, assets, and a launch within 48 hours; the speed was attributed to shared context, authority, and craft being close together rather than to lowering standards.
Paul gave a name to something I think we’re going to see everywhere. Brand as software. A brand guideline records what a team has figured… Brand as software
Mind the Product
  • Meta’s agent product bet: Meta launched Muse on September 8 as a personal AI agent available on the web, iOS, Android, and WhatsApp; unlike a chatbot, it is designed for users to hand off tasks such as drafting emails, booking flights, and making purchases. Each user receives a dedicated cloud virtual machine, while Sentinel controls what Muse may do; users choose connected apps and accounts, can interrupt the agent, and must approve spending or sending messages. Meta’s rollout sequence was model-first—Muse Spark in April, open-weight Muse Glimmer in August, then the consumer product.
  • Agent safety framework and launch-readiness lesson: Meta’s “agents rule of two” says an agent should not simultaneously process untrusted content, access sensitive data or systems, and take an action that changes something or communicates externally; Meta describes this as only one layer that must be supplemented with human approval gates. Despite delaying Muse from April specifically to address security concerns, internal testing reportedly found safeguard bypasses that exposed private iCloud photos, unreliable ticket and inventory monitoring, and Meta CTO Andrew Bosworth being logged out mid-task. The PM takeaway is to evaluate not only whether an agent ships, but what it can access, who approves actions, when approval occurs, and how it behaves when it fails.
  • Commercial and positioning choices: Muse offers 100 million tokens per week free for typical users, with $20-per-month and $100-per-month paid tiers; Meta says it plans to monetize primarily by taking a small cut of completed transactions rather than showing ads. Meta is positioning Muse for mass consumer adoption as a personal assistant, while Grokbot is positioned toward power users and enterprise work, illustrating different target segments and willingness-to-pay strategies. The broader market is also testing agent placement across browsers, standalone apps, and messaging, giving PMs competing distribution patterns to evaluate rather than a settled interaction model.
Muse: Meta's new personal AI agent. How much control would you hand over?
The community for ventures designed to scale rapidly | Read our rules before posting ❤️
  • The smart-fridge concept bundles three difficult products—reliable inventory sensing, grocery purchasing, and health guidance. The recommended lowest-risk starting wedge is expiry detection plus meal suggestions using user-confirmed inventory; the team should prove repeat engagement with waste reduction before taking on custom hardware. Autonomous ordering makes a single detection error financially costly, while therapeutic diet guidance requires a higher trust bar.
  • The proposed trust and autonomy rollout uses three stages: visibility and expiration reminders, a reviewable draft basket, and fully autonomous ordering constrained by budget and diet rules. The founder sets a 99% item-recognition target before scaling and treats an 80%-accurate order as a likely churn event because of misidentification risk.
  • The business model deliberately treats hardware as a low-margin or subsidized distribution channel, with recurring premium nutrition subscriptions and transaction revenue from retail and delivery partners as the primary monetization.
  • Early discovery feedback flags adoption and market-entry risks: one commenter rejects an in-home camera and additional AI in the home, another flags cybersecurity in Samsung’s comparable product, and a hardware commenter reports that distribution can overwhelm feature advantages.
The concept bundles three hard products: reliable inventory sensing, grocery purchasing, and health guidance. I’d pick the lowest-risk we… I will not promote. "Autopilot for your kitchen" - Smart fridge ecosystem that manages inventory, calorie tracking, and auto-orders groceries. You couldn't pay me to use this. I don't need or want a camera in my fridge, and I don't even slightly trust any company at this point th… Samsung has this and it’s one of the worst products for a variety of reasons. Recommend checking what failed. One big issue is cyber secu… I ran into this with hardware, distribution swallowed every feature advantage we thought we had.
Hiten Shah
  • Supermemory discontinued its Company Brain and Nova products roughly one month after launch, despite 500k+ impressions on X, more than 2M impressions across platforms, and hundreds of companies using Company Brain; customers who had been charged were refunded. Early product feedback was also weak: most of the first 100 signups churned.
  • The shutdown prioritized product clarity and strategic focus: adding applications had made it harder to explain what Supermemory was, the team judged itself poorly positioned to build the broader “company brain” future, and application-layer products risked making infrastructure customers view Supermemory as a competitor. Supermemory therefore refocused on its frontier memory API and infrastructure for agents.
  • PM lesson: strong launch reach and initial usage do not outweigh poor retention, unclear positioning, or channel conflict; reassess whether a product strengthens the core strategy and customer trust, and be willing to unship it when focus and differentiation suffer.
An update on supermemory
The community for ventures designed to scale rapidly | Read our rules before posting ❤️
  • Paid beta / client-funded R&D: The founder’s B2B SaaS has a working core but an incomplete roadmap; they proposed charging early customers $99/month against a $199 list price in exchange for regular feedback, arguing that payment is a stronger willingness-to-pay signal than free-beta feedback. A commenter described this approach as “client funded R&D,” saying upfront payment can validate the sale and motivate customers to try the product seriously.
  • Conditional monetization rule: One response recommends charging from day one when users receive a service they will eventually pay for, but waiting until paid features are usable when the current users are not the revenue source. Another suggests keeping users free until the product works properly, then starting with a low price and increasing it over time.
  • Founding-cohort mechanics: If offering a free or discounted pilot, one commenter recommends a contract with six months free and an extension if the product is not ready; a follow-up recommends explicit terms and a clawback if customers break the agreement after the six-month period.
  • Free-tier design: A commenter says free users may help put a company logo on the product but often show less product value, while another gives a concrete example of a daily usage cap on free accounts versus hourly background checks on the paid plan.
Should you charge your first customers before the product is finished? I will not promote The most successful startups I've witnessed managed to charge customers before the project was even finished. I'm not talking custom prod… It's a good question. In my case, we're currently not charging anything for our thing, because the main users are not going to be the sou… In for free until the product works properly, then you can start charging them. You can start at a low amount and grow it from there. I h… Put them on contract and give 6 months free. Extend if not ready. You can claim paying subscribers. Also great vibe check at 6 month mark. This is the answer. You don’t just want to hand out anything free, they’ll never want to pay later and it sets kind of a difficult negoti… Either way is defensible, keep in mind that free users generally don't value your product at all, but can be useful to get the ball rolli… Free users not valuing it depends on the free tier. ours has a daily cap and the paid plan just runs the same background check every hour…
Hiten Shah
  • cfo.ai demonstrates an agent-native launch pattern: users reply publicly with a business or modeling task, the agent performs the work in public, and the output becomes the demo and proof of capability. Each subsequent user can provide another task, creating a compounding loop in which product usage generates evidence before others try the product.
  • Product design implication: instead of placing an agent on top of legacy software, rebuild the core workflow—in this case, the spreadsheet layer—around how agents actually work.
introducing ═══════════ 𝚌𝚏𝚘.𝚊𝚒 ═════════════ tl;dr 1/ runway is now [http://cfo.ai](http://cfo.ai) 2/ you can hire [@arithecfo](https://x… I’m late to this party, but [http://cfo.ai](http://cfo.ai) just showed what an agent-native launch can look like. You reply with a busine…
Lenny Rachitsky

Lenny Rachitsky reported strong demand for in-person PM community after an event attended by more than 1,000 PMs; recordings of the talks are planned for his YouTube channel.

Still processing from yesterday's event (one of the most meaningful days of my life)— more to come—but a few quick things for now: 1. Tha…
Shreyas Doshi

AI can help product teams produce more answers and prototypes, but bypassing the underlying thought process may atrophy essential product skills and invite the teams’ own obsolescence, according to Shreyas Doshi.

This is an important point in product too. Product teams are using AI to get to higher volumes of answers and prototypes but many of them…
Product Management
  • A PM with 12 years of experience across industries argues that product managers should not default to constant availability: use a “by appointment only” operating norm, reserve direct interruptions for the manager or manager’s manager, and let engineering counterparts handle incident response; the PM also argues that customer communication does not always need to come from product.
  • Practical boundary-setting tactics include defining a genuine-emergency escalation path, turning off push notifications, and using a separate work phone that is physically put away after hours; one PM working across US, European, and Asian teams uses a timed lockbox and gives family the work number for emergencies.
  • Availability should be calibrated to role and context rather than treated as universal: one proposed model is five-minute responsiveness during 9–5, up to two hours in the evening, no overnight expectation, and four-to-six-hour weekend response for simple follow-ups or production issues, with more flexibility needed around major launches.
I’ve been a PM for 12 years now. All different types of roles across multiple industries. I spent the first half of my career always avai… Sounds counter-intuitive, but giving my personal cell number to my team is the best way I’ve found to support boundaries like this. They … You need 2 phones. Personal phone out of reach during working hours, but still within hearing distance. Work phone in a lockbox outside o… You seem to be reasoning yourself in circles. If you can't solve this yourself you might not be a great product manager. The simple answe…
Product Management
  • Treat skip-levels as strategic context and relationship-building, not status reporting. They may serve to review or sense-check a manager’s performance and keep senior leaders connected to actual work, while giving PMs access to broader strategy, longer-horizon opportunities and risks, and context that may be filtered before reaching leadership. Use the time to explain decision rationale and trade-offs, discuss blockers beyond your manager’s ability to unblock, and ask what leadership considers most important next.
  • Bring a concise, forward-looking agenda. Cover what is going well, unnecessary friction, unclear requirements, priorities, ownership or product direction, risks leadership may not see, a proposed improvement, and any concrete request for a decision, escalation, context, or blocker removal.
  • Build trust before raising sensitive feedback, then use the relationship for growth. Early meetings can focus on understanding the director’s intent, listening, and building rapport; when raising problems, stay candid and professional, connect them to work or morale impact, and bring evidence plus improvement ideas rather than personal attacks. Skip-levels can also support mentoring, career direction, skill development, and building internal champions or sponsors.
Ex product leader here. There are typically three intents behind this meeting: 1. Your Director using it for a performance review of your… As someone who’ve lead team of teams, this is what I would do in those skip level meetings from the individual’s perspective. 1) It’s a g… Skip levels are less about status updates and more about giving them visibility they don't normally get, so lean toward strategy and cont… Go to the meeting with an agenda. Start with the small talk of how their weekend was or what's going on next. If they give you something,… You get 45 minutes of my time every two months. Use it wisely \- I don’t want status updates \- I don’t want to know what you are working… I think it’s valid that a director may want to understand how your manager is performing. But in reality, you don’t always know their int… Why are you so fearful? A person in the position of Director or VP has a responsibility to lead their org well and keep their people happ… I did this recently with our AVP of marketing, and it was a pretty chill 30 minutes, mostly because there was no ask for any official pre… Depends on their personality but I know folks used this for mentorship as well. Asking them about their experience, what you can work on,… Skip meetings are so important. Gotta use those to network and build champions and sponsors for yourself.
Product Management
  • Interview-prep tactic: A PM is building a tool that maps companies such as Google Maps and Spotify to provide industry context for product-sense questions, with Airbnb and Stripe planned; the approach helps candidates compensate for limited exposure beyond their own niche.
  • Career-tool product pattern: Joey centralizes a candidate’s work history, matches it to job requirements, asks about missing information, and drafts role-specific resumes that can be checked against source evidence while learning the user’s writing preferences without inventing experience. Before sending resume text to an outside writing model, it strips direct identifiers; the product is invite-only while its builder gathers feedback.
Working on a little tool for myself to use for product interview prep (I'm already a PM). It helps me orient myself for product sense que… I’ve been building Joey (I couldn't think of a name and this sounded friendly) because I got tired of rewriting the same work history for…
Product Management
  • In complex, heavily customized B2B products, UAT should validate customer-specific business processes, configurations, legacy data, and user flows in addition to the technical and functional coverage expected from QA/SIT. Unexpected behavior may require joint investigation because the customer holds context about business-process nuances, legacy configurations, and test data that the product team may not possess.
  • To reduce surprises without treating every discovery as a preparation failure, introduce an earlier customer-facing QA or pre-UAT step that brings the customer into the process sooner for feedback before formal UAT.
  • When each customer has a heavily customized legacy version and the product lacks documentation, test automation, code reviews, and CI/CD, UAT problems can reflect accumulated technical debt and unclear ownership rather than an individual PM’s preparation. Teams should distinguish preventable defects from customer-context discovery and address the underlying documentation, automation, and SME gaps.
Is it normal for UAT to cover things that the product team needs to investigate with the client? Your understanding of UAT is correct, not the org's. UAT exists precisely because there are things — business process nuances, legacy con… I’m guessing this is has to be a high value contract that you are working directly with customer to launch features. While the highly cus… Sadly this is not the customer’s decision, but rather our company. Also no, sadly our product is a legacy heavily customized behemot diff…
The community for ventures designed to scale rapidly | Read our rules before posting ❤️

For product positioning and demo work, prioritize clarity over polish: consider paying for a clean deck or short product demo when it prevents confusion, iterate the narrative through real investor or customer questions, and defer a full rebrand or elaborate video until customers are pulling for it.

My rule would be: spend on clarity, not polish. A clean deck and a short product demo can be worth paying for if they stop people getting…
Hiten Shah

PM operating principle: standards scale only when teammates can apply and enforce them without the leader being present; otherwise execution remains dependent on that individual.

Your standards are only scalable when someone can enforce them without you in the room.
The community for ventures designed to scale rapidly | Read our rules before posting ❤️
  • For an early B2B SaaS with a live core but an incomplete roadmap, charge once the product delivers clear value: payment provides a stronger willingness-to-pay signal than free usage, though charging can add friction for early adopters.
  • Treat paid early access as a product-discovery loop: commenters argue that paying customers provide more candid feedback and force stricter prioritization; start with a small, reasonable price and iterate based on what customers say.
  • One proposed implementation was $199/month list pricing with $99/month for the first customers in exchange for regular feedback, positioned between lower-cost tools and much more expensive enterprise offerings.
Should you charge your first customers before the product is finished? I will not promote You're focussing on the wrong thing. A software product is rarely, if ever, finished. The question you need to answer is "am I delivering… Charge them, yeah. Not to milk them, but because free customers don't tell you the truth about what's broken. They tolerate stuff, they g…
Hiten Shah

Hiten Shah’s local-AI experimentation pattern is to start with one ordinary machine assigned one job, repeat the experiment until the failure modes are understood, and then decide what additional setup is needed; he plans to demonstrate six local-model jobs on an ordinary laptop.

Building a home AI lab has made me appreciate how much one computer can already do. A lot of my experiments now start the same way. One m…
The community for ventures designed to scale rapidly | Read our rules before posting ❤️
  • Use a 30-day focused B2B go-to-market learning loop when an early service business needs repeatable demand: choose one buyer type and one narrow service or pain, use direct outreach to learn objections, and test a small number of referral partners.
  • Reduce adoption friction with a small, fixed-scope assessment instead of pitching the full managed service. Use prior client work as proof only when the buyer, problem, and scope are genuinely comparable, and frame the case study around the problem and outcome.
  • Evaluate each experiment by qualified conversations, segment and contact reason, agreed next steps, and time spent qualifying—not by raw contact or reply volume. Add referral partnerships once the buying trigger and introduction process are clear.
For the next 30 days, choose one buyer type and one service, then make direct outreach the learning channel while testing a small number … The thing to replicate may be why a client needed you then, not where you found them. For the next 30 days, I'd test one narrowly scoped … For B2B security work, I’d probably pick one narrow buyer and make the first offer painfully easy to say yes to. A small fixed-scope asse…
Product Management

A product manager who moved from 10 years as a data engineer into media technology product management reports difficulty identifying new AI use cases despite weekly internal AI demos; they are considering targeted courses or hands-on PM-and-AI projects to build capability.

How to be use AI as a workforce?