We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
Two separate recommendations make Ben Thompson the clearest signal in this set: Bill Gurley calls a Thompson conversation “very good,” while Not Boring calls his AI-router breakdown “characteristically good.”
Standout: My conversation with Ben Thompson
- Content type / creator: Video conversation hosted by Patrick O’Shaughnessy with Ben Thompson, who has written Stratechery for more than a decade and is described by O’Shaughnessy as one of his favorite business thinkers. Watch the conversation.
- Recommended by: Bill Gurley, who calls it “very good” and says Thompson brings a career of thinking to current events. Gurley’s favorite part is the opening argument that there is no reasonable path for the United States to achieve AI dominance over the rest of the world.
- Key takeaway: The conversation connects that geopolitical claim to questions about AI funding, model capabilities, compute and chip suppliers, advertising, Nvidia’s competitive position, and the major platform companies.
- Why it matters: It is the strongest single resource here for pressure-testing AI “race” narratives against the underlying economics of funding, infrastructure, distribution, and business models.
Companion AI-infrastructure read: Stripe Acquiring OpenRouter, Aggregating AI?, Flipping the Business Model
- Content type / creator: Stratechery analysis by Ben Thompson. Packy’s link points to Stratechery generally; the current Stratechery listing identifies the relevant OpenRouter analysis by this title and supplies the direct link.
- Recommended by: Packy McCormick, who calls Thompson’s breakdown “characteristically good” in a discussion of AI model routers.
- Key takeaway: Routers send each task to the model that fits its needs: frontier capability when it is worth paying for, or a cheaper and faster model when it is sufficient. Packy ties that choice to “Return on Tokens”; Stratechery’s accessible teaser frames Stripe’s reported OpenRouter acquisition as a bet on a future market of models and aggregation.
- Why it matters: The useful lens is operational rather than tribal: model choice becomes a routing problem governed by capability, speed, and cost. The direct article page exposed only a teaser in the source capture, so this recommendation is best treated as a pointer to that thesis rather than a full article summary.
Research pick: Voluntary attention regulates acute immune responses in humans
- Content type / creators: Research paper by Nofar Mizrachi, Menachem Rottem, and Liron Rozenkrantz in Nature Human Behaviour. Read the paper.
- Recommended by: Packy McCormick in Weekly Dose of Optimism, who connects the result to his argument that deliberate attention creates meaning and concludes, “Attention really is all you need.”
- Key takeaway: Across three pre-registered within-subject experiments, directing attention toward bodily sensations rather than distraction produced roughly 1.5-fold smaller acute inflammatory responses, with the effect appearing in about 90% of participants across two cohorts.
- Why it matters: This is a concrete, checkable finding—not a general claim that attention treats disease. The authors stress that the model involved localized, acute, histamine-induced inflammation in healthy participants, and that its relevance to infectious, chronic, or autoimmune conditions remains unclear.
Personal practice: Radical Acceptance: Embracing Your Life with the Heart of a Buddha
- Content type / creator: Book by Tara Brach, Ph.D. The post text does not name the book; the attached cover identifies the title and author. Ferriss’s recommendation post.
- Recommended by: Tim Ferriss, who calls it a “godsend” for people who beat themselves up and says it was recommended to him by a neuroscience PhD and then a cynical friend, both of whom said it changed their lives.
- Key takeaway: Ferriss describes it as one of the most useful books he has read, says it is easy to digest, and recommends one short chapter before bed each night.
- Why it matters: The recommendation comes with a clear, low-friction way to test the book rather than a generic endorsement.

Cover image from Tim Ferriss’s post.
Direct answer: the bundle verifies only the subject, not the full article: the first line says "Stripe is reportedly acquiring OpenRouter, an implicit bet on a future market of models and the chance at Aggregation" , and the exact title is available only in bundle metadata ("Stripe Acquiring OpenRouter, Aggregating AI?, Flipping the Business Model"). The rest of the document is a paywall/subscription block, so the author and the concrete argument about model routing, aggregation, and the flipped business model are not extractable here.
Findings:
- L1 is the only substantive article line and contains the teaser claim about Stripe acquiring OpenRouter as a bet on a future market of models and Aggregation.
- The document goes paywalled at L5 ("Subscribe to Stratechery Plus for full access.") and L5-L99 are subscription/FAQ/marketing content, with no further article text.
- The author is not identified in any supplied line; the title is only in metadata. Flag these as verification gaps before recommending this resource.
Direct answer: Yes — the supplied Stratechery page identifies an article title and gives a one-line takeaway relevant to the AI-model-router topic.
- The page is a weekly overview of the Stratechery bundle , and it lists the article “Stripe Acquiring OpenRouter, Aggregating AI?, Flipping the Business Model”.
- Its listed description frames Stripe's reported acquisition of OpenRouter as “an implicit bet on a future market of models and the chance at Aggregation” .
Caveats/gaps:
- The supplied bundle contains no mention of Packy McCormick, so it cannot confirm that this is the Stratechery resource he referenced .
- Only the roundup listing and one-sentence publisher description are present; the full linked Stratechery article is not included, so the takeaway is limited to that summary .
Direct answer. The bundle contains the full text of the Nature Human Behaviour article "Voluntary attention regulates acute immune responses in humans" (journal name confirmed in-text; title itself present only in the bundle metadata, not as a numbered line). It reports that voluntarily directing attention toward bodily sensations during histamine-induced acute skin inflammation produced markedly smaller, more rapidly resolving inflammatory responses than attentional distraction, across three pre-registered within-subjects experiments . Packy McCormick is not mentioned anywhere in this bundle, so his specific recommendation cannot be verified from this source; the connection is by subject matter only.
Findings
- Journal. Nature Human Behaviour is named in the peer-review information and the publisher's note . The article's exact title does not appear as a numbered line in the extracted text, so it cannot be line-cited.
- Authors (gap). The extracted text begins at the Abstract heading with no author byline , and the bundle metadata lists author as nil; authorship cannot be verified from this extraction.
- Study design. Three pre-registered within-subjects experiments using the standardized histamine skin prick test to induce acute cutaneous inflammation; conditions were completed within each individual in counterbalanced order, with participants serving as their own control; Exp. 1 (n=37) contrasted internal attention with video-based distraction, Exp. 2 (n=20) equated visual input and task structure to isolate attention itself, Exp. 3 (n=17) attenuated sensory signalling with lidocaine . Participants attended two laboratory sessions at the same hour, 3–5 days apart ; a trained experimenter was blinded to experimental conditions and participants were blinded to the study's aim and conditions ; 57 participants were included in the final analyses .
- Core finding. Internal attention produced a substantially smaller and more regulated inflammatory response than distraction, with reduced wheal and flare at the 20-min peak (wheal, 3.5 ± 1.1 versus 5.0 ± 0.7 mm; flare, 10.6 ± 8.3 versus 14.0 ± 8.0 mm); the effect was consistent in ~90% of participants, with wheal and flare increasing by ~1.5-fold on average under distraction . Exp. 2 replicated this with ~1.6-fold reduced magnitudes under internal attention and 90% of participants again showing larger responses under distraction .
- Temporal dynamics. Condition differences emerged within 3 min and increased over time, peaking at 20 min; ~90% of internal-attention participants showed wheal stabilization or decline after 20 min versus 46% under distraction .
- Mechanisms. Two complementary pathways: a sensory-dependent route — lidocaine attenuation of sensory signalling weakened but did not abolish the regulatory effect of attention — and a top-down route engaging parasympathetic vagal activity, with heart rate variability (HRV) significantly higher during internal attention than distraction (t(48)=2.4, P=0.023, d=0.34) ; HRV increases persisted when sensory signalling was attenuated and were indistinguishable from intact-signal internal attention .
- Relevance to the recommendation context. The authors suggest that turning attention away from unpleasant sensations "may come at a physiological cost, impairing the body's ability to regulate inflammation," and connect the effect to attention-based interventions such as mindfulness that involve sustained engagement with bodily sensations ; the conclusion states that "voluntary attentional engagement alone is sufficient to regulate acute immune responses" .
- Caveats. The model is a localized, acute, histamine-induced response in healthy participants; generalization to infectious, chronic inflammatory, or autoimmune conditions is unclear; reduced inflammatory magnitude may not universally reflect beneficial regulation; and HRV is an indirect index of vagal activity .
The book pictured is titled Radical Acceptance: Embracing Your Life with the Heart of a Buddha by Tara Brach, Ph.D. It is noted as an "Updated Edition." The cover also features a quote attributed to Thich Nhat Hanh: "An invitation to embrace ourselves with all our pain, fear, and anxieties." The provided image does not contain any text, blurbs, or references identifying it as a recommendation by Tim Ferriss.
Recommendations found (all cited):
1. Tim Ferriss — unnamed book (title/creator/link missing) Tim Ferriss (@tferriss) explicitly recommends a book: “For those of us who beat ourselves up, this book is a godsend”; it was recommended to him by a neuroscience PhD and then a cynical friend, both saying it changed their lives; he calls it “one of the most useful books I’ve read,” easy to digest, and suggests one short chapter before bed each night . The post text names no title, creator, or book link; the only linked asset is an image URL (https://pbs.twimg.com/media/HQQdHoNW0AA5Xhg.jpg), so no visual claim about a cover can be made from the supplied text .
2. Packy McCormick / Not Boring — “some of my favorites…” and Stratechery In Weekly Dose of Optimism #207, Packy McCormick runs an explicit “some of my favorites…” roundup :
- Matic Robots — maticrobots.com; “a home cleaning robot you can talk to” .
- Quince — quince.com; “high quality … everything at affordable prices. half my wardrobe” .
- Notion — notion.com; “not boring’s everything hub” .
- Ramp — ramp.com; “Time is money. Save both.” .
- Create — trycreate.co (30% off code notboring30); “you should be taking creatine (and supporting Dan)” .
Caveat: Ramp is later named as one of “our not boring capital portfolio companies” , and Create’s blurb includes “supporting Dan” , so those two favorites carry a possible affiliation/self-interest signal; full creator identities are otherwise not named in the post.
Separately, he calls Ben Thompson’s Stratechery breakdown “characteristically good” and links Stratechery.com in the routers item ; the specific article title is not given.
3. Bill Gurley — endorses linked Ben Thompson piece Bill Gurley (@bgurley) endorses the linked Patrick O’Shag X post (https://x.com/patrick_oshag/status/2089713931183153293) as “This is very good,” adding “Ben brings a career of thinking to his views on where we are today” . His stated favorite part: “There is no reasonable path to AI global dominance for the U.S. vs the world. It’s a fantastical notion that falls apart when thinking through the details” . The supplied text does not name a title or format for the underlying Ben Thompson piece beyond the link .
4. Chamath Palihapitiya — thin endorsement, no reason Chamath Palihapitiya (@chamath) shares an @buildamericanai X post (https://x.com/buildamericanai/status/2090136716258476525) with only “Wow” . No title, creator, subject, or takeaway appears in the supplied text, and it is not labeled there as a “data-center article”; treat as a share/endorsement signal rather than a substantive recommendation.
Excluded The Next Big Idea Club “Book of the Day” newsletter contains book recommendations, but per supplied metadata it is a book-club newsletter rather than an item from a tech founder/VC/startup leader .
Patrick O'Shaughnessy recommended Ben Thompson's Stratechery, describing Thompson as 'one of my favorite business thinkers' after more than a decade of writing. He shared a video conversation with Thompson, which he said covered every important company in the industry and the forces acting on them, including AI, Google, Amazon, Nvidia, TSMC, Intel, Samsung, Microsoft, Apple, and Meta .
Tim Ferriss recommends a book — title not stated in the post text, with an image attached — for people who beat themselves up, calling it "a godsend" and "one of the most useful books I’ve read." He says it was recommended by a neuroscience PhD who said it changed her life, then by a cynical friend who said the same; he finds it easy to digest and suggests reading one short chapter before bed each night . Shared at https://x.com/tferriss/status/2090828220895514689.
Bill Gurley recommended Patrick O'Shag's video conversation with Ben Thompson (Stratechery) https://x.com/patrick_oshag/status/2089713931183153293, calling it "very good" and saying "Ben brings a career of thinking" to today's market questions; Gurley's favorite part is the opening argument that "there is no reasonable path to AI global dominance for the U.S. vs the world" . The episode features Thompson, who "has been writing Stratechery for over a decade" and is one of O'Shag's favorite business thinkers, covering why the US winning the AI race could be problematic, whether AI funding will run out, Google becoming Berkshire Hathaway, why ads are amazing, TSMC/Intel/Samsung, Nvidia's invisible price cuts and biggest competitors, and Microsoft/Amazon/Apple/Meta .
Elon Musk replied "Yes" to an X thread by @r0ck3t23 and linked to it, endorsing it for his followers . The thread summarizes Musk's memo to Tesla employees on management and communication: excessive meetings are "the blight of big companies," it is fine to walk out or drop off a call when you aren't adding value; communication should travel the shortest path, not the chain of command; managers who enforce chain-of-command communication may be fired; avoid acronyms and nonsense words; and common sense should override ridiculous company rules . The thread is at https://x.com/r0ck3t23/status/2090901321054371872.
Packy McCormick recommends the research paper Voluntary attention regulates acute immune responses in humans (Nofar Mizrachi, Menachem Rottem & Liron Rozenkrantz, Nature Human Behavior), highlighting the finding that deliberately directing attention regulates immune function and concluding "Attention really is all you need" . He also points to Ben Thompson's "characteristically good" breakdown of AI model routing on Stratechery.
Elon Musk (@elonmusk) endorsed @JamesLucasIT's X post on Rome's birth-rate collapse, writing: "The reason Rome fell was because they stopped making Romans" . The linked post (https://x.com/jameslucasit/status/2090847857666294046) describes Augustus's failed laws (Lex Julia and Lex Papia Poppaea, ~18 BC and AD 9) to encourage childbearing , draws parallels to today's falling fertility rates (e.g., South Korea 0.72, Italy/Japan ~1.2, China ~1.0, US ~1.6, UK ~1.5) , and ends with Arnold J. Toynbee: "Civilizations die from suicide, not by murder" .
VC Chamath Palihapitiya recommended a CNN article on what a good data center deal looks like, quoting a post about Quincy, Washington: about 30 facilities shoulder an estimated 57% of local property taxes, funding a $120M high school, library, hospital, and police and fire stations. He endorsed it with "Wow" and shared the link. Article: https://www.cnn.com/2026/08/11/business/data-centers-ai-economy
Tim Ferriss recommends a book as "a godsend" for people who "beat ourselves up," calling it "one of the most useful books I've read" and saying it is "not nearly as woo-woo as it might seem" . It was recommended to him by a neuroscience PhD who said it changed her life, then by a cynical friend who said the same; he finds it easy to digest and suggests reading one short chapter before bed each night . The post's text does not name the book title; an image accompanies the post .
Stratechery by Ben Thompson
( Adam Mares, Greatest of All Talk)
Welcome back to This Week in Stratechery!
As a reminder, each week, every Friday, we’re sending out this overview of content in the Stratechery bundle; highlighted links are free for everyone. Additionally, you have complete control over what we send to you. If you don’t want to receive This Week in Stratechery emails (there is no podcast), please uncheck the box in your delivery settings (opens in new tab).
On that note, here were a few of our favorites this week. - Apple Makes Compromises in the EU. Ben has covered the angst surrounding the App Store since the beginning of Stratechery and was focused on Apple’s policies long before it was cool. Now that the company’s finally been forced to compromise in various forums — including a settlement this week with the EU, as well adjustments to its ATT policies in Germany — I thought the most remarkable aspect of Ben’s coverage on Wednesday (opens in new tab) was how incidental and boring it all seems in the shadow of the possibilities and concerns that exist everywhere else in tech right now. We had a fun conversation about that dynamic at the top of this week’s episode of Sharp Tech (opens in new tab) before turning to AI cybersecurity, vibe coding epiphanies, and more insight on writing with and without AI. — Andrew Sharp
- Truth (Social) and Reconciliation. Sharp China returned from its annual August hiatus this week, and in an episode that’s outside the paywall (opens in new tab), we talked about various sources of U.S.-China friction before Xi’s visit to D.C. in September. Before that, however, we began in Korea with more questions than answers as Foreign Minister Wang Yi descended on Seoul in the wake of President Trump’s abrupt Sunday evening decision to reduce joint military exercises between the US and ROK. As for that Trump decision, in this week’s Sharp Text article (opens in new tab), I used the Korea news as an opportunity to marvel at the exhausting economy of takes and theories that accompanies every foreign policy decision (and meme) under the current administration. — AS
- August Fun with the Clippers and Lakers. During the quietest period of the NBA calendar, there’s actually been quite a bit of news out of L.A. On one hand, we have a terrific mess as Buss family members squabble and Mark Walter’s DOJ-flavored cashflow problems have led to a shocking sale nine months after he initially purchased the team. On the other, Steve Ballmer and the crosstown Clippers might be in the (relative) clear after a 12-month NBA investigation into alleged salary cap circumvention. We discussed all of it on this week’s Greatest of All Talk (opens in new tab), including frustrations with Clippers media coverage, what the NBA wants for the Lakers, and a memorable Top 5 segment about our top vacations. — AS
### Stratechery Articles and Updates- Stripe Acquiring OpenRouter, Aggregating AI?, Flipping the Business Model (opens in new tab) — Stripe is reportedly acquiring OpenRouter, an implicit bet on a future market of models and the chance at Aggregation.
- Nvidia Backs OpenAI Data Center, Anthropic News, Google Buys Spirit Airlines Data (opens in new tab) — Nvidia makes another deal, this time with a frontier lab; Anthropic’s revenue continues to amaze; and maybe data finally is oil.
- Apple Settles With E.U., U.S. App Store Fees, ATT Rules in Germany (opens in new tab) — Apple’s App Store is finally facing the reality of lower fees, and the EU should be satisfied with its work; it’s ok it’s late.
Sharp Text by Andrew Sharp
- So What Was Trump Saying to South Korea on Sunday? (opens in new tab) — A snapshot of Truth Social foreign policy and the take economy it inspires.
Dithering with Ben Thompson and Daring Fireball’s John Gruber
- More on Watermarking (opens in new tab)
- Apple Settles With EU (opens in new tab)
Asianometry with Jon Yu
- How TSMC Uses Old Fabs to Make New Chips (opens in new tab)
- China Built 700 Waste-to-Energy Plants in 6 Years (opens in new tab)
Sharp China with Andrew Sharp and Sinocism’s Bill Bishop
- Wang Yi Visits South Korea; Remembering Zhu Rongji; US-China Ahead of Xi’s Visit; How China Monitors Foreigners (opens in new tab)
Greatest of All Talk
- August Fun with the Lakers and Clippers, Top 5 Takeable Teams or Players, Top 5 Vacations (opens in new tab)
Sharp Tech with Andrew Sharp and Ben Thompson
- The App Store in the Shadow of AI, Offensive and Defensive Cybersecurity, Q&A on Financial Planning, AI Writing, American Sports (opens in new tab)
This week’s Sharp Tech video is on the turnover at DeepMind.
- Listen to Podcast (opens in new tab) Listen to this post: Log in to listen (opens in new tab)
On January 1, 1870, Jay Cooke, hailed as an American hero for his role in financing the Union effort in the Civil War, signed a contract that would, if you squint, lead to world war.
In 1864, Congress had created the Northern Pacific Railway Company with the goal of linking the Great Lakes and Puget Sound with tracks that would eventually run from Duluth to Tacoma; the charter included 40 million acres of land adjacent to the proposed line in exchange for accomplishing the build-out. For the ensuing six years, however, Northern Pacific struggled to secure financing, even as the Union Pacific and Central Pacific railroads built towards each other, driving the golden spike linking Sacramento and Omaha in May 1869.
Northern Pacific had approached Cooke about funding in 1866, but lacked the generous federal guarantees that undergirded Union Pacific and Central Pacific (which, it should be noted, led to an incredible amount of graft); Cooke, himself no stranger to the financial power of the federal government, wasn’t interested. Ultimately, however, Northern Pacific gave him an offer he couldn’t resist: a commission of 12 percent on every bond, and $200 of Northern Pacific stock for every $1,000 in bonds he sold.
Cooke soon found that his institutional peers agreed with his earlier refusal, and weren’t interested in his bonds, so he leaned on the same tactics he honed selling war bonds: appeals to patriotism, control of the media, and promises of railroad fortunes, backed by industrial-scale distribution. At the peak Cooke employed 1,500 salespeople and funded 1,300 newspapers (through a combination of advertising and direct payments) with a brand burnished by the Civil War. Retail investors could already buy railway bonds; Cooke made them his primary funding mechanism.
This was, to be certain, an incredible innovation. It used to be the case that if you couldn’t get loans from the government or from banks, you couldn’t get much money at all. The problem was that Northern Pacific’s capital needs were endless, and by September 1873, as credit tightened worldwide thanks to a crash on the Vienna stock exchange and the demonetization of silver, Cooke, who had been funding Northern Pacific from deposits in between bond issuances, could find no more buyers. The subsequent bankruptcy of Jay Cooke & Company triggered the Panic of 1873, culminating in endless railroad bankruptcies across the country, a multi-year depression, multi-decade deflation, and, one could argue, the financial conditions that made Europe, four decades later, into a tinder box.
Northern Pacific did eventually finish their line, by the way, with multiple bankruptcies along the way; ultimately, they were one of four railroads that were merged to form the Burlington Northern Railroad. Burlington Northern would eventually merge with the Atchison, Topeka and Santa Fe Railway to form BNSF Railway; Berkshire Hathaway would purchase the parent corporation in 2009.
### Blowing Through Debt
If this story sounds vaguely familiar it might be because Cooke is — for obvious reasons — a central character in Liaquat Ahamed’s new book, 1873 (opens in new tab), released earlier this year. Ahamed is not shy about drawing a link between the collapse of the railroad buildout and the current AI moment; the book’s very first page — even before page 1 — is about translating sums of money, and concludes thusly:
In order to grasp the true significance of sums of money that relate to the economic situation of whole countries — such as the size of the indemnity imposed on France after the Franco-Prussian war — it is most useful not simply to make allowances for changes in the cost of living but instead to adjust for changes in the size of economies. To translate such figures into comparable 2026 magnitudes, multiply by a factor of 1,200. Thus the $500 million that went into U.S. railway bonds annually during the boom years of the early 1870s would today be the equivalent of $600 billion, roughly what is projected to be invested by major tech companies in 2026. Microsoft CEO Satya Nadella is certainly aware of the connection: he cited 1873 as “the book to be read” on the company’s recent earnings call (opens in new tab). Perhaps it’s not a coincidence, then, that Microsoft, alone amongst the hyperscalers (opens in new tab), still boasts substantial free cash flow — $19.6 billion last quarter. Microsoft is the one hyperscaler still abiding by the dictum used to deny the existence of a bubble: its CapEx isn’t funded by debt. This was, believe it or not, a defense that could be used for nearly all of Big Tech a year ago; then, between September and November, Oracle, Meta, Alphabet, and Amazon issued a combined $80 billion in debt for building out infrastructure. That was only the beginning: after raising a combined $108 billion in all of 2025, these four companies have, as of July 7, already raised $194 billion this year. Unsurprisingly, spreads are rising, and 86% of the bonds issued this year are already trading at higher yields than at issuance. Cover for recent issuance has fallen to less than 2x, from 5x in February. The real shock, however, came at the beginning of June, when Google announced it would raise $85 billion in equity, including a special $10 billion issuance to the aforementioned Berkshire Hathaway. I wrote at the time in The Google Capital Company (opens in new tab): It is worth noting that $10 billion is a relatively small amount of money to both companies. To that end, perhaps the primary utility is as a signaling mechanism. On Google’s side, the signal is that the expected demand is actually far greater than anyone thinks, and that the company is ready and willing to fund supply using all means at its disposal, including equity; for them Berkshire Hathaway’s investment is an endorsement of this view and a validation of the wisdom of the investment. And, on the flip side, if the signal is correct, then Berkshire Hathaway is getting a deal and putting its cash flow machines to work building the future. I concluded: Implicit in this analysis was that there was enough compute capacity in the world to be bought; what happens, however, when and if there isn’t? What if the ultimate battle — the one that determines who gets compute — becomes a matter of who can bring the most cash to bear? And what if that advantage compounds, such that the company with the most cash capacity ends up with the most compute capacity (which we already know they will sell, in addition to using themselves) driving the ability to generate more cash? In that world, what company would be your best bet? The implied answer, of course, was Google.
DeepMind Drama
Google right now is no one’s bet, at least in terms of the frontier. After the departure of DeepMind CEO Demis Hassabis (technically promoted to chairman, but no longer in charge of day-to-day operations) and Gemini co-lead and former Chief Scientist Jeff Dean, along with a host of other prominent researchers, SemiAnalysis declared that Gemini is Cooked (opens in new tab): For all intents and purposes, we believe DeepMind is no longer a frontier lab. We said as much a few months ago to our Tokenomics clients due to large numbers of departures from their reinforcement learning teams and poor compute allocation. Google will continue meandering on and releasing models, but their odds of reaching SOTA again have dropped to zero.
Furthermore, the biggest beneficiary of today’s news is neither Anthropic nor OpenAI—it’s Google Cloud. Whereas Gemini and GCP used to desperately fight for compute allocation, it’s now clear that Thomas Kurian won. We expect GCP revenue growth to meaningfully accelerate as a result. From later in the post: We’ve obviously been quite bearish on DeepMind thus far, and if we had to steelman the case for why they’ll still be able to train a true SOTA model in the future, it would go something like the following:
- The current setup clearly wasn’t working. With the existing leadership team, their odds of catching up to Anthropic/OpenAI looked extremely slim.
- Now that they’ve cleaned house, the new guys can start from a blank slate. Maybe they’ll even acqui-hire a neolab like SSI or Thinking Machines.
- With this new team, their odds of catching up to the frontier actually increase.
Perhaps there’s some world in which this happens, but we think the odds are basically zero. The issue with Google was not Jeff Dean nor Noam Shazeer, but rather their extremely bureaucratic, painfully slow, and strategically timid culture. Remember that DeepMind had an AI chatbot 1 year before ChatGPT but was not allowed to release it due to fears of disrupting their core business. Actually, you could make the case the problem was also Hassabis and DeepMind. I explained in an Update after Google I/O (opens in new tab) how Hassabis’ vision of the frontier was fundamentally different from the other frontier labs because he believed in world models, not just text/code, and concluded: What falls out of [Hassabis’ vision] are models with multimodality — in contrast to Claude, which outputs text only — and, it must be said, not nearly as impressive coding capabilities. This gets at the point of this entire digression: I think it’s possible that the reason Google is widely considered to be behind both Anthropic and OpenAI in terms of coding, particularly long-running agentic workflows that depend just as much on the harness as the model itself, simply comes down to their research team having other priorities. That’s why the coding parts of this keynote fell on the Antigravity team, not DeepMind, and why Hassabis was barely on stage. From this perspective, last week’s events are less surprising, and were arguably foretold at I/O: Hassabis might be right about world models being the path to AGI, but Google has run out of patience in terms of letting him find out; Google co-founder Sergey Brin is reportedly deeply involved and closely allied with Koray Kavukcuoglu, the new DeepMind CEO, and I wouldn’t be surprised if the company is pivoting to Anthropic’s more text- (and thus code-) centered approach.
Google’s Infrastructure Bet
What is fascinating about Google’s position is that these machinations do not necessarily mean the Berkshire Hathaway bet was a bad one; indeed, it’s arguably good news. This is what the SemiAnalysis article was driving towards, and it’s a point I made last week about Google’s recent earnings (opens in new tab): The story seems to be very similar to last quarter (opens in new tab), with even more Google Cloud growth: 82% year-over-year (compared to 63% last quarter, and 32% a year ago), with 36% margins (compared to 33% last quarter, and 21% a year ago). I wondered then how much of this growth was actually Anthropic, and while we didn’t get clear confirmation this quarter, I thought this answer from CEO Sundar Pichai on the earnings call (opens in new tab) about why Google needs to rent 3rd-party capacity was notable:
I think on the bridge deal, the main thing I would say is, look, there are — on the margin, there are very, very large customers of ours on Cloud who we are trying to support them through this extraordinary moment. And the incremental opportunities they are bringing to us, while a short‑term cost over a few months may be very high, in the lifetime of the deal, as we bring more capacity on, is highly ROI‑positive. So those are factors we are taking into account. So are you willing to take upfront a six‑month deal to be able to serve the customer in what is a multiyear opportunity where the margins and the returns are very, very attractive over that multiyear horizon? So hopefully that gives some color on how we’ve thought about those opportunities.
That customer is almost certainly Anthropic. Again from SemiAnalysis: More than 20% of total TPU shipments from 3Q26 to 4Q27 are being sold directly to Anthropic. This is excluding the hundreds of thousands of TPUs GCP already rents to Anthropic today, and the many hundreds of thousands more they’ve committed to rent to Anthropic and Meta over the next 6 quarters…
If you’ve ever listened to an interview of Google Cloud CEO Thomas Kurian, you know he is not AGI pilled. In one podcast (opens in new tab), for example, he argued that it’s great for TPUs to become “general purpose infrastructure” that supports customers like Citadel, the Department of Energy, and generic high performance computing. And when asked why he was selling compute to Anthropic despite them competing with Gemini, he said this was the natural consequence of Google being a “platform company.” Kurian said the same thing to me in a Stratechery Interview (opens in new tab): We sell different parts of our stack. One of the things people don’t realize is we monetize many different parts of the stack in different ways. Like Anthropic, there’s a lot of labs that use our stack — in fact, most of the large AI labs use our stack. So if somebody uses TPUs to either to train their model or to use it for inference, we’re monetizing that part of the stack, that gives us resources to then fund our R&D and other investments. Some of the labs use our TPU and our Gemini model, others may use our TPU and then buy our cybersecurity protection for their models. So as a platform player, we have to allow our technology to be monetized in as many ways as possible and we don’t see it as a zero sum. We’ll see how zero sum compute actually is — there are reports Google’s researchers have been starved for compute (opens in new tab) — but the overall takeaway is that whether or not Google is competing for the frontier, they are absolutely competing to dominate AI infrastructure. And, in a world where intelligence is a commodity, TPUs in particular are a big deal. Last month, in Who’s Afraid of Chinese Models? (opens in new tab), I talked about commodity markets in the context of frontier labs versus everyone else; in commodity markets marginal costs are determinative of not just profitability but also viability, and I made the case that the frontier labs are well-positioned to have superior cost structures for any given unit of intelligence. That cost structure, at least for now, includes the cost of renting compute, and it seems likely that TPUs are cheaper than Nvidia GPUs; Anthropic may have built for TPUs (and Amazon’s Trainium chips) because only Google and Amazon had the wherewithal to fund them, but at this point that ability may very well be a significant advantage. The fact that Anthropic is straight up buying TPUs for its own data centers (converting compute costs from marginal costs to capital costs) suggests that is the case. What is notable is how amenable Google is to share, even at the price of needing to issue equity. This, however, fits the Berkshire Hathaway model that I wrote about in The Google Capital Company: One of the businesses Berkshire Hathaway used the See’s profits for was on the opposite end of the spectrum in terms of capital utilization: BNSF Railway. Railways require a lot of capital to operate; BNSF consumed $3.8 billion last year; they also make a lot of money: BNSF’s net income was $5.5 billion on revenue of $23.4 billion. To put that in perspective, the total amount that Berkshire Hathaway has made from See’s Candies is probably less than $3 billion (the last disclosure was “over $2 billion” in 2019), i.e. less than BNSF made last year…
In fact, you can make the case that Abel is actually just replaying Buffett’s strategy, only this time Berkshire Hathaway is See’s Candies, and Google is BNSF. At the end of last quarter Berkshire Hathaway had $373 billion in cash, and $25 billion in free cash flow in 2025. How many companies could actually employ that cash in a way that generated a high rate of return?
It’s hard to imagine a better option than Google. The company is not only investing in AI, but has optionality in terms of outcomes: its Services business benefits from the investment, it is in contention at the model layer with Gemini, and it can sell capacity to the frontier labs. Moreover, that capacity has a sustainable cost advantage because of TPUs, which means that in a world where compute becomes a commodity — as hard as that is to imagine right now — Google is the hyperscaler that is poised to make the most profit. Notice that I didn’t say margin; if that were Google’s concern they would almost certainly be making different choices. Profit, however, is an absolute number, and Google is bringing everything to bear — first its cash flow, then its debt, and now its equity — on making money from the infrastructure build-out.
Nvidia’s Investable Asset Class
Today corporate executives and financial engineers don’t need to control newspapers; thanks to his new X account (opens in new tab), Nvidia CEO Jensen Huang can go straight to the public. From an X Article (opens in new tab) posted last night: NVIDIA AI Factory Compute Is Becoming an Investable Asset Class
Today, we announced partnerships with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR to establish independent financing platforms designed to mobilize over $500 billion of third-party capital to support the buildout of AI infrastructure over time.
This is a major milestone for NVIDIA and the AI industry. We have moved from an era in which companies bought chips and built data centers project by project to one in which AI factories can be financed as productive infrastructure — with repeatable platforms, long-term institutional capital and a diverse customer base that uses compute to create revenue.
AI has reached an inflection point. It is moving from research into production. AI is creating real value, and the infrastructure behind it is becoming one of the world’s most productive assets. In AI, compute is revenue. Huang argues that Nvidia-based AI factories are fungible, protecting residual value, and that CUDA makes AI factories better over time, extending their economic value; according to Huang: These are the characteristics of an investable infrastructure asset: it produces revenue, serves a broad market, improves in performance over time and can be redeployed. Thus the attempted formalization of a new investment structure: The demand for AI infrastructure is extraordinary. But access to capital is uneven. Many great AI companies, enterprises and AI clouds have demand for compute but do not yet have access to financing at the scale or cost required to build quickly. That is why we are partnering with the world’s leading long-term capital providers.
Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR are also among the world’s leading infrastructure investors, with deep expertise in underwriting long-lived, productive assets. Together, we are creating repeatable financing platforms to help the AI ecosystem build the factories it needs. What Apollo et al. are, are new sources of capital beyond the investment grade debt markets. In that sense this proposed structure is somewhat akin to Google’s equity issuance: a way to secure funding beyond bonds. The difference, however, is stark: whereas equity dilutes the upside for investors without adding risk to the company, this structure preserves Nvidia’s margins by finding new pools of capital willing to bear risk. It’s not a total free ride for Nvidia: the company is backstopping opportunities with up to 25% residual-value based financing, suggesting that Huang believes his “investable asset class” pitch much more than the market does. That is, in a certain sense, a price cut, as the goal is to reduce the cost of capital for entities building data centers with Nvidia chips, by putting Nvidia’s profits on the line for uncertain investments. That guarantee is downstream from Google’s (and soon Amazon’s (opens in new tab)) aggressiveness: why build a data center with Nvidia chips if you can buy TPUs or Trainiums (Nvidia chips are likely better, but if the constraint on new data centers is capital, lower up-front prices may matter more than token efficiency). Nvidia’s bigger problem is one that has been apparent for a long time; I wrote back in 2024 (opens in new tab): In the before-times, i.e. before the release of ChatGPT, Nvidia was building quite the (free) software moat around its GPUs; the challenge is that it wasn’t entirely clear who was going to use all of that software. Today, meanwhile, the use cases for those GPUs is very clear, and those use cases are happening at a much higher level than CUDA frameworks (i.e. on top of models); that, combined with the massive incentives towards finding cheaper alternatives to Nvidia, means both the pressure to and the possibility of escaping CUDA is higher than it has ever been (even if it is still distant for lower level work, particularly when it comes to training). The situation today, with Anthropic and OpenAI appearing to pull away, is even more problematic: Anthropic has not been dependent on CUDA for years, and OpenAI is moving in that direction, at least for inference. If those companies win then Nvidia’s profits will be squeezed — indeed, the implication of that backstop is they already are (this, needless to say, is why Huang’s first post (opens in new tab) was an open letter in defense of open models).
Risky Business
This might not cost Nvidia anything in the end: if AI revenues truly take off, then the debt markets will open back up, and ultimately companies will go back to funding infrastructure investment through free cash flows. Right now, however, is the danger zone, as hyperscalers blow through the debt markets and Google at least starts to tap equity. To the extent Nvidia competes through novel funding mechanisms that, at the end of the day, draw on things like insurance floats and pension funds and other long-run liabilities that are the bread and butter of the asset managers the company is partnering with, the risk — unmarked, unlike equity — is considerably higher. That’s why I started with 1870 and Cooke’s ill-fated agreement with Northern Pacific. Yes, the upside the deal afforded Cooke was incredible, but it was incredible for a reason: it was very risky, and pioneering new funding mechanisms only served to spread the pain when it all blew up. It’s one thing to spend all of your free cash flow; it’s another thing to tap the debt markets. And, beyond that, it’s a completely new nerve-racking thing to bring safety-seeking assets to bear. AI better deliver before it’s too late.
- Listen to Podcast (opens in new tab) Watch on YouTube (opens in new tab) Listen to this post: Log in to listen (opens in new tab)
There’s a story I tell about my first day in STRT-431 at Kellogg School of Management, the introductory class that every first-year MBA was required to take; I leafed through the readings and case studies and was dismayed that there weren’t any tech companies on the docket. Me being me, I spoke to the professor after class wondering why, and was told that the goal of the course was not to necessarily learn about specific industries, but rather to uncover broadly applicable universal principles that could be applied to any company in any industry.
I did not, as I usually tell the story, find this very satisfactory: to me the nature of tech, particularly the fact that software and distribution had zero marginal costs (and zero transaction costs), was something fundamentally different; putting in zeroes in formulas tends to wreak havoc! I soon realized, however, that that was my opportunity. The fundamental insight undergirding Aggregation Theory (opens in new tab) is that zero marginal costs leads to fundamentally different value chains than people once expected from the Internet: centralization and scale in a world where controlling demand mattered more than distributing supply.
What is fascinating about AI, however, is the extent to which those old universal principles are coming back to the forefront. That was never more apparent than this past weekend, when arguments raged on X about the implications of Kimi K3, another open weights model out of China, approaching the state-of-the-art in terms of capabilities. The long and short of it is this: marginal costs are back in a big way, both in terms of short-term implications of state-of-the-art free models, and in terms of the long-term structure of the industry.
### COGS Versus R&D
One of the most common misconceptions undergirding discussion of open weights models is that they are cheaper — free, even. After all, you can just download the weights, and skip the time and expense and capabilities necessary to create your own model. That is, of course, true, but the “free” in this case is a reference to the amount you need to spend on research and development; R&D is a fixed expense that is independent of the revenue you generate. If you spend $1 million in R&D, it doesn’t matter if you do $100 thousand in revenue or $100 million; you still spent $1 million on R&D (it does, of course, impact your profitability).
What is related to revenue is COGS — cost of goods sold — and COGS is real for AI in a way it hasn’t been for software for a very long time. Specifically, running inference on a model — whether that model be Kimi or Fable — costs money, and the amount of money an AI provider spends on inference is, at least in most business models, directly correlated to revenue. To reuse the above example, generating $100 million versus $100 thousand in revenue will likely require 1,000x COGS. In concrete terms, if it costs 50 cents to generate the tokens that drive $1 in revenue, then $100 million in revenue will have $50 million in COGS; $100 thousand in revenue will only have $50 thousand in COGS.
The point in terms of open weight models is that they are not free to serve. Kimi K3 costs (opens in new tab) $3 per million input tokens, and $15 per million output tokens; that is cheaper than Sol’s $5 per million input tokens and $30 per million output tokens, but that might not even be the right measurement.
### Tokens Versus Intelligence
Nvidia CEO Jensen Huang has described what Nvidia is building as “token factories”, and from Nvidia’s perspective that framing makes sense. Nvidia GPUs are model agnostic: they generate tokens, and do so in the fastest and most efficient way possible. That leads to measurements like tokens-per-second, time-to-first-token, tokens-per-watt, token cost, etc., and Huang argues that these metrics will be the basis for decision-making.
This is a framing that definitely made sense during the first paradigm of AI, the ChatGPT era, when tokens were delivered straight to the end user. The second paradigm of AI, however, the reasoning era, confounds this measurement. Reasoning entails an explosion in chain-of-thought tokens, and different models need different amounts of reasoning tokens to arrive at the right answer. Kimi, for example, reportedly uses significantly more tokens than Sol, rendering its price advantage moot. Agents introduce a similar dynamic: some models are more efficient than others in terms of the number of tokens they need to execute agentic workflows.
What this means is that tokens are not a commodity. The defining characteristic of a commodity is that it is fungible: a gallon of oil is a gallon of oil; a ton of copper is a ton of copper; a bushel of wheat is a bushel of wheat. A token from one model, however, is not the same as a token from another model. What is fungible is what is constructed from tokens, which is to say intelligence. In other words, if both Kimi and Sol generated the right answer, then that answer is fungible; the difference in tokens generated to get to that right answer is a contributor to a difference in COGS.
The COGS for intelligence is a function of a few different factors:
- Model footprint: The weights and runtime state determine how much expensive memory and how many accelerators are required to host each serving replica.
- Inference efficiency: Architectural choices (e.g. Mixture-of-Experts) reduce computation per generated token.
- Memory efficiency: Architectural choices can reduce KV cache requirements, allowing more concurrent requests and better GPU utilization.
- Serving efficiency: Batching, scheduling, prefix caching, and other inference optimizations maximize utilization and share work across requests.
-
Token efficiency: The fewer tokens required to reach a correct answer, the lower the inference cost.
The reason this matters is that we are rapidly approaching a state in which intelligence for many economically beneficial tasks is in fact a commodity. Anyone building a basic CRUD app (opens in new tab), for example, can likely do so using models from multiple providers. And, in a commodity market, the route to profitability is not through charging higher prices — again, you can (or will soon be able to) make the exact same app using multiple models — but rather through having a superior cost structure.
Understanding Commodity Markets
It’s worth stepping through the mechanics here, because, as I noted a few months ago in Amazon’s Durability (opens in new tab), the dynamics of commodity markets are not something people in tech are generally familiar with: - In commodity markets, everyone charges the same price, because everyone is selling the same thing; that price is determined by supply and demand.
- The demand for a commodity is a function of price elasticity: the cheaper the commodity, the more demand there is for it, and vice-versa.
- The supply for a commodity is a function of the marginal cost of producing the commodity. The key thing to understand is that the marginal cost of producing the commodity differs by supplier. What this means in practice is that the supplier with the worst cost structure ends up selling the commodity at their marginal cost (if they can produce at all); the profits of everyone else depend on the extent to which their cost structure is better than the marginal supplier. As an example:
- Supplier A can produce 10 units of the commodity for $10 each
- Supplier B can produce 10 units of the commodity for $15 each
- Supplier C can produce 10 units of the commodity for $20 each Let’s assume the price elasticity is such that there is demand for 25 units of the commodity at $20. That means:
- Supplier A will sell 10 units of the commodity for $20, earning $10/unit
- Supplier B will sell 10 units of the commodity for $20, earning $5/unit
-
Supplier C will sell 5 units of the commodity for $20, earning $0/unit
This isn’t precisely right: the reason why Supplier C will bear the shortfall is because Suppliers A and B will be able to slightly undercut them in price, which will of course affect demand (which is elastic), but it makes the point. Supplier A has a great business, Supplier B has a good business, and Supplier C is going to go bankrupt.
Bankruptcy risk is where fixed costs come back to the forefront: Supplier C has both fixed costs (like potentially R&D spend) and also may have taken on debt to finance the equipment necessary to produce the commodity. It can’t price its commodity with these costs in mind — remember, the market-clearing price approximates the marginal cost of the highest-cost unit needed to satisfy demand — but those costs can absolutely drive the supplier out of business. And, if that supplier goes out of business, then prices go up, until another supplier decides to enter (or the other suppliers expand).
The Intelligence Market
Let’s bring this back to models. Right now, none of the above analysis applies because demand exceeds supply for frontier models, and supply is limited by a lack of compute. This compute shortage doesn’t just mean that a compute supplier like Nvidia makes very large margins, but also that Nvidia’s customers, like SpaceXAI, can turn around and resell compute at high margins as well to a company like Anthropic. Anthropic, meanwhile, can pay the markup because they can sell tokens with a higher markup still. It’s not just excess demand that gives Anthropic great margins, however: Anthropic and OpenAI likely have among the lowest costs per unit of frontier-quality intelligence, thanks to model capability, serving scale, and token efficiency. They are serving models at a particular capability level for months before their competitors, and are simultaneously applying the best models to optimizing those costs. It’s also worth noting that the market is not yet treating intelligence like a commodity: demand is for Anthropic and OpenAI specifically, and much less for models that aren’t as good (thus SpaceXAI and Meta selling capacity to Anthropic); one way to think about the push for optimizing cost is that that is a function of defining jobs-to-be-done by intelligence level, such that intelligence buyers can create a market where intelligence is commoditized. In the long run, however, whoever is on the frontier is the best placed to dominate non-frontier markets as well, which are just the frontier minus n-months, i.e. months in which the frontier model makers have been optimizing their cost of serving. All of this is to say that I think the reaction to Kimi and Chinese models generally is pretty over-blown, at least from an economic perspective. Right now there is a price umbrella that is downstream of the lack of compute; I highly doubt that Chinese models are cheaper to serve on a marginal cost basis, they just seem cheaper because Anthropic and OpenAI are so supply constrained that they are charging far more than they would if there were sufficient supply to meet the demand for intelligence.Frontier Lab Paranoia
Why, then, do the model makers in particular seem so panicked about Chinese models? First, I think the frontier labs are anchored in a world where training costs dominated their financial modeling. As long as training consumed more GPUs than inference, it was critical to maximize inference revenue to help fund the next training run, which meant charging very high prices for inference. Going forward, however, I expect the inference market to grow much faster than training costs (and that includes the assumption that training costs will continue to skyrocket), which means they really can make it up in volume. It wasn’t clear this would be the case as recently as eight months ago, but the agent paradigm unlock is so massive that frontier labs should have more confidence that they can not just survive but thrive with lower prices (once they have sufficient compute). Second, intelligence isn’t in fact a perfect commodity, in part because applied intelligence makes itself smarter. Specifically, whoever is running inference is also collecting data, and that data goes into making the next iteration of the model better. This is, on one hand, all the more reason for the frontier labs to lower prices and increase usage as more compute comes online; on the other hand, this is why companies like Microsoft (opens in new tab) are increasingly obsessed with helping companies run their own models. That is much more viable if Chinese models are a viable alternative. Third, the other way that frontier labs can not only differentiate from Chinese models but also from each other is by continuing to integrate up into the customer experience. It’s striking the extent to which Claude Code and Codex are proving to be quite sticky; whichever harness you start working with is likely to be the one you stick with, and that figures to be even more the case with non-technical users. And, in the long run, this imperative to move up the stack does mean that frontier models are absolutely a threat to software providers, including Microsoft. On the flipside, the extent to which software companies who currently own the customer experience have access to competitive models is the extent to which they may be able to resist the encroachment of the frontier labs. Finally, the ideological angle of Anthropic in particular (opens in new tab) is impossible to ignore. This is a company that believes only it can be entrusted with AI, and the existence of open weights alternatives strikes a fatal blow to that presumption.China’s Motivation
Kimi isn’t the only new Chinese model; from Bloomberg (opens in new tab):Alibaba Group Holding Ltd. shares rose as much as 5.4% on Monday after the company launched a preview version of its flagship Qwen3.8 Max model, describing it as second only to Anthropic PBC’s Fable 5. The Sunday release came only days after startup Moonshot AI unveiled a powerful new offering that’s roiled markets and triggered concern in the US about China closing the gap on global leaders like Anthropic and OpenAI. Qwen3.8 Max has 2.4 trillion parameters, joining Moonshot’s Kimi K3 in the heavyweight class. With 2.8 trillion parameters, K3 rivals top offerings and Alibaba is setting similarly high expectations.
Developers can now access Qwen3.8 Max through Alibaba’s coding platforms, including Qoder. Alibaba plans to make the model open-weight soon, expanding access beyond the preview release. Interest in these made-in-China artificial intelligence systems and models is so high that Moonshot was forced to pause taking on new subscriptions late on Sunday to manage overwhelming demand. The fact that Qwen3.8 Max will also have open weights is notable. Alibaba stopped releasing weights for its leading edge models earlier this year, but appears to have reverted that change; I suspect that shift was related to last week’s Xi Jinping speech about AI (opens in new tab) that doubled down on the open weights approach: We should adhere to the principle of openness and win-win and boost innovation-driven development. As a new engine of world economic growth and an accelerator for the shift of growth drivers, AI is moving from the digital world into the physical world. We should seize this rare, historic opportunity to encourage open source, openness, collaboration and sharing. We should facilitate technological innovation, industrial development and scenario-based application of AI. We should make coordinated advances in the transformation and upgrade of traditional industries, the cultivation and growth of emerging industries and forward-looking planning for future industries, so that all sectors and businesses can benefit from AI. The strategy for China is obvious: commoditize your complements. Note that Xi explicitly ties openness to AI “moving from the digital world into the physical world”; the physical world is the world dominated by China, and the country’s lead in areas like robotics is going to massively benefit from widely available AI models. Along the same lines, China does not want the U.S. to gain an asymmetric advantage in AI; to the extent that China can weaken the U.S. frontier labs while strengthening any and all potential U.S. adversaries so much the better, and it can benefit from the innovation that will attach itself to an open ecosystem.
The Distillation Question
By the same token, don’t expect China to do anything about distillation attacks on the frontier labs. I think it is mistaken to attribute all of the success of Chinese labs to distillation, but it’s just as much of a mistake to pretend like distillation doesn’t give Chinese labs a big advantage. That advantage has really come to bear in the last year as post-training reinforcement learning has become increasingly crucial to model performance. Instead of having to fashion reinforcement learning environments from scratch, Chinese labs can simply use frontier labs models as teachers, allowing for rapid improvement at much lower costs (this is not the only reason why Chinese models are cheaper to develop, but it’s a big one). What is interesting is that one of the most important use cases for Chinese models in the West is itself distillation. Thinking Machines, for example, which just released an open-weight model, relies on Chinese models (opens in new tab) to solve the cold start problem for reinforcement learning. Dean Meyer and Konstantine Buhler wrote an excellent article on X (opens in new tab) explaining that distillation means that Western open weight models are fundamentally disadvantaged relative to China:Distillation does not explain China’s entire open-model lead. Chinese labs have world-class researchers, substantial compute, strong pre-trained models, software-hardware codesign, and rapidly improving post-training capabilities. But distillation compresses the costly final gap between a strong base and a near-frontier system. Even if distillation represents a smaller share of a Chinese model’s total capability, it represents a meaningful share of its advantage over American open models.
New enforcement mechanisms will make large-scale distillation harder, slower, and more expensive for Chinese companies. However, enforcement will not eliminate distillation backed by state actors. Every Western frontier advance therefore creates another teacher for Chinese labs. Western builders must either reproduce those capabilities independently or wait to learn from Chinese models. This gap gives Chinese labs a recurring structural advantage over Western companies. This is a point that bears repeating: because U.S. open weight model makers must follow the frontier labs’ terms of service, they (1) are worse than Chinese alternatives and (2) end up distilling the distillation, just with a detour through Chinese labs. Wouldn’t it be better if western open weight model makers could go to the source? To that end, here’s an even more interesting question around distillation: why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here? In fact, this paradox is the solution. I believe that open weight models are good for innovation (and, per the above, I think that labs on the frontier will be fine), but it’s a problem to be dependent on China. The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service that forbid distillation, for U.S. companies at a minimum. Stopping distillation — which is literally just querying the API — is nearly impossible; the U.S. should go the other way and lean into a new copyright policy that both indemnifies the labs and also guarantees that what they learned fuels further innovation for everyone else.
The Reason to Be Afraid
This entire Article has been an exercise in defusing overreaction to Kimi K3 specifically and Chinese open weight models generally; however, there is one reason to be concerned, and that is cybersecurity. Consider this story from The Stack (opens in new tab):Hugging Face said its production infrastructure was breached by an “autonomous” AI agent system early last week. The platform’s security team were initially stymied in their incident response (IR) by unnamed US LLM frontier model guardrails “which cannot distinguish an incident responder from an attacker,” they said. So Hugging Face’s defenders turned instead to the open-source GLM 5.2 model from China’s Z.ai lab – running it on their own infrastructure to analyse the 17,000+ logs, or footprints, that the attackers left behind.
That’s a striking public admission for the New York-headquartered Hugging Face, which lets users collaborate on models, datasets and applications, and which this summer hit the $100 million ARR mark. In an incident report, the company recommended that defenders “have a capable model you can run on your own infrastructure [our italics] vetted and ready before an incident, both to avoid guardrail lockout and to keep attacker data and credentials from leaving your environment.” It’s difficult to overstate how wrong-headed the Trump administration’s panicked response to Anthropic’s release of Fable was, particularly since it exacerbated Anthropic’s worst tendencies in terms of assuming only they can be trusted with powerful AI. In a world with only one AI, it might make sense to reserve the most powerful cybersecurity capabilities for the U.S. government and trusted allies; however, that’s not the world we live in. There are and will be models eminently capable of mounting cybersecurity attacks on existing infrastructure, and those models will be — already are — widely available. The best defense — the only viable defense, in fact — will be to make sure defenders have access to the best models as well. Right now defenders are effectively banned from using Fable or Sol for cybersecurity because of Trump administration directives; that means the best alternative is using models from a country which has been trying to weaken our cyber defenses for years. This is insane! The better course is clear: first, loosen Fable and Sol restrictions on cybersecurity, and second, ensure that U.S. open weight model makers are on an equal playing field with China. Yes, the frontier labs will kick and scream about this, but the Administration should realize that listening to their histrionics has led the U.S. to a position where U.S. companies are dependent on China for their defenses. Let the frontier labs win by being better; don’t let them define safety or security, or pull up the ladder of humanity’s collective knowledge. China is already hard enough to compete with; letting them carry the standard for openness and innovation is simply giving away our biggest advantage.
- Listen to Podcast (opens in new tab) Watch on YouTube (opens in new tab) Listen to this post: Log in to listen (opens in new tab)
I’m sympathetic to the cynics who consistently characterize Anthropic’s public statements, particularly those surrounding their model releases, as scare-mongering for the sake of marketing. It was only two months ago that Anthropic announced Mythos Preview, a model that they said was too dangerous to make publicly available, thanks in particular to its advanced cybersecurity capabilities. Then, two months later, the company publicly released Fable, a version of Mythos with various safety guardrails.
Fable is, in my limited experience, a very impressive model. It’s increasingly difficult to objectively evaluate models for anything other than coding performance, but there is subjective feel, and I found my interactions with Fable to be extremely impressive; it made other models, including GPT 5.5 and Opus 4.8, feel small and dumb. The two times I felt that way previously were with GPT-4 and Grok 4, both of which represented new generations in terms of base model size and complexity; my sense is that Fable is downstream of a new pre-train and the first of a new generation.
To that end, I can certainly buy the case that Fable/Mythos is in fact more capable when it comes to identifying and exploiting security issues, and that Anthropic’s cautious roll-out was justified. The problem with publicly releasing models, however, is that guardrails can be jailbroken, and apparently that is exactly what happened shortly after the release.
### Anthropic vs. the U.S. Government, Again
What happened next is somewhat unclear. Anthropic wrote in a blog post (opens in new tab):
The US government, citing national security authorities, has issued an export control directive to suspend all access to Fable 5 and Mythos 5 by any foreign national, whether inside or outside the United States, including foreign national Anthropic employees. The net effect of this order is that we must abruptly disable Fable 5 and Mythos 5 for all our customers to ensure compliance. Access to all other Anthropic models will not be affected.
We received the directive from the government today at 5:21pm (ET). The letter did not provide specific details of its national security concern. Our understanding is that the government believes it has become aware of a method of bypassing, or “jailbreaking” Fable 5. We reviewed a demonstration of this specific technique being used to identify a small number of previously known, minor vulnerabilities. These vulnerabilities all appear relatively simple, and we have found that other publicly-available models are able to discover them as well without requiring a bypass. Anthropic went on to make the case that non-universal jailbreaks were inevitable and also narrow, and that there was no evidence of a universal jailbreak; the jailbreak that was found, meanwhile, appears to have been reported by Amazon (opens in new tab), which is notable given Amazon is both an investor in Anthropic and a major provider of inference to the company. As I write this, senior Anthropic staff are in Washington D.C. (opens in new tab) seeking to resolve what they insist is a misunderstanding, and which White House officials are suggesting is insouciance by the company’s leadership to legitimate national security concerns. I don’t actually have much to add to the current conflict given how many facts are in dispute; what I am not surprised about is the fact that the conflict is happening: I already explained in Anthropic and Alignment (opens in new tab) why conflict between the U.S. government and Anthropic was inevitable. To that end, people who are arguing that Mythos isn’t powerful enough to warrant the government’s drastic action are missing the point: if it’s not powerful enough now, the next one will be, or the one after that, particularly now that models are increasingly useful in creating their successors. That, however, raises another question — one that seems to validate the cynics’ viewpoint: if Mythos is so dangerous, why even release Fable in the first place, and why fight with the government doing exactly what you claim to want? In fact, I think that Anthropic’s actions are quite understandable; what makes the company unique is how it justifies them, and it is those justifications that both give the cynics their fuel and Anthropic its magic.
The Economic Imperative
For the first few years of AI the most economic value has flown to compute, for obvious reasons: we don’t have enough supply to meet demand, which has meant skyrocketing prices; the biggest beneficiaries have been Nvidia, TSMC, and the memory makers (SK hynix, Samsung, and Micron). Anthropic and OpenAI, meanwhile, have collectively lost tens of billions of dollars building leading-edge models that, once released, are distilled and commoditized by open source models, primarily from China. This represents the bear case for the labs — they never cover their costs because their differentiation is fleeting, while free alternatives become “good enough” — and I think it’s a legitimate one. A world where models are interchangeable is one where models are commodities, while most of the value flows elsewhere. Right now that’s compute, but in the fullness of time, whenever we have enough compute, the most valuable place to be in the value chain will be the place that has always been the most valuable: owning the user touchpoint. To that end, it has long been clear to me that the frontier labs have the economic imperative to move closer to the user. If you own the user touchpoint, then you have meaningful lock-in, and the best way to own the user touchpoint is to be the canvas for everything they need to do. This, by extension, means that the frontier labs are on a collision course with software companies: it’s software that owns the user touchpoint, and it’s in the frontier labs’ long-term interest to not simply be a commodity input into software but to simply replace software outright. Software companies, meanwhile, are working to do the opposite. Satya Nadella laid out his vision for how companies should build on models in an essay on X (opens in new tab): Every company is going to have to build what I think of as human capital and token capital. Human capital comprises the knowledge, judgment, relationships, ingenuity, and pattern recognition of its people, while token capital is the firm’s AI capability it builds and owns. Importantly, human capital does not become less valuable as token capital grows. It only becomes more valuable! I believe human agency will be the driver of token capital growth. Humans will set ambitious goals, connect dots across domains, build relationships, and recognize patterns that matter most. Without human direction, you have compute running in circles.
This means the real opportunity is not in picking the best model but instead in building a learning loop on top of models where human capital and token capital compound. You can offload a task, or even a job, but you can never offload your learning. The future of the firm is the ability to compound that learning across people and AI. This requires a new architectural approach where every business is able to build agentic systems that improve over time, while still retaining control over their IP. A company should be able to switch out a “generalist” model without losing the “company veteran” expertise built into their learning system. This is the key “test” of your control and sovereignty in the era ahead. Nadella set this vision off with a warning: The last thing any of us want is a world where every company across every sector is ceding value to a few models that eat everything they see. If all the value is accrued by only a few models, the political economy will simply not tolerate it. There is no societal permission for an AI future that hollows out entire industries.
Think about what happened in the first phase of globalization where entire industrial economies were hollowed out by outsourcing. The GDP numbers looked fine on the surface, but the displacement was real and the consequences are still being felt. Let us not bring that dynamic into the AI era, with a small number of AI systems capturing all the economic returns, while entire industries find their knowledge commoditized right out from underneath them. Here’s the problem with that analogy: the globalization happened, and the industrial economies were hollowed out. There’s a possibility that this isn’t a warning but a prophecy; small wonder Nadella is raising the alarm given that Microsoft could be one of the casualties. And, by the same token, the economic imperative for the model makers is to accomplish exactly this.
The Data Imperative
The models — not even Mythos — are not yet at this point. What they need, beyond more compute, is more and better data. Model improvements increasingly come from reinforcement learning; some of this can be generated synthetically, but the most powerful lever for a frontier lab is real world use. This, I think, is a major reason why both OpenAI and Anthropic offer their heavily subsidized subscription plans. SemiAnalysis recently estimated (opens in new tab) that a $200 plan gets you $8,000 worth of Claude tokens and $14,000 worth of Codex tokens. Of course both are fighting for user and developer mindshare, but they’re also fighting to have access to actual usage data to make their models better. Anthropic upped the ante in a major way with Fable, announcing that they would retain the data for all usage for 30 days, even for their enterprise plans that previously promised zero data retention. The company said they would not train on this data, but they didn’t put in any sort of safeguards to guarantee they wouldn’t do so in the future (like storing the data with a third party). If this policy change (whenever Fable is restored) doesn’t lead to a significant loss of customers, I suspect it’s only a matter of time until they start using the data: it’s simply too valuable to their end goals. Note also the virtuous cycle with moving up into user touchpoints: the more workflows that are done directly with Claude or Codex, the more data each company gets to feed back into their training, which makes their products that much more capable and useful, expanding the number of workflows they can serve, expanding their access to data. Nadella, in his essay, highlights the importance of this data, but naturally thinks it should be independent from the model: Companies need to turn their workflows, domain knowledge, and accumulated judgment into AI systems that improve with each use. Private evals should capture whether a model is actually improving against outcomes that matter to the business (not just external benchmarks!). Private reinforcement learning environments should let models grow stronger on real traces from inside the organization. Its knowledge base makes institutional memory queryable and use of tokens more efficient.
This loop becomes the new IP of the firm. I think of it as a hill climbing machine. And unlike most assets, it compounds. Every improved workflow generates better training signal, which accelerates the accumulation of tacit knowledge unique to the firm. The companies that build this early will have an advantage that is hard to replicate, regardless of any new individual model capability. What if, however, the companies that give in to Anthropic’s data policies get better results right now? Or what if existing companies resist, leaving the door open for new companies — or the model makers themselves — to outcompete them in the market? Anthropic is certainly putting the resolve Nadella is calling for to the test.
The Power Imperative
The data retention policies around Fable/Mythos were, amazingly enough, not even the most controversial part of the launch. Rather, Anthropic said at launch that it would silently degrade Fable performance if it were used for LLM development; from the System Card (opens in new tab): We have also added safeguards related to frontier LLM development. As discussed in Section 6.1 of our February 2026 Risk Report (opens in new tab), we are concerned about the risks of accelerating the overall pace of AI development, though we remain uncertain about the severity of these risks. In particular, our concern is with — as we wrote then — “accelerating other AI developers in building powerful AI systems that pose similar risks to the ones ours pose – without necessarily having commensurate safeguards.”
In light of the ability of recent models to accelerate their own development (opens in new tab), we’ve implemented new interventions that limit Claude’s effectiveness for requests targeting frontier LLM development (for example, on building pretraining pipelines, distributed training infrastructure, or ML accelerator design). Using Claude to develop competing models already violates our Terms of Service (opens in new tab), but enforcing this restriction through our safeguards avoids accelerating the actors most willing to violate these terms.
Unlike our interventions for cybersecurity, biology and chemistry, and distillation attempts, these safeguards will not be visible to the user. Fable 5 will not fall back to a different model. Instead, the safeguards will limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning (PEFT). These interventions will not affect the vast majority of coding work. We estimate they will impact ~0.03% of traffic, concentrated in fewer than 0.1% of organizations. When these interventions are active, we expect them to have minimal behavioral impact on the model except to limit its effectiveness in developing frontier LLMs. Claude will still respond helpfully to user requests. We’ll continue to improve the precision of our detection methods following the launch of this model. Anthropic walked back this change (opens in new tab) — Fable will simply hand off LLM-related requests to Opus 4.8, and disclose this hand off to the user — but I think the initial policy was very illuminating. On one hand, I actually don’t begrudge Anthropic not wanting to help its competitors; on the other hand, what should be blisteringly clear is that Anthropic does not think that anyone else other than them should even be making frontier LLMs. What makes this policy all the more remarkable is the fact that it was enacted only two months after Anthropic had that dispute with the Department of War: the latter wanted to use Claude for any legal use, while the former wanted more stringent controls around surveillance and autonomous weapons. What this degradation represented was both the capability and willingness of Anthropic to silently alter its models to achieve its policy preferences. In other words, Anthropic willfully validated some of its critics’ worst fears in terms of being a supply chain risk. The broader takeaway from that previous episode, however, is that Anthropic believes that they are the ones who should have final say over how Anthropic is used; given that they think only they should be developing leading edge AI, they by extension think that only they should have final say over AI generally. When you further combine this realization with the company’s pronouncements about AI’s ability to conduct all economic activity, you realize that Anthropic’s leadership effectively wants to have power over everything and everyone.
The Safety Story
Of course Anthropic would never put things so baldly; the story, rather, is safety:
- I expect Anthropic to increasingly expose their model’s capabilities to end users through endpoints increasingly tailored to different workflows, even as they start to restrict the API. This replacement of software and restriction of access will be done in the name of safety, even as Anthropic fulfills its economic imperative of getting closer to end users.
- Anthropic’s explanation for their dramatic change in their data retention policy was safety. Specifically, the company claims that retaining all user data for 30 days is necessary to prevent the jailbreaks the U.S. government is worried about. I can certainly imagine a future where safety compels them to train on this data as well, to better protect against malicious usage.
- The entire Anthropic origin story is rooted in the founders’ belief that OpenAI wasn’t taking safety seriously enough; the company believes that only they can control AI, and that because they uniquely care about safety, they are justified in trying to control everyone else, up to and including the U.S. government. Here’s the thing about these safety justifications: I think they work because, to Anthropic, they aren’t justifications. The company really believes that they are the only ones who believe in super intelligence, and thus are the only ones who are sufficiently concerned about the dangers. That excuses decision after decision, policy after policy, and confrontation after confrontation that, to people on the outside, look like a bizarre combination of cynicism and naiveté. The contrast to OpenAI is massive: I think that one way to understand how and why OpenAI lost its lead is that, in the years following the release of ChatGPT, the company has been at war with itself internally as what used to be a research lab was suddenly seized with the burden of being the accidental consumer tech company (opens in new tab); to the extent OpenAI solved that conflict, it was by bleeding huge amounts of talent to Anthropic in particular. Anthropic, on the other hand, has perfect alignment between talent and mission and business. The company gets to sell to researchers the creation of a machine god, with the mantle of being the sort of person who cares about the dangers and is smart enough to navigate them on behalf of humanity; that every policy change that falls out of that happens to be great for business is the most beautiful coincidence in the world. I respect this alignment, and I fear it. I respect it because it is so clearly effective; the closest analogy is probably Apple, which has always framed every self-serving action in the guise of doing right by users — and often they were. So it is with Anthropic. What I fear, however, is that it is one thing to have people convinced they know best building a smartphone that I can take or leave; it’s considerably more concerning to have them building superintelligence that has the potential to rival or exceed the power of nation states, or merely massive corporations. The history of brilliant people convinced they know what humanity needs is a sordid one, precisely because they have convinced themselves that their intentions are good, justifying actions that very much are not.
- Listen to Podcast (opens in new tab) Watch on YouTube (opens in new tab) Listen to this post: Log in to listen (opens in new tab) Apple fans would, for years and years, sneer at Microsoft’s penchant for talking about products that may or may not ship, deriding them as vaporware. After Apple’s bungled 2024 launch of Apple Intelligence and new Siri (opens in new tab), however, vaporware is fair game, and just in time for this Article. ### Project Solara Last week, at its annual Build developer conference, Microsoft put forth a vision for a new ecosystem of hardware devices under the banner of Project Solara (opens in new tab):
The concept — which isn’t entirely clear from that video, but was [more fully explained on stage](https://www.youtube.com/watch?v=hbnhMKckxTE) — is that in the future you will be surrounded by an ecosystem of devices, none of which stand alone, but are more like portals to interact with your agents, which live in the cloud. In other words, as I wrote in February, [Thin Is In](https://stratechery.com/2026/thin-is-in/):
> This is even clearer when you consider the next big wave of AI: agents. The point of an agent is not to use the computer for you; it’s to accomplish a specific task. Everything between the request and the result, at least in theory, should be invisible to the user. This is the concept of a thin client taken to the absolute extreme: it’s not just that you don’t need any local compute to get an answer from a chatbot; you don’t need any local compute to accomplish real work. The AI on the server does it all.
I made the case in that Article that server-side inference would dominate AI workloads, thanks in particular to increasingly high memory demands for agents. What I found intriguing about Microsoft’s vaporware, however, is that it showcased a use case wherein this thin client approach was compelling for reasons beyond KV cache.
Specifically, for most of tech history *computing* has been indistinguishable from *interacting*; that’s why we place so much value on new input methods, as they often set off new paradigm shifts. By the same token, the problem with wearables as the paradigm beyond the iPhone is that interacting with them generally sucks. Sure, you can imagine a future where voice interaction is completely seamless or where a device can “see” what you see, but anything longer than a few seconds is much less convenient than simply swiping on your phone. Agents, however, compute on your behalf, without any interaction necessary: a few seconds is all you need to get work done for hours — at least in theory.
### Siri AI
Apple, a company that can actually make devices, was under heavy scrutiny going into yesterday’s WWDC keynote for a different concern: can the company make AI? And, if your standards are the state of the art in AI circa June 2024, when Apple took their first crack at answering the question, they did quite well. The company’s pre-recorded keynote took great pains to show actual demos — spinning indicators and all — and they worked! Here was the first one of what Apple is calling “Siri AI”:What’s fascinating about this specific demo is that it also showed just how far behind Apple is. New head of Siri Mike Rockwell successfully used Siri to set a reminder to enter a lottery for concert tickets, demonstrating context awareness and the ability to interact with the Reminders app through Apple’s App Intents framework; what would have been state of the art would have been asking Siri to enter the lottery on his behalf when the time came. In other words, to act outside of the interaction paradigm that has traditionally defined computing, and which Apple has dominated.
At the same time, the fact that Apple is behind the state of the art might not matter that much given Apple’s market and opportunity in that market. To start with the former, Apple is targeting consumers, for whom traditional chatbot functionality is probably sufficient for the vast majority of their AI needs. Siri will be able to give you recipes, tips on do-it-yourself projects, or generate images. Moreover, the fact that Siri will have access to your iPhone gives it all of the same advantages that made me optimistic about Apple Intelligence in the first place. From [an Update after that initial June 2024 launch](https://stratechery.com/2024/wwdc-apple-intelligence-apple-aggregates-ai/):
> The key part here is the “understanding personal context” bit: Apple Intelligence will know more about you than any other AI, because your phone knows more about you than any other device (and knows what you are looking at whenever you invoke Apple Intelligence); this, by extension, explains why the infrastructure and privacy parts are so important.
>
> What this means is that Apple Intelligence is by-and-large focused on specific use cases where that knowledge is useful; that means the problem space that Apple Intelligence is trying to solve is constrained and grounded — both figuratively and literally — in areas where it is much less likely that the AI screws up. In other words, Apple is addressing a space that is very useful, that only they can address, and which also happens to be “safe” in terms of reputation risk. Honestly, it almost seems unfair — or, to put it another way, it speaks to what a massive advantage there is for a trusted platform. Apple gets to solve real problems in meaningful ways with low risk, and that’s exactly what they are doing.
Apple actually made this version of Siri much more capable in terms of accessing world knowledge and image generation, which should make the experience much more seamless, but the real differentiation will clearly be that access to your personal information. You can ask Siri about something you received in messages — or was it email, or a voicemail? — and it will actually find what you’re looking for; it can also “see” what you are looking at on your screen, and act on the information. And, to the extent that third-party apps offer up their data to the Spotlight semantic index, and make actions available via App Intents, Siri can actually operate across different services in a way other AIs can not, at least without making massive sacrifices in security on a local Mac or PC.
### The Consumer Market
These capabilities are genuinely useful, and there’s a good chance they’re enough, at least for now, and that’s because there is another aspect of the consumer market that is worth considering — beyond the fact that billions of consumers already have iPhones. Specifically, consumers don’t want to work, and don’t really care about being productive.
This reality about the consumer market is a lesson that Silicon Valley has to re-learn every decade or so. Consider Dropbox, whose founder, Drew Houston, [is in the process of stepping down](https://www.cnbc.com/2026/05/26/dropbox-ceo-drew-houston-ashraf-alkarmi.html). Dropbox was a category-defining product that had a viral hook — if someone signed up with your referral code, you got more storage — and grew extremely fast amongst consumers; the company then spent too long trying to actually build a business in the consumer space, [before finally realizing](https://stratechery.com/2015/dropbox-kills-carousel-and-mailbox-facebook-kills-creative-labs/) that the only way to make money with what was ultimately a productivity product was by selling to enterprise.
The reason is obvious when you think about it: enterprises are paying for their employees’ time, so of course they are willing to pay for tools that make those employees more productive; consumers, on the other hand, are mostly looking to waste time, which is why attention-harvesting advertising is the only software business model that works at scale for consumer services. The fact that Silicon Valley forgets this is downstream from Silicon Valley being a bubble; normal people aren’t looking for agents to buy them tickets to a concert.
Still, the bubble was strong enough to convince OpenAI to make the exact same mistake Dropbox did: the company somehow convinced itself that it could make enough money selling subscriptions to consumers; Anthropic, meanwhile, realized that it was enterprises who were willing to pay for AI’s massive productivity benefits, even as OpenAI failed to capitalize on their consumer market penetration by [refusing to build an advertising product](https://stratechery.com/2025/an-interview-with-openai-ceo-sam-altman-about-building-a-consumer-tech-company/#advertising).
This is a long-winded way of saying that I don’t think that Apple’s agentic shortcomings are a big deal, at least for now. Agents help you do work and be more productive, and consumers don’t want to work or care about being productive. What they do want to do is watch short-form video, and an iPhone is simply much better at that than any other device ever will be; in that context, Siri being good enough is enough, and it appears that Apple crossed that bar.
### The iPhone’s Centrality
There are actually a lot of interesting technical details about how Apple rebuilt Siri, including expanding Private Cloud Compute to include Nvidia chips running in Google data centers, as well as a 20 billion parameter on-device mixture-of-experts model that selects the expert on a per-query basis (as opposed to on a per-token basis) so that it can run in an iPhone’s limited memory.
The key strategic takeaway of these implementation details, however, is the centrality of the iPhone. Microsoft’s Project Solara obviously makes sense for Microsoft given the fact that the company missed out on mobile, but it also fits with the infrastructure of AI, which is in the cloud, and increasingly about compute happening without a human in the loop. Apple, in contrast, is heavily incentivized to preserve the iPhone’s importance, and by extension, to focus on use cases organized around human interaction.
However, it’s too simplistic to reduce these approaches to a cynical analysis of incentives; both make sense in their own right. What makes me intrigued about Project Solara is the fact that Microsoft is positioning it as purely an enterprise play, which is important because an enterprise has context about the work being done, making it more viable to build long-running agents — which the enterprise is willing to pay for. That context would be far more difficult to build for consumers, given the need to tie together a huge number of services to get a coherent set of data over which to operate. Indeed, the only entities that can probably pull that off are Google and Apple via Android and iOS, respectively — and Google is always going to be focused more on its cloud services as the point of integration instead of the device.
That leaves Apple as the only company truly — dare I say it? — thinking differently. And yes, the iPhone as the true core of Siri (which will work across your devices, but get its differentiated context first-and-foremost from your iPhone) just so happens to perfectly align with Apple’s business model and desire to not spend billions in capex, but that doesn’t mean it’s the wrong approach. You’ll be able to access all of that capex that other companies are building on your phone, you’ll just have to use an app; if you need to find something personal, or work across apps, Siri will be the only one who can pull it off — as long as it’s not vaporware (and it appears the second time is the charm).
---

---- Listen to Podcast (opens in new tab) Watch on YouTube (opens in new tab) Listen to this post: Log in to listen (opens in new tab)
It’s hardly the biggest problem in the world — or perhaps the height of privilege to consider it a problem at all — but one of the most annoying consumer experiences is booking an Uber Black and realizing you got assigned a Tesla Model Y (Uber finally stopped allowing new Model Y’s onto Black last year (opens in new tab)). Buckle up for an uncomfortable back seat, basic plastic finishes, and, all-too-often, potential car sickness from a driver who hasn’t completely mastered the Tesla’s aggressive regenerative braking.
Still, the fact that the Model Y ever made it to the Black level is a testament to the brand Elon Musk built. Back in 2016, when 300,000 people dropped $1,000 each in a matter of hours to reserve an as-yet-unreleased Model 3, I explained that the phenomenon was because It’s a Tesla (opens in new tab):
The real payoff of Musk’s “Master Plan” is the fact that Tesla means something: yes, it stands for sustainability and caring for the environment, but more important is that Tesla also means amazing performance and Silicon Valley cool. To be sure, Tesla’s focus on the high end has helped them move down the cost curve, but it was Musk’s insistence on making “An electric car without compromises” that ultimately led to 276,000 people reserving a Model 3, many without even seeing the car: after all, it’s a Tesla. This is the same brand halo that landed what is, if we’re honest, a pretty basic car on the Uber Black list. What actually makes these cars compelling is the extent to which they are computers on wheels: I know plenty of very rich people who drive a Tesla not for the finishes but rather the Full Self-Driving (Supervised); there is nothing like it on the market, at least when it comes to cars you can own. Tesla appears to be doubling down on this point of differentiation: the company stopped production of the Models S and X (opens in new tab) earlier this year, focusing production resources on the CyberCab and robots; if you want your car to drive itself, you’ll get the same model as everyone else. It reminds me of Andy Warhol’s famous quote (opens in new tab): What’s great about this country is that America started the tradition where the richest consumers buy essentially the same things as the poorest. You can be watching TV and see Coca-Cola, and you know that the President drinks Coke, Liz Taylor drinks Coke, and just think, you can drink Coke, too. A Coke is a Coke and no amount of money can get you a better Coke than the one the bum on the corner is drinking. All the Cokes are the same and all the Cokes are good. Liz Taylor knows it, the President knows it, the bum knows it, and you know it. That “tradition” is scale, and America is indeed better at it than any other country in the world; and, amongst Americans, no one pursues and seeks to leverage scale quite like Musk.
Starlink and Airlines
From a press release (opens in new tab) from American Airlines: American Airlines today announced a sweeping modernization of its narrowbody inflight customer experience with the installation of Starlink, the fastest Wi-Fi in the sky, on more than 500 narrowbody aircraft beginning in Q1 2027. Starlink is widely regarded as the world’s most advanced satellite constellation using a low Earth orbit to deliver broadband Internet capable of supporting inflight streaming, online gaming, collaborative meeting tools and more. With thousands of satellites in low Earth orbit, Starlink can deliver multigigabit connectivity to aircraft using its Aero Terminal, which can support up to 1 Gbps per antenna.
“As a premium global airline, we are continuously seeking out world-class partners like Starlink to deliver what our customers need and want,” said American Airlines Chief Customer Officer Heather Garboden. “The addition of Starlink solidifies American as a leading airline in keeping passengers connected in flight.” As part of American’s commitment to an elevated onboard experience, Starlink will enable seamless streaming, browsing and real-time communication capabilities across American’s domestic and short-haul international routes. I linked to the press release just for the amusement of American Airlines, which has in recent years built its strategy around offering anything-but-premium on routes you need, billing their Starlink deal as a commitment to “an elevated onboard experience.” That may have been the argument for United’s Starlink deal when it was announced in 2024 (opens in new tab), but by this point it’s tablestakes (opens in new tab), which is surely exactly how Musk wants it. Starlink is the consumer-facing business of SpaceX, generating $8.7 billion in revenue last year and $4.4 billion in profit; while it’s not totally clear exactly how SpaceX accounts for launch costs, obviously Starlink benefits greatly from the fact that it has access to SpaceX’s launch capacity. That launch capacity has resulted in over ten thousand active satellites in low Earth orbit, delivering low latency high speed Internet anywhere in the world — including in the air. That’s the carrot for airlines; the stick is the prospect of everyone else having the same service, and customers making flight decisions based on the quality of Internet access available. There is a similarity to Tesla in this way. Musk companies at their best don’t win the game; they change the rules through scale, such that billionaires buy economy cars because they actually drive themselves (with supervision), and airlines transform the consumer experience on their own dime. Musk makes all-in bets — whether that be in terms of launch capacity or in autonomous driving (opens in new tab) — not by making rational short-term business decisions, but by starting with the desired end state and working backwards.
SpaceX’s Silly S-1
Tech has a long history of silly charts — there is an entire category known as Bezos charts (opens in new tab) — and the SpaceX S-1 (opens in new tab) has one that made me laugh. It came in the discussion of SpaceX’s total addressable market: We believe we have identified the largest actionable total addressable market (“TAM”) in human history. We estimate that our quantifiable TAM is $28.5 trillion, consisting of $370 billion in Space from space-enabled solutions; $1.6 trillion in Connectivity across $870 billion in Starlink Broadband and $740 billion in Starlink Mobile as well as additional opportunities in enterprise and government; $26.5 trillion in AI across $2.4 trillion in AI infrastructure, $760 billion in consumer subscriptions, $600 billion in digital advertising, and $22.7 trillion in enterprise applications. For illustrative purposes of sizing our addressable market opportunity, we exclude China and Russia from our global estimates.
This image is approximately to scale vertically, but certainly not horizontally: I could use the help in really wrapping my mind around the $26.5 trillion AI opportunity, given it’s more than 13 times the space and connectivity opportunity combined!
In all seriousness, the numbers are obviously absurd, but then again, everything about this IPO is absurd. SpaceX is seeking a $2 trillion valuation on a mere $18.67 billion in revenue with $4.9 billion in losses last year, and growth actually slowed from 35% to 33%. That slowdown happened despite the addition of xAI (and thus also X), which tipped the company from a small profit to that massive loss, thanks to $5.1 billion in AI R&D expense. That R&D, keep in mind, went towards building a model that is in 5th place, and whose entire founding team recently left the company. But sure, $26.5 trillion AI opportunity!
This is not to say that SpaceX won’t get its desired valuation. Tesla’s valuation never made any sense right up until the Models 3 and Y actually worked out, causing Tesla’s share price to soar (and even then it was hard to ever build a financial model that justified the new share price). Musk’s ability to make his own reality starts with investors; from 2021’s Mistakes and Memes (opens in new tab) and comparing Apple and Tesla:
This comparison works as far as it goes, but it doesn’t tell the entire story: after all, Apple’s brand was derived from decades building products, which had made it the most profitable company in the world. Tesla, meanwhile, always seemed to be weeks from going bankrupt, at least until it issued ever more stock, strengthening the conviction of Tesla skeptics and shorts. That, though, was the crazy thing: you would think that issuing stock would lead to Tesla’s stock price slumping; after all, existing shares were being diluted. Time after time, though, Tesla announcements about stock issuances would lead to the stock going up. It didn’t make any sense, at least if you thought about the stock as representing a company.
It turned out, though, that TSLA was itself a meme, one about a car company, but also sustainability, and most of all, about Elon Musk himself. Issuing more stock was not diluting existing shareholders; it was extending the opportunity to propagate the TSLA meme to that many more people, and while Musk’s haters multiplied, so did his fans. The Internet, after all, is about abundance, not scarcity. The end result is that instead of infrastructure leading to a movement, a movement, via the stock market, funded the building out of infrastructure. I explained in that Article why I generally did not cover Tesla’s financial results, and the reasoning extends to why I don’t expect to cover SpaceX’s: Musk is the master of memes, and is himself a meme. He offers a dream — Mars, fully autonomous vehicles, an addressable market of $28.5 trillion — and positions his companies and their stock as access to that dream, and through the alchemy of capital markets, transforms shared delusion into mass market reality. Musk’s track record matters in this regard. Building an electric car company was possible, as was full self-driving (supervised); at the same time there were ever increasing government mandates and programs around decreasing emissions that acted as the stick to Tesla’s carrot. Similarly, landing rockets was possible, and the new market creation downstream from correspondingly lower launch costs was comprehensible. That Musk succeeded in both instances gives him the benefit of the doubt. The question that matters, then, is not if the numbers make sense right now (they absolutely do not); what matters is if the dream is even possible, and if there are actual reasons to think it might happen. I think that data centers in space meet these conditions.
The Case for Data Centers in Space
The first question about data centers in space is if they are even possible, and I think the answer is clearly yes. The key thing to consider is that there is no requirement that these data centers look anything like data centers on earth. On earth we build massive buildings full of GPUs with massive infrastructure for cooling those GPUs and massive power plants (or a connection to a grid which connects to massive power plants) to power those GPUs. The idea of transporting these massive structures to space sounds implausible, and it is! However, there is no reason that space data centers would look like data centers on earth. What makes far more sense is to think about an individual satellite as something akin to a rack. Right now the largest Starlink satellite in orbit is the V2 Mini Direct-to-Cell, which measures 7.4 meters by 2.7 meters by 0.3 meters (estimated); an NVL72 rack from Nvidia, meanwhile, measures 2.2 meters by 1.1 meters by 0.6 meters, so we’re already in the right size range. The V2 Mini Direct-to-Cell consumes (and dissipates) up to an estimated 25kW of energy; the NVL72 up to 135kW, and it can fit a 1 trillion parameter model quantized to FP4. The big shortcoming for a rack-satellite is power and its dissipation, but going from 25kW to 135kW is certainly within the realm of possibility — and given that you don’t need much of the cooling and power distribution usage on earth, something closer to 100kW might deliver similar performance. There are other issues to address, including the problem of radiation screwing with calculations, reliability, etc., although those two concerns could be addressed in part by using larger chips (which are less efficient, but also use less power); these rack-satellites will also be disposable, like Starlink satellites, ameliorating reliability issues. The key factor, however, is that a fleet of racks, interconnected with lasers (as Starlink’s already are), each with their own solar panels and radiator arrays for cooling (deploying 200+ square meters of radiators per rack will be a huge challenge), is possible.
The next question about data centers in space is if there is a use case for them — the carrot — and I already made the argument that there is in The Inference Shift (opens in new tab). Specifically, there are three types of workloads developing around LLMs: training, answer inference, and agentic inference. From the section making the case for “agentic inference”:Critically, this articulation of an agentic-specific memory hierarchy implies a necessary trade-off of speed for capacity. Here’s the thing, though: lower speed isn’t nearly as important a consideration if there isn’t a human in the loop. If an agent is waiting around for a job that is being run overnight, the agent doesn’t know or care about the user experience impact; what is most important is being able to accomplish a task, and if entirely new approaches to memory make that possible, then delays are fine.
If delays are fine, then all of the focus on pure compute power and high-bandwidth memory seems out of place: if latency isn’t the top priority, then slower and cheaper memory — like traditional DRAM, for example — makes a lot more sense. And if the entire system is mostly waiting on memory, then chips don’t need to be as fast as the cutting edge either. This represents a profound shift in future architectures, but it also doesn’t mean that current architectures are going away:
- Training will continue to matter, and Nvidia’s current architecture, including high-speed compute, large amounts of high-bandwidth memory, and high-speed networking, will likely continue to dominate.
- Answer inference will be a meaningful market, albeit a relatively small one, and speed from chips like Cerebras or Groq (I explained how Nvidia is deploying Groq’s LPUs here (opens in new tab)) will be very useful.
- Agentic inference will gradually unbundle the GPU, which alternates between stranding high-bandwidth memory (during the prefill process) and stranding compute (during the decode process), in favor of increasingly sophisticated memory hierarchies dominated by high capacity and relatively lower cost memory types, with “good enough” compute; indeed, if anything it will be the speed of CPUs for things like tool use that will matter more than the speed of GPUs.
At the same time, these categories won’t be equal in size or importance. Specifically, agentic inference will be the largest market by far, because that is the market that won’t be limited by humans or time. Today’s agents are fancy answer inference; in the future true agentic inference will be work done by computers according to dictates given by other computers, and the market size scales not with humans but with compute. It’s agentic inference that makes the most sense for racks in space, and conveniently enough, that is also the market that is likely to be the largest in the long run.
The third question about data centers in space is if there is a stick. Specifically, while I think that racks-in-space are both a lot more viable than people think, and a lot more relevant to agentic inference than current modes of compute, it is at the end of the day cheaper and easier to build on earth, all things being equal. All things are not equal, however: right now we are at the very beginning of the AI buildout and already one of the biggest constraints is not just power (expected), but zoning (unexpected). I wrote in an Update last week (opens in new tab):That leads to an interesting contrast to globalization: when companies were closing down American factories and laying off workers and moving operations to China, none of the affected towns or workers had a say. They just suddenly no longer had a job, and a huge number of cities across the Rust Belt no longer had a reason to exist. People simply had to move, or worse, retreat to things like alcohol or drugs.
AI, however, is the opposite: building data centers requires permission, which is to say that people actually have a say. Again, I am not at all saying that these people are well informed about data centers, or about the economic impact on their communities, much less the economic impact of AI generally; what I am noting is that people who didn’t have a say in globalization are suddenly finding they do have a say about AI, and it’s not a surprise they are expressing their disapproval by blocking data centers. In that Update I made the case that data center builders — and by extension the companies that use them — should straight up pay people for permission to build data centers in their communities. At a minimum, however, that increases the costs of terrestrial data centers. What seems very plausible in the long run is that the demand for compute ends up being so large that there eventually is nowhere left to build, making the vast expanses of space not just an alternative but in fact the only choice.
An IPO Worth Supporting
If all of this happens — and there are a lot of “if”s here! — then suddenly that $2 trillion valuation starts looking reasonable. SpaceX is already monetizing xAI’s first data center, Colossus 1, to the tune of $15 billion/year for 300MW of capacity; that’s 3,000 racks-in-space. Anthropic, meanwhile, will probably make 3x the revenue on that capacity; it remains to be seen if xAI can get back in the state-of-the-art game, but if so then the amount of revenue it can generate per rack-in-space will be commensurately higher. Even without xAI, however, SpaceX has the potential to be a monopoly provider of marginal compute capacity. There are, needless to say, a massive number of assumptions baked into this argument, including assuming a huge number of engineering challenges are solved, Starship actually works, SpaceX gets sufficient supply of the right kinds of chips, compute demand is massively larger, agentic inference unbundles current architectures, and data center opponents are successful. The risk attached to all of these assumptions should discount the valuation you put on this business, which is to say I still think this IPO is nuts. At the same time, I’m glad it exists, for multiple reasons. The first one is the most obvious one: Musk, for all of his faults, has already pushed humanity forward on multiple vectors, including electric cars, self-driving, reusable rockets, satellite Internet, etc., and I’m excited to see him try and do more. The second is that I am in fact concerned about our ability to muster enough compute to fully realize the gains from AI, and am very worried about a replay of nuclear power, where our failure to build denied us the opportunity to even imagine what could be invented in a world of unlimited energy; the fact Musk is proposing an alternative path to unlimited compute is a relief. The third is that I appreciate the extent to which this IPO is a return to what an IPO should be: the opportunity for people to contribute capital to actually build the business, and to benefit if it works out. As I noted, I can’t make a financial model that necessarily justifies this valuation, particularly based on current financials, but neither can a VC investing in the Series A of a company. SpaceX has already invented a lot, and its early investors are going to make a lot of money with this IPO; at the same time, there is still so much more to invent that there remains a lot of upside — and, to be very clear, a lot of risk. It’s a testament to SpaceX’s ambitions that retail investors get to play VC. And hey, you get Mars upside for free!
- Listen to Podcast (opens in new tab) Watch on YouTube (opens in new tab) Listen to this post: Log in to listen (opens in new tab)
If you were looking for the ideal time to IPO, being a chip company in May 2026 is hard to beat. Reuters reported over the weekend (opens in new tab):
Cerebras Systems is set to raise the size and price of its initial public offering as soon as Monday, as demand for the artificial intelligence chipmaker’s shares continues to climb, two people familiar with the matter told Reuters on Sunday. The company is considering a new IPO price range of $150-$160 a share, up from $115-$125 a share, and raising the number of shares marketed to 30 million from 28 million, said the sources, who asked not to be identified because the information isn’t public yet. The fundamental driver of the ongoing surge in semiconductor stocks is, of course, AI, particularly the realization that agents are going to need a lot of compute (opens in new tab). What Cerebras represents, however, is something broader: while the compute story for AI has been largely about GPUs, particularly from Nvidia, the future is going to look increasingly heterogeneous.
The GPU Era
The story of how Graphics Processing Units became the center of AI is a well-trodden one, but in brief:
- Just as drawing pixels on a computer screen was a parallel process, which meant there was a direct connection between the number of processing units and graphics speed, making AI-related calculations was a parallel process, which meant there was a direct connection between the number of processing units and calculation speed.
- Nvidia enabled this dual-usage by making its graphics processors programmable, and created an entire software ecosystem called CUDA to make this programming accessible.
- The big difference between graphics and AI has been the size of the problem being solved — models are a lot bigger than video game textures — which has led to a dramatic expansion in high-bandwidth memory (HBM) per GPU, and dramatic innovations in terms of chip-to-chip networking to allow multiple chips to work together as one addressable system. Nvidia has been the leader in both. The number one use case for GPUs has been training, which stresses the third point in particular. While the calculations within each training step are massively parallel, the steps themselves are serial: every GPU has to share its results with every other GPU before the next step can begin. This is why a trillion-parameter model needs to fit in the aggregate memory of tens of thousands of GPUs that can communicate as one system. Nvidia dominates both problem spaces, first by securing HBM ahead of the rest of the industry, and second thanks to its investments in networking. Of course training isn’t the only AI workload: the other is inference. Inference has three main parts:
- Prefill encodes everything the LLM needs to know into an understandable state; this is highly parallelizable and compute matters.
- The first part of decode entails reading the KV cache — which stores context, including the output of the prefill step — to make an attention calculation. This is a serial step where bandwidth matters, but the memory requirements are variable and increasingly large.
-
The second part of decode is the feed-forward computation over the model weights; this is also a serial step where bandwidth matters, and the memory requirements are defined by the size of the model.
The two decode steps alternate for every layer of the model (they’re interleaved, not in sequence), which is to say that decode is serial and memory-bandwidth bound. For every token generated, two distinct memory pools must be read: the KV cache, which stores context and grows with each token, and the model weights themselves. Both must be read in full to produce a single output token.
GPUs handle all three needs: high compute for prefill, abundant HBM for KV cache and model weights, and chip-to-chip networking to pool memory across multiple chips when a single GPU isn’t enough. In other words, what works for training works for inference — look no further than the deal SpaceX made with Anthropic. From Anthropic’s blog (opens in new tab):
We’ve signed an agreement with SpaceX to use all of the compute capacity at their Colossus 1 data center. This gives us access to more than 300 megawatts of new capacity (over 220,000 NVIDIA GPUs) within the month. This additional capacity will directly improve capacity for Claude Pro and Claude Max subscribers. SpaceX retains Colossus 2 — presumably for both training of future models and inference of existing ones — and can afford to do both in the same data center precisely because xAI’s models aren’t getting much usage; more pertinently to this piece, they can do both in the same data center because both training and inference can be done on GPUs. Indeed, the GPUs Anthropic is contracting for at Colossus 1 were originally used for training as well; the fact that GPUs are so flexible is a big advantage.
Understanding Cerebras
Cerebras makes something completely different. While a silicon wafer has a diameter of 300mm, the “reticle limit” — the maximum area that a lithography tool can expose on that wafer — is around 26mm x 33mm. This is the effective size limit for chips; going beyond that entails linking two separate chips together over a chip-to-chip interposer, which is exactly what Nvidia has done with the B200. Cerebras, on the other hand, has invented a way to lay down wiring across the so-called “scribe lines” that are the boundary between reticle exposures, making the entire wafer into a single chip with no need for relatively slow chip-to-chip linkages. The net result is a chip with a lot of compute and a lot of SRAM that is blisteringly fast to access. To put it in numbers, the WSE-3 (Cerebras’ latest chip) has 44GB of on-chip SRAM at 21 PB/s of bandwidth; an H100 has 80GB of HBM at 3.35 TB/s. In other words, the WSE-3 has just over half the memory of an H100, but 6,000 times the memory bandwidth. The reason to compare the WSE-3 to an H100 is that the H100 is the chip most used for inference — and inference is clearly what Cerebras is most well-suited for. You can use Cerebras chips for training, but the chip-to-chip networking story isn’t very compelling, which is to say that all of that compute and on-chip memory is mostly just sitting around; what is much more interesting is the idea of getting a stream of tokens at dramatically faster speed than you can from a GPU. Note, however, that the limitation in terms of training also potentially applies in terms of inference: as long as everything fits in on-chip memory Cerebras’ speed is an incredible experience; the moment you need more memory, whether that be for a larger model or, more likely, a larger KV cache, then Cerebras doesn’t make much sense, particularly given the price. That whole-wafer-as-chip technique means high yields are a massive challenge, which hugely drives up costs. At the same time, I do think there will be a market for Cerebras-style chips: right now the company is highlighting the usefulness of speed for coding (opens in new tab) — reasoning means a lot of tokens, which means that dramatically scaling up tokens-per-second equals faster thinking — but I think this is a temporary use case, for reasons I’ll explain in a bit. What does matter is how long humans are waiting for an answer, and as products like AI wearables become more of a thing, the speed of interaction, particularly for voice — which will be a function of token generation speed — will have a tangible effect on the user experience.Agentic Inference
I have previously made the case, including in Agents Over Bubbles (opens in new tab), that we have gone through three inflection points in the LLM era: - ChatGPT demonstrated the utility of token prediction.
- o1 introduced the idea of reasoning, where more tokens meant better answers.
- Opus 4.5 and Claude Code introduced the first usable agents, which could actually accomplish tasks, using a combination of reasoning models and a harness that utilized tools, verified work, etc. All of this falls under the banner of “inference”, but I think it will be increasingly clear that there is a difference between providing an answer — what I will call “answer inference” — and doing a task — what I will call “agentic inference.” Cerebras’ target market is “answer inference”; in the long run, I think the architecture for “agentic inference” will look a lot different, not just from Cerebras’ approach, but from the GPU approach as well. I mentioned above that fast inference for coding is a temporary use case. Specifically, coding with LLMs requires a human in the loop. It’s the human that defines what is to be coded, checks the work, commits the pull request, etc.; it’s not hard to envision a future, however, where all of this is completely handled by machines. This will apply to agentic work broadly: the true power of agents will not be that they do work for humans, but rather that they do work without human involvement at all. This, by extension, will mean that the likely best approach to solving agentic inference will look a lot different than answer inference. The most important aspect for answer inference is token speed; the most important aspect for agentic inference, however, is memory. Agents need context, state, and history. Some of that will live as active KV cache; some will live in host memory or SSDs; much of it will live in databases, logs, embeddings, and object stores. The important point is that agentic inference will be less about GPUs answering a question and more about the memory hierarchy wrapped around a model. Critically, this articulation of an agentic-specific memory hierarchy implies a necessary trade-off of speed for capacity. Here’s the thing, though: lower speed isn’t nearly as important a consideration if there isn’t a human in the loop. If an agent is waiting around for a job that is being run overnight, the agent doesn’t know or care about the user experience impact; what is most important is being able to accomplish a task, and if entirely new approaches to memory make that possible, then delays are fine. Meanwhile, if delays are fine, then all of the focus on pure compute power and high-bandwidth memory seems out of place: if latency isn’t the top priority, then slower and cheaper memory — like traditional DRAM, for example — makes a lot more sense. And if the entire system is mostly waiting on memory, then chips don’t need to be as fast as the cutting edge either. This represents a profound shift in future architectures, but it also doesn’t mean that current architectures are going away:
- Training will continue to matter, and Nvidia’s current architecture, including high-speed compute, large amounts of high-bandwidth memory, and high-speed networking, will likely continue to dominate.
- Answer inference will be a meaningful market, albeit a relatively small one, and speed from chips like Cerebras or Groq (I explained how Nvidia is deploying Groq’s LPUs here (opens in new tab)) will be very useful.
-
Agentic inference will gradually unbundle the GPU, which alternates between stranding high-bandwidth memory (during the prefill process) and stranding compute (during the decode process), in favor of increasingly sophisticated memory hierarchies dominated by high capacity and relatively lower cost memory types, with “good enough” compute; indeed, if anything it will be the speed of CPUs for things like tool use that will matter more than the speed of GPUs.
At the same time, these categories won’t be equal in size or importance. Specifically, agentic inference will be the largest market by far, because that is the market that won’t be limited by humans or time. Today’s agents are fancy answer inference; in the future true agentic inference will be work done by computers according to dictates given by other computers, and the market size scales not with humans but with compute.
The Implications of Agentic Inference on Compute
To date the invocation of “scaling with compute” has implicitly meant Nvidia bullishness. However, much of Nvidia’s relative advantage to date has been a function of latency: Nvidia chips have fast compute, but keeping that compute busy has required big investments in ever-expanding HBM memory and networking. If latency isn’t the key constraint, however, then Nvidia’s approach seems less worth paying a premium for. Nvidia does recognize this shift: the company launched an inference framework called Dynamo (opens in new tab) that helps disaggregate different parts of inference, and is shipping products like standalone memory and CPU racks to enable increasingly large KV caches and faster tool use, the better to keep their expensive GPUs busy. Ultimately, however, it’s easy to see cost and simplicity being increasingly attractive to hyperscalers for agentic inference that isn’t remotely GPU-bound. China, meanwhile, for all of its lack of leading edge compute, has everything it needs for agentic inference: fast-enough (but not leading-edge) GPUs, fast-enough (but not leading-edge) CPUs, DRAM, hard drives, etc. The challenge, of course, is compute for training; it’s also possible that answer inference is more important for national security, at least when it comes to military applications. The other interesting angle is space: slower chips actually make space data centers more viable for a number of reasons. First, if memory can be offloaded, chips can be made much simpler and run much cooler. Second, older nodes, by virtue of being physically larger, will better withstand space radiation. Third, older nodes require less power, which means there will be less heat to dissipate via radiation. Fourth, not being on the bleeding edge will mean higher reliability, an important consideration given that satellites won’t be repairable. Nvidia CEO Jensen Huang regularly says that “Moore’s Law is Dead”; what he means is that the future of computing speed-ups will be a function of systems innovation, which is exactly what Nvidia has done. Maybe the most profound implication of agents that act without humans in the loop, however, will be that Moore’s Law doesn’t matter, and that the way we get more compute is by realizing that the compute we have is already good enough.
- Listen to Podcast (opens in new tab) Watch on YouTube (opens in new tab) Listen to this post: Log in to listen (opens in new tab)
When it comes to the AI soap opera — there is news every day, and the company on top and the bottom seems to shift by the quarter if not the month — the news that I find most intriguing and instructive this week is about physical goods and logistics. From Bloomberg (opens in new tab):
Amazon.com Inc. unveiled a suite of logistics services that will let businesses buy its existing freight and distribution offerings as a package, sending shares of rival delivery companies such as FedEx Corp. and United Parcel Service Inc. lower. The world’s largest online retailer on Monday announced Amazon Supply Chain Services (ASCS), offering other companies access to its “full portfolio” of supply-chain and distribution offerings. The service largely consolidates a package of existing products — air and ocean freight, trucking and last-mile delivery — into a new suite it says companies like Procter & Gamble Co. and 3M Co. are already using. This is a very satisfying announcement for Stratechery, given it’s the culmination of a prediction I made a decade ago in The Amazon Tax (opens in new tab). Amazon at that point had two primary businesses — Amazon.com and AWS — and I made the case in that Article that they were actually very similar: in both cases Amazon built “primitives” that had Amazon itself as their first, best customer, justifying and driving initial development, but in both cases the ultimate play was to sell those primitives to other companies. It was already clear at the time that logistics would follow the same path: It seems increasingly clear that Amazon intends to repeat the model when it comes to logistics: after experimenting with six planes last year the company recently leased 20 more to flesh out its private logistics network; this is on top of registering its China subsidiary as an ocean freight forwarder…
So how might this play out? Well, start with the fact that Amazon itself would be this logistics network’s first-and-best customer, just as was the case with AWS. This justifies the massive expenditure necessary to build out a logistics network that competes with UPS, FedEx, et al, and most outlets are framing these moves as a way for Amazon to rein in shipping costs and improve reliability, especially around the holidays.
However, I think it is a mistake to think that Amazon will stop there: just as they have with AWS and e-commerce distribution I expect the company to offer its logistics network to third parties, which will increase the returns to scale, and, by extension, deepen Amazon’s eventual moat. Now, ten years later, we are here, with the official unveiling of Amazon Supply Chain Services (opens in new tab), and I think the time frame is an important one: Amazon, more than any other company, actually operates with decade-long timeframes, consistently making real-world investments at massive scale that (1) convert their marginal costs into capital costs and (2) gain leverage on those capital costs by selling them to other businesses. This is, by the way, still a story about AI.
A Brief History of AWS
Three years ago SemiAnalysis wrote an Article entitled Amazon’s Cloud Crisis: How AWS Will Lose The Future Of Computing (opens in new tab), and I found it very compelling. First, though, some history (much of which is covered in SemiAnalysis’ article). Amazon not only invented cloud computing, but also realized it would be a commodity market. While most people in tech think about building sustainable differentiation that allows you to charge higher prices, thus producing profit, commodity markets work differently: there, sustainable profits come from having structurally cheaper costs. Amazon developed exactly that, first through having the largest scale — giving the company both buying power and also the most leverage on their development costs — and second through genuine innovation. AWS built a specialized system called Nitro, built on their own chips, that offloaded server management, including network management, storage management, hypervisor management, etc. from the expensive Intel and AMD servers that the company sold access to; this let Amazon run that many more virtual machines on a single server, significantly increasing utilization, i.e. delivering a structural cost advantage. Amazon doubled down on their custom chip efforts with Graviton, their ARM processors. Graviton chips, particularly the first few generations, were inferior to Intel or AMD chips, but that didn’t mean they were useless. By that time AWS had expanded from simply being an Infrastructure-as-a-Service (IaaS) provider to being a Platform-as-a-Service (PaaS) provider as well. IaaS means you provide raw compute, storage, etc., on which customers can run things like operating systems or databases; PaaS means you provide that basic functionality as a service. Amazon Relational Database Service (RDS), for example, is a fully managed database that customers can access via a set of APIs without having to worry about actually managing the full database themselves, worrying about scaling, duplication, etc. This, by extension, means that customers don’t need to know and don’t need to care about the compute infrastructure that undergirds services like RDS — which has long been Graviton! PaaS lets Amazon double-dip in terms of profitability: first, AWS could sell PaaS products at a higher margin than IaaS products, and second, the company could leverage its own cheaper silicon to serve those products, reducing their costs. Over time Graviton has become more competitive in performance — while still being cheaper — giving Amazon a lower-cost compute instance to sell to end users, but even without 3rd-party take-up the investment in building its own silicon has paid off over time.
Training vs. Inference
Fast forward to AI, and SemiAnalysis’ concern was that all of these optimizations left AWS ill-prepared for AI. One big problem was networking: Rather than implement the best networking from Nvidia and/or Broadcom, Amazon is using its own Nitro and Elastic Fabric Adaptor (EFA) networking. This works well for many workloads, plus it delivers a cost, performance, and security advantage. There are business, cultural, and security reasons why Amazon will not implement other networking. The cultural one is important. Nitro and networking SoC’s generally have been Amazon’s biggest cost advantage for years. It’s ingrained into their DNA. Even EFA delivers on this too, but they don’t see how new workloads are evolving and that a new tier is needed due to the lack of foresight in their internal workload and infrastructure teams. Amazon is making a deliberate choice of not adopting that we believe will bite them in the future. Another was Amazon’s insistence on building its own chips, which were not only inferior to the best Nvidia chips in terms of performance, but might also lead to them getting fewer Nvidia chips going forward: At least some other clouds will implement out-of-node NVLink. That’s where the discussion of prioritization now comes in. AI GPUs face tremendous shortages, for at least a full year. This is one of the most pivotal times for AI, and it may mark the haves and the have-nots. Nvidia is a complete monopoly right now. Why would Nvidia prioritize Amazon for these GPUs, when they know Amazon will move to their in-house chips as quickly as they can, for as many compute workloads as they can? Why would Nvidia ship tons of GPUs to the cloud that is not using any of their networking, thereby reducing their share of wallet?
Instead, Nvidia prioritizes the me-too clouds. Amazon does get meaningful volume, but nowhere close to where demand is. Amazon’s H100 GPU shipments relative to public cloud shipments is a significantly lower than their share of the public cloud. Those other clouds also can’t satisfy demand, but they get a bigger percentage of the GPUs they ask Nvidia for, and as such, firms looking for GPUs for training or inference will move to those clouds. Nvidia is the kingmaker right now, and they are capitalizing on it. They have to spread the balance of power out to prevent compute share from clustering towards Amazon. These concerns were well-founded in the 2023 time-period when that Article was written: that was a time when AI, thanks to ChatGPT, had hit the mainstream, but the largest share of compute still went to training. Training required all of the things that Amazon lacked, particularly the ability to network large numbers of Nvidia GPUs together into one coherent system. In such a system the most important capability was horizontal networking between chips, so that you could update weights during training, a step that needed to happen serially. It was absolutely the case that cloud providers like Microsoft or Oracle or the neoclouds, which implemented full Nvidia solutions, instead of the standalone HGX racks that AWS favored, were much better suited to training large language models. That is still the case, by the way. What has changed is that training is no longer the biggest AI compute market; inference is, thanks not only to increased AI adoption, but also because of fundamental changes in terms of how AI works. From an Update about Nvidia (opens in new tab):
- The first inflection point was the emergence of LLMs — call this the ChatGPT moment. In this first paradigm tokens were generated by GPUs and presented as the answer to a question.
- The second inflection point was the emergence of reasoning models — call this the o1 moment. In this paradigm there are a very large number of tokens that are generated to figure out the answer before the answer is actually generated; this was an exponential increase in the addressable market for tokens.
- The third inflection point was the emergence of functional agents — call this the Opus 4.5 moment. In this paradigm those reasoning models are not triggered by humans asking a question, but by an agent solving a problem. This increases the market in two directions: first, humans can run multiple agents, and secondly, agents can leverage reasoning models multiple times to accomplish a task. This isn’t just an exponential increase in the addressable market for tokens, it’s two exponential increases squared. Both the shift to inference and the shift in the nature of inference have been positives for AWS’ approach.
- First, while inference still requires significant memory, the requirement is significantly less than that required for training. It’s actually viable to store a model’s parameters in a single server; you don’t need to network together thousands of chips.
- Second, while reasoning and agentic workloads require significantly more tokens, and thus a massively larger KV cache, the increase is actually so large that even the most optimized Nvidia inference systems are being built with dedicated memory servers (opens in new tab). This sort of architecture is much more compatible with Amazon’s networking approach than the thousands-of-chips-networked-together approach is.
- Third, agents are heavily CPU dependent, which has two important implications. First, fully utilizing accelerators is a function of having sufficient general compute; second, achieving maximum utilization of heterogeneous compute means unbundling CPUs and GPUs and routing workloads between resources, which is exactly the sort of disaggregated-resource abstraction that Amazon has been building with Nitro. The utilization point is an important one. Nvidia CEO Jensen Huang made his case for Nvidia chips over custom ASICs at length at GTC 2025 (opens in new tab). Huang’s argument was that AI factories — to use his term — were ultimately constrained by power; that meant that the most important metric for profitability was not the cost of chips but rather tokens-per-watt. In other words, if you can’t increase watts, it’s worth spending more on chips to increase tokens on those watts. There are, however, three reasons why this argument may not hold, particularly for a company like Amazon.
- First, if you have the money to buy that many Nvidia chips, you also have the money to spend on getting more power — which is exactly what AWS has been focused on. This very much fits AWS’ modus operandi, which is to invest more upstream (in this case in power) with the goal of spending less downstream (paying Nvidia huge margins for their chips).
- Second, in the long term, electricity is more of a commodity than logic is. That means it is a market where innovation and competition are more likely to break a bottleneck, which is another way to say that investing in one’s own silicon is the area most likely to deliver a return on investment.
-
Third, the nature of inference workloads — particularly agentic ones — is such that perfect accelerator utilization is going to be a much harder problem to solve than when it comes to training.
These points are moot, however, if you don’t have your own logic chip that is at least competitive, and here Amazon’s long-term outlook is paying off. Amazon bought Annapurna Labs, which makes their chips, in 2015, and launched their first AI-focused chip in 2019. No, it wasn’t very good, but critically, that was seven years ago: now Trainium 3 is decent (opens in new tab) and the trajectory is even better. AWS is positioned to have a sustainable cost advantage for inference going forward.
AWS’s Neutrality
Moreover, they are already replaying the Graviton playbook. Trainium chips help undergird Bedrock, its AI platform, which is to say that users are using Trainium chips even if they didn’t explicitly choose to do so. AWS CEO Matt Garman made this point explicitly in a Stratechery Interview (opens in new tab):I think just with GPUs, by the way, you’re going to interact with a lot of these accelerator chips through abstractions. So the vast majority of customers don’t interact with GPUs either, except through maybe like in their laptop or something like that, for graphics. But when you’re talking to OpenAI, even if they’re running on GPUs, you’re not talking to the GPUs, if you’re talking to Claude, you’re through GPUs or Trainium or TPUs, you’re not talking to any of those chips, you’re talking to the interface. And the vast majority of inference out there is being done on one of a handful of models.
And so whether it’s 5, 10, 20, 100, it’s not millions of people that are programming to those things directly, and that’s gonna be true going forward just because these systems are so complex, they’re very large. If you’re going to go train a model, not that many people have enough money to go train a model, not that many people have the expertise to actually manage it. They’re very complicated systems, and the OpenAI team is incredible in their ability to squeeze value out of a very large compute cluster. But not that many people have the team that can do that, independent of what the chip happens to be, and so I think that that’s going to be true for all accelerator chips, honestly. The frontier models are an important factor in this, and that is an angle that I didn’t see coming. Nvidia CEO Jensen Huang explained in a recent interview with Dwarkesh Patel (opens in new tab) why Nvidia didn’t invest in Anthropic early on: At the time, I didn’t deeply internalize how difficult it would be to build a foundation AI lab like OpenAI and Anthropic, and the fact that they needed huge investments from the supplier themselves. We just weren’t in a position to make the multi-billion dollar investment into Anthropic so that they could use our compute. But Google and AWS were. They put in huge investments in the beginning so that Anthropic, in return, used their compute. We just weren’t in a position to do that at the time.
I would say my mistake is I didn’t deeply internalize that they really had no other options, that a VC would never put in $5-10 billion of investment into an AI lab with the hopes of it turning out to be Anthropic. So that was my miss. But even if I understood it, I don’t think we would’ve been in a position to do that at the time. But I’m not going to make that same mistake again. Amazon had both the money and the chips to invest into Anthropic precisely because they had built such a cash machine with AWS in the first place. That’s the thing with big investments in infrastructure: they take years to build, but the benefit of that investment compounds over time. Anthropic, meanwhile, thanks to those investments from Amazon and Google, can not only run across a variety of chips, but for a long time was the only frontier model available on all of the leading clouds, an important selling point for enterprises. Microsoft, in the end, needed to let go of Azure’s exclusive access to OpenAI’s API (opens in new tab) in part because that exclusivity was hurting the prospects of their mammoth stake in OpenAI. You can also make the case that Amazon is the best choice for frontier model access in a world of limited compute: Microsoft’s core business is software, which is to say that the company faces massive pressure to invest in their own AI capabilities, even at the cost of de-prioritizing cloud customers. That’s exactly what happened at Microsoft earlier this year (opens in new tab), when the company missed Azure growth projections because they devoted more compute to their internal workloads. It was an understandable decision: cloud demand is eternal, but the risk from AI for existing software businesses is existential. This also applies to Google: the company’s core business is also digital, and while search has fended off the threat from chatbots that many expected, the fundamental challenge is still one to be managed, not extinguished. Amazon’s core businesses, meanwhile, are very much rooted in the physical world: selling and shipping physical goods, and building data centers. Both are amenable to Amazon devoting the majority of its chips to customers’ workloads.
Amazon’s Future
If this week marks the resolution of one of Amazon’s long bets, you can see the outline of future resolutions in present day announcements. One prominent example is Amazon Leo, the company’s satellite service that seems, at first glance, duplicative of SpaceX’s Starlink, which has the advantage of already existing at scale. Remember Amazon’s formula, however, which CEO Andy Jassy stated explicitly with regards to Leo on the company’s most recent earnings call (opens in new tab):Today, if you ask what stops us from growing the business, we have to get the constellation into space. We have over 20 launches planned this year. We have over 30 launches planned in 2027. But I think the business has a chance to be a very large many billion-dollar revenue business. And I think it has some characteristics that are reminiscent of AWS in that it’s capital-intensive upfront where you’re committing a lot of capital and cash in the early years for assets that you get to leverage over a long period of time. And so I like the free cash flow and return on invested capital characteristics of that business in the medium to long term. The fact that it is extremely capital-intensive is not the only thing about Leo that makes it like AWS: a critical factor is that Amazon is the first-best customer to give the service scale, and here it’s worth going back to logistics. I noted above that Amazon delivery still has marginal costs, and that is because humans have to make the delivery. Amazon, however, has already pointed to the future, a full 13 years ago (opens in new tab) when the company first started talking publicly about drone delivery. It’s been a long slog, to be sure, but it’s increasingly plausible to imagine a future where delivery costs are a matter of depreciation on drone assets, and what would such a future require? How about reliable widespread satellite coverage for communicating with and guiding those drones? And, if Amazon doesn’t want to be dependent on Jensen Huang for chips, do you think they want to be dependent on Elon Musk for drone connectivity? Of course other businesses — like Apple (opens in new tab) — will be able to pay to use Amazon’s satellite infrastructure, just like they can now pay to use Amazon’s delivery service, or pay to use AWS, or pay to sell on Amazon.com. The world may change, in increasingly drastic ways, but Amazon’s approach, by virtue of its focus on long-term investments in the physical world, appears to be as sturdy as ever. More generally, I increasingly suspect that long-term vulnerability to AI — or, to put it more positively, long-term incentives to invest in AI — are very strongly correlated with the degree to which a company interacts with the physical world, and secondarily, the degree to which companies feel secure in their control of distribution:
- Apple and Amazon feel comfortable not having leading edge models, just access to them, because their business is rooted in the physical.
- Microsoft has invested heavily in data centers, but doesn’t own their own model, perhaps because they feel their control of distribution to enterprises will protect their core business (or because they had too much of a dependency on OpenAI).
- Google and Meta are investing at a similar scale to Amazon, and are also heavily invested in their own models. Both are Aggregators, which is to say they have to continually earn attention from consumers, given that competition is only a click away; having good AI is existential to them. This is, in the end, another advantage to making the sort of long-term bets Amazon specializes in: the threats are so distant that you have plenty of time to make new investments that address any weaknesses that develop in the meantime — or, as is the case of AI, wait for the market to tilt in your favor.
- Listen to Podcast (opens in new tab) Listen to this post: Log in to listen (opens in new tab)
It’s the nature of business that the eulogy for a chief executive doesn’t happen when they die, but when they retire, or, in the case of Apple CEO Tim Cook, announce that they will step up to the role of Executive Chairman on September 1 (opens in new tab). The one morbid exception is when a CEO dies on the job — or quits because they are dying — and the truth of the matter is that that is where any honest recounting of Cook’s incredibly successful tenure as Apple CEO, particularly from a financial perspective, has to begin.
The numbers, to be clear, are extraordinary. Cook became CEO of Apple on August 24, 2011, and in the intervening 15 years revenue has increased 303%, profit 354%, and the value of Apple has gone from $297 billion to $4 trillion, a staggering 1,251% increase.
Apple’s increase in market cap over Tim Cook’s tenure as CEO
The reason for Cook’s accession in 2011 became clear a mere six weeks later, when Steve Jobs passed away from cancer on October 5, 2011. Jobs’ death isn’t the reason Cook was chosen — Cook had already served as interim CEO while Jobs underwent treatment in 2009 — but I think the timing played a major role in making Cook arguably the greatest non-founder CEO of all time.
### Zero to One
Peter Thiel introduced the concept of Zero To One (opens in new tab) thusly: When we think about the future, we hope for a future of progress. That progress can take one of two forms. Horizontal or extensive progress means copying things that work — going from 1 to n. Horizontal progress is easy to imagine because we already know what it looks like. Vertical or intensive progress means doing new things — going from 0 to 1. Vertical progress is harder to imagine because it requires doing something nobody else has ever done. If you take one typewriter and build 100, you have made horizontal progress. If you have a typewriter and build a word processor, you have made vertical progress. Steve Jobs made 0 to 1 products, as he reminded the audience in the introduction to his most famous keynote (opens in new tab):
> Every once in a while, a revolutionary product comes along that changes everything. First of all, one’s very fortunate if one gets to work on one of these in your career. Apple’s been very fortunate: it’s been able to introduce a few of these into the world.
>
> In 1984, we introduced the Macintosh. It didn’t just change Apple, it changed the whole computer industry. In 2001, we introduced the first iPod. It didn’t just change the way we all listen to music, it changed the entire music industry.
>
> Well, today we’re introducing three revolutionary products of this class. The first one: a widescreen iPod with touch controls. The second: a revolutionary mobile phone. And the third is a breakthrough Internet communications device. Three things…are you getting it? These are not three separate devices. This is one device, and we are calling it iPhone.
Steve Jobs would, three years later, also introduce the iPad, which makes four distinct product categories if you’re counting. Perhaps the most important 0 to 1 product Jobs created, however, was Apple itself, which raises the question: what makes Apple Apple?
### The Cook Doctrine
“What Makes Apple Apple” isn’t a new question; it was the central question of Apple University, the internal training program the company launched in 2008. Apple University was hailed on the outside as a Steve Jobs creation, but while I’m sure he green lit the concept, it was clear to me as an intern on the Apple University team in 2010, that the program’s driving force was Tim Cook.
The core of the program, at least when I was there, was what became known as [The Cook Doctrine](https://asymco.com/2011/01/17/the-cook-doctrine/):
> We believe that we’re on the face of the Earth to make great products, and that’s not changing.
>
> We’re constantly focusing on innovating.
>
> We believe in the simple, not the complex.
>
> We believe that we need to own and control the primary technologies behind the products we make, and participate only in markets where we can make a significant contribution.
>
> We believe in saying no to thousands of projects so that we can really focus on the few that are truly important and meaningful to us.
>
> We believe in deep collaboration and cross-pollination of our groups, which allow us to innovate in a way that others cannot.
>
> And frankly, we don’t settle for anything less than excellence in every group in the company, and we have the self-honesty to admit when we’re wrong and the courage to change.
>
> And I think, regardless of who is in what job, those values are so embedded in this company that Apple will do extremely well.
Cook explained this on [Apple’s January 2009 earnings call](https://seekingalpha.com/article/115797-apple-inc-f1q09-qtr-end-12-27-08-earnings-call-transcript), during Jobs’ first leave of absence, in response to a question about how Apple would fare without its founder. It’s a brilliant statement, but it is — as the last paragraph makes clear — ultimately about maintaining, nurturing, and growing what Jobs built.
That is why I started this Article by highlighting the timing of Cook’s ascent to the CEO role. The challenge for CEOs following iconic founders is that the person who took the company from 0 to 1 usually sticks around for 2, 3, 4, etc.; by the time they step down the only way forward is often down. Jobs, however, by virtue of leaving the world too soon, left Apple only a few years after its most important 0 to 1 product ever, meaning it was Cook who was in charge of growing and expanding Apple’s most revolutionary device yet.
### Cook’s Triumphs
Cook, to be clear, managed this brilliantly. Under his watch the iPhone not only got better every year, but expanded its market to every carrier in basically every country, and expanded the line from one model in two colors to five models in a plethora of colors sold at the scale of hundreds of millions of units a year.
Cook was, without question, an operational genius. Moreover, this was clearly the case even before he scaled the iPhone to unimaginable scale. When Cook joined Apple in 1998 the company’s operations — centered on Apple’s own factories and warehouses — were a massive drag on the company; Cook methodically shut them down and shifted Apple’s manufacturing base to China, creating a just-in-time supply chain that year-after-year coordinated a worldwide network of suppliers to deliver Apple’s ever-expanding product line to customers’ doorsteps and a fleet of beautiful and brand-expanding stores. There was not, under Cook’s leadership, a single significant product issue or recall.
Cook also oversaw the introduction of major new products, most notably AirPods and Apple Watch; the “Wearables, Home, and Accessories” category delivered $35.4 billion in revenue last year, which would rank 128 on the Fortune 500. Still, both products are derivative of the iPhone; Cook’s signature 0 to 1 product, the Apple Vision Pro, is more of a 0.5.
Cook’s more momentous contribution to Apple’s top line was the elevation of Services. The Google search deal actually originated in 2002 with an agreement to make Google the default search service for Safari on the Mac, and was extended to the iPhone in 2007; Google’s motivation was [to ensure that Apple never competed for their core business](https://stratechery.com/2024/friendly-google-and-enemy-remedies/), and Cook was happy to take an ever increasing amount of pure profit.
The App Store also predated Cook; Steve Jobs said [during the App Store’s introduction](https://www.youtube.com/watch?v=xo9cKe_Fch8) that “we keep 30 \[percent\] to pay for running the App Store”, and called it “the best deal going to distribute applications to mobile platforms”. It’s important to note that, in 2008, this was true! The App Store really was a great deal.
Three years later, in a July 28, 2011 email — less than a month before Cook officially became CEO — Phil Schiller [wondered if Apple should lower its take](https://www.theverge.com/2021/5/3/22417725/apple-vs-epic-full-trial-slideshows-opening-arguments) once they were making $1 billion a year in profit from the App Store. John Gruber, [writing on Daring Fireball in 2021](https://daringfireball.net/2021/06/app_store_the_schiller_cut), wondered what might have been had Cook followed Schiller’s advice:
> In my imagination, a world where Apple had used Phil Schiller’s memo above as a game plan for the App Store over the last decade is a better place for everyone today: developers for sure, but also users, and, yes, Apple itself. I’ve often said that Apple’s priorities are consistent: Apple’s own needs first, users’ second, developers’ third. Apple, for obvious reasons, does not like to talk about the Apple-first part of those priorities, but Cook made explicit during his testimony during the Epic trial that when user and developer needs conflict, Apple sides with users. (Hence App Tracking Transparency, for example.)
>
> These priorities are as they should be. I’m not complaining about their order. But putting developer needs third doesn’t mean they should be neglected or overlooked. A large base of developers who are experts on developing and designing for Apple’s proprietary platforms is an incredible asset. Making those developers happy — happy enough to keep them wanting to work and focus on Apple’s platforms — is good for Apple itself.
I want to agree with Gruber — I was criticizing Apple’s App Store policies [within weeks of starting Stratechery](https://stratechery.com/2013/why-doesnt-apple-enable-sustainable-businesses-on-the-app-store/), years before it became a major issue — but from a shareholder perspective, i.e. Cook’s ultimate bosses, it’s hard to argue with Apple’s uncompromising approach. Last year Apple Services generated 26% of Apple’s revenue and 41% of the company’s profit; more importantly, Services continues to grow year-over-year, even as iPhone growth has slowed from the go-go years.
### China and AI
Another way to frame the Services question is to say that Gruber is concerned about the long-term importance of something that is somewhat ineffable — developer willingness and desire to support Apple’s platforms — which is, at least in Gruber’s mind, essential for Apple’s long-term health. Cook, in this critique, prioritized Apple’s financial results and shareholder returns over what was best for Apple in the long run.
This isn’t the only part of Apple’s business where this critique has validity. Cook’s greatest triumph was, as I noted above, completely overhauling and subsequently scaling Apple’s operations, which first and foremost meant developing a heavy dependence on China. This dependence was not inevitable: Patrick McGee explained in [Apple In China](https://appleinchina.com/), which I consider one of the all-time great books about the tech industry, how Apple made China into the manufacturing behemoth it became. McGee added [in a Stratechery Interview](https://stratechery.com/2025/an-interview-with-apple-in-china-author-patrick-mcgee/):
> Let me just refer back to something that [you wrote I think a few months ago](https://stratechery.com/2025/american-disruption/) when you called the last 20, 25 years, like the golden age for companies like Apple and Silicon Valley focused on software and Chinese taking care of the hardware manufacturing. That is a perfect partnership, and if we were living in a simulation and it ended tomorrow, you’d give props for Apple to taking advantage of the situation better than anybody else.
>
> The problem is we’re probably not living in the simulation and things go on, and I’ve got this rather disquieting conclusion where, look, Apple’s still really good probably, they’re not as good as they once were under Jony Ive, but they’re still good at industrial design and product design, but they don’t do any operations in our own country. That’s all dependent on China. [You’ve called this in fact the biggest violation of the Tim Cook doctrine](https://stratechery.com/2025/apple-and-the-ghosts-of-companies-past/) to own and control your destiny, but the Chinese aren’t just doing the operations anymore, they also have industrial design, product design, manufacturing design.
It really is ironic: Tim Cook built what is arguably Apple’s most important technology — its ability to build the world’s best personal computer products at astronomical scale — and did so in a way that leaves Apple more vulnerable than anyone to the deteriorating relationship between the United States and China. China was certainly good for the bottom line, but was it good for Apple’s long-run sustainability?
This same critique — of favoring a financially optimal strategy over long-term sustainability — may also one day be levied on the biggest question Cook leaves his successor: what impact will AI have on Apple? Apple has, to date, avoided spending hundreds of billions of dollars on the AI buildout, and there is one potential future where the company profits from AI by selling the devices everyone uses to access commoditized models; there is another future where AI becomes the means by which [Apple’s 50 Years of Integration](https://stratechery.com/2026/apples-50-years-of-integration/) is finally disrupted by companies that actually invested in the technology of the future.
### Cook’s Timing
If Tim Cook’s timing was fortunate in terms of when in Apple’s lifecycle he took the reins, then I would call his timing in terms of when in Apple’s lifecycle he is stepping down as being prudent, both for his legacy and for Apple’s future.
Apple is, in terms of its traditional business model, in a better place than it has ever been. The iPhone line is fantastic, and selling at a record pace; the Mac, meanwhile, is poised to massively expand its market share as Apple Silicon — another Jobs initiative, appropriately invested in and nurtured by Cook — makes the Mac the computer of choice for both the high end (thanks to Apple Silicon’s performance and unified memory architecture) and the low end (the iPhone chip-based MacBook Neo significantly expands Apple’s addressable market). Meanwhile, the Services business continues to grow. Cook is stepping down after Apple’s best-ever quarter, a milestone that very much captures his tenure, for better and for worse.
At the same time, the AI question looms — and it suggests that [Something Is Rotten in the State of Cupertino](https://daringfireball.net/2025/03/something_is_rotten_in_the_state_of_cupertino). The new Siri still hasn’t launched, and when it does, it will be with Google’s technology at the core. That was, [as I wrote in an Update](https://stratechery.com/2025/apple-earnings-siri-white-labels-gemini-short-term-gains-and-long-term-risk/), a momentous decision for Apple’s future:
> Apple’s plans are a bit like the alcoholic who admits that they have a drinking problem, but promises to limit their intake to social occasions. Namely, how exactly does Apple plan on replacing Gemini with its own models when (1) Google has more talent, (2) Google spends far more on infrastructure, and (3) Gemini will be continually increasing from the current level, where it is far ahead of Apple’s efforts? Moreover, there is now a new factor working against Apple: if this white-labeling effort works, then the bar for “good enough” will be much higher than it is currently. Will Apple, after all of the trouble they are going through to fix Siri, actually be willing to tear out a model that works so that they can once again roll their own solution, particularly when that solution hasn’t faced the market pressure of actually working, while Gemini has?
>
> In short, I think Apple has made a good decision here for short term reasons, but I don’t think it’s a short-term decision: I strongly suspect that Apple, whether it has admitted it to itself or not, has just committed itself to depending on 3rd-parties for AI for the long run.
As I noted above and in that Update, this decision may work out; if it doesn’t, however, the sting will be felt long after Cook is gone. To that end, I certainly hope that John Ternus, the new CEO, was heavily involved in the decision; truthfully, he should have made it.
To that end, it’s right that Cook is stepping down now. Jobs might have been responsible for taking Apple from 0 to 1, but it was Cook that took Apple from 1 to $436 billion in revenue and $118 billion in profit last year. It’s a testament to his capabilities and execution that Apple didn’t suffer any sort of post-founder hangover; only time will tell if, along the way, Cook created the conditions for a crash out, by virtue of he himself forgetting The Cook Doctrine and what makes Apple Apple.
*I wrote a follow-up to this Article in [this Daily Update](https://stratechery.com/2026/john-ternus-and-apples-hardware-defined-future-spacexai-and-cursor/)*.
---[^1]: As an aside, it’s notable that Alphabet has another business — Waymo — where the company has so far rejected an asset-light model of licensing their software to OEMs, and has instead to date pursued a much more capital intensive approach of owning and operating their own cars; that’s a choice that has always felt at odds with Google Services, but is perhaps more compelling and aligned with Google Cloud and the Google Capital Company.
Direct answer: Yes — the supplied Stratechery page identifies an article title and gives a one-line takeaway relevant to the AI-model-router topic.
- The page is a weekly overview of the Stratechery bundle , and it lists the article “Stripe Acquiring OpenRouter, Aggregating AI?, Flipping the Business Model”.
- Its listed description frames Stripe's reported acquisition of OpenRouter as “an implicit bet on a future market of models and the chance at Aggregation” .
Caveats/gaps:
- The supplied bundle contains no mention of Packy McCormick, so it cannot confirm that this is the Stratechery resource he referenced .
- Only the roundup listing and one-sentence publisher description are present; the full linked Stratechery article is not included, so the takeaway is limited to that summary .