We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
Top pick: books from Periodic Labs' co-founders
Liam Fedus and Dogus Cubuk co-found and co-lead Periodic Labs, a startup building "synthesis superintelligence." At the end of their Generalist episode, each recommended one book :
- The Beginning of Infinity by David Deutsch, recommended by Fedus. Asked what book he'd give everyone on Earth, he admitted it's "probably a popular one in Silicon Valley." He said it has "a spirit of optimism" and gets at "the unboundedness of science" .
- Subtle Is the Lord by Abraham Pais, recommended by Cubuk. It's a biography of Einstein by a friend of his who was also a physicist. Cubuk said most biographers aren't experts in their subject's field, but this one covers "very precise technical detail about what he actually did." He called it both "an incredible friendship book" and "an incredible physics book" .
Why it's the top pick: both founders gave specific reasons for their choices. Cubuk's point about the biographer's technical expertise also connects to how Periodic Labs thinks about science: Fedus says AI has mostly been trained on "the final artifacts of science" rather than on how the science actually unfolded .
Bill Gates's three books
At the end of his Ezra Klein Show interview, Gates named three books. Gates shared the episode himself, calling AI either "the greatest tool for opportunity and equality we've ever seen—or the cause of immense, unprecedented harm" . The titles and authors below come from the transcript, and some names may be garbled:
- The Correspondent by Virginia Evans. He "literally binged" this novel over the weekend and called it "very touching" and "very upbeat" .
- Into the Wood Chipper by "Nicholas Enrich." Gates described it as being about how USAID got put "into the wood chipper" .
- The Infinity Machine by "Sebastian Malib." He called it an AI book and a history of Demis Hassabis (transcribed as "Dean Deusibus") and DeepMind. He said it gives "a sense of the whole founding of the AI industry" and does "an incredible job" .
Essays and articles
- "Personal AI Should Actually Be Personal" by Ryan Sarver, recommended by Garry Tan. Sarver starts from Amazon blocking Meta's Muse agent from its retail site. He argues that an agent hosted by a platform "has two principals," and that one option should be one you own: open source, with your data in storage you control . Tan's takeaway: personal AI is "in its Homebrew computer club era," and truly personal AI means you "own your own skills and memory" .
- "Free the models: harness design at the frontier", from Replit's AI team, recommended by Amjad Masad. The authors are Daniel Furman, Jacky Zhao, Vaibhav Kumar, Ed Sioufi and Michele Catasta . Masad calls it a guide to building a harness that reaches "frontier performance at a fraction of the cost" . The post argues that the main model should pick its subagents' tier and effort itself, rather than a router making that choice. It frames this as an instance of Sutton's bitter lesson: "the less the harness should decide for them" . Masad runs Replit, so this is promotion of his own team's work.
- Richard Hanania on which experts to trust, recommended by Bill Gurley by way of @dawallach. Gurley says "everyone watching" the AI-doomer debate "should read about Tetlock's work." His point is that generalists ("foxes") predict better than specialists ("hedgehogs"), so researchers' "we know more" argument "works against you" .
Also noted
- Patrick O'Shaughnessy's interview with Anthropic CFO Krishna Rao (YouTube). O'Shaughnessy called it "worth revisiting today" . Topics include how Rao allocates compute across Trainium, TPUs and GPUs, what investors misunderstand about model companies, and why returns to frontier intelligence keep rising . O'Shaughnessy is the host, so this is self-promotion.
In response to the show's request for three book recommendations, Bill Gates named:
- The Correspondent by Virginia Evans: Gates said he had binged it that weekend and described the fiction as “very touching” and “very upbeat,” and apropos to what they had just discussed.
- Into the Wood Chipper by Nicholas Enrich: Gates described it as about someone who skips a party and decides to put USAID “into the wood chipper.”
- The Infinity Machine by Sebastian Malib (author name as transcribed): Gates called it an AI book and praised its account of the AI industry's founding.
Tony Fadell recommends listening to his audiobook, Build: An Unorthodox Guide to Making Things Worth Making, narrated by Roger Wayne, instead of calling him for career and startup advice; he says it contains much of the advice he gives daily to new grads, CEOs, executives, and interns .
Bill Ackman pointed readers to Warren Buffett’s work (“go back and read Warren Buffett”) as a case study: Buffett did not foresee internet-driven disruption, such as Wikipedia disrupting World Book, underscoring how difficult it is to assess disruption risk in the AI era.
Balaji pointed readers to Joe Gebbia’s talk accompanying the America.gov launch; the linked post announces the site and includes a video. He praised the talks and site launch as having the feel of an Apple event and described the effort as rebooting government like a tech startup, highlighting tangible user-facing features as a key takeaway.
Bill Gurley recommends reading about Tetlock’s forecasting work in the AI-doomer debate, arguing that generalist “foxes” are better at predicting broad implications than specialist “hedgehogs”; the shared excerpt says domain expertise alone does not ensure good forecasting habits . His post points to David Wallach’s post, which links to a Richard Hanania article: article.
Garry Tan endorsed and amplified @ruima’s X post arguing against PAUSD restrictions on math acceleration and ceilings on students who want to advance. Tan called Palo Alto “ground zero” for what he described as restrictions on honors classes and personalized education, and urged parents to support merit and advanced placement. Read @ruima’s post.
Patrick O’Shaughnessy said his interview with Anthropic CFO Krishna Rao was “Worth revisiting today,” linking the YouTube video. Rao’s first podcast appearance covers compute allocation, what investors misunderstand about model companies, returns to frontier intelligence, platform-versus-application strategy, and how Anthropic uses Claude internally. O’Shaughnessy said the conversation offers a rare view inside a consequential company at a pivotal moment and highlighted Rao’s unusual answer to his recurring “kindest thing” closing question.
Bill Gates shared a link to a lengthy conversation with Ezra Klein about AI, saying the world is not awake to its dangers and that AI could bring unprecedented opportunity and equality or immense harm, depending on humanity’s choices: https://b-gat.es/4AFZZqC.
Patrick Collison points to Bay Atlas (https://bayatlas.vercel.app) as an example of AI enabling richer, higher-density online maps; he says current services elide detail and that he wants to “wallow in cartographic filigree.”
Garry Tan endorsed Ryan Sarver (@rsarver)’s article, “Personal AI Should Actually Be Personal”: Tan says personal AI is still in its “Homebrew computer club era” and that users need to own their own skills and memory. Sarver argues that personal AI should be open and user-owned, with information under the user’s control and the user as its sole principal.
Patrick O’Shaughnessy recommended his video conversation with Instinct founder Noah Shinn, describing it as his first long conversation about the company and closing with “Enjoy!” The conversation covers personal AI agents, trust and privacy, Instinct’s business model, and growth. Watch the conversation. O’Shaughnessy shared it alongside his takeaway that using personal agents was making him see them as “the real everything store.”
Mike praised an unnamed write-up on agent-native architectures, saying he used it to create a skill and valued its encapsulation of the principle that anything a human can do, an agent should also be able to do.
Tobi (@tobi) amplified @redaction’s post featuring a YouTube video from @chasmmm about gamers rewriting entire games in Rust and combining them; Tobi reacted, “We haven’t seen anything yet.”
- Periodic Labs co-founder Liam Fedus recommended The Beginning of Infinity by David Deutsch to everyone, saying he enjoyed its optimistic spirit and its idea of the unboundedness of science.
- Periodic Labs co-founder Dogus Cubuk recommended Subtle Is the Lord by Abraham Pais, an Einstein biography. He praised its precise technical account of Einstein’s work and its perspective as a book written by a friend; he said it would appeal to readers interested in either friendship or physics.
Elon Musk posted “Also sprach Zarathustra” with a link to James Douma’s post . Douma praised what he called an “AI homily” for its substantive lyrics, nuance for critics, and story/message, and linked onward to an @_brightmirror post . Musk gave no further explanation in his post .
Amjad Masad endorsed “Free the models: harness design at the frontier,” shared via @pirroh’s post and written by Daniel Furman, Jacky Zhao, Vaibhav Kumar, Ed Sioufi, and Michele Catasta; the full post is here. Masad framed it as guidance for building a harness that reaches frontier performance at a fraction of the cost; the article argues that a composable harness should let models choose how to execute work rather than impose a rigid approach.
Guillermo Rauch praised Bay Atlas’s cartographic detail, saying, “You can just ship things (and wallow in cartographic filigree)” while linking to the showcase post . The linked post presents Bay Atlas as an example of AI making richer, higher-density online maps possible .
Morgan Housel shared The Psychology of Money Podcast content about why home prices are so high, connecting the topic to Yogi Berra’s joke that “no one goes there anymore, it’s too crowded.”
Jason Lemkin promoted a new SaaStr article, “What 17,000 Subscription Apps Tell Us About Free Trial Length”. He highlighted the article’s finding that 30-day trials on annual plans convert 86% better, contrasting this with B2B’s common 14-day trials, which he says persist because Salesforce and HubSpot used that length 15 years ago.
Jason Lemkin recommends a free guide to AI Sales Agents at saastr.ai; he gives no further reason for the recommendation.
Free the models: harness design at the frontier
Free the models: harness design at the frontier

Model routers are everywhere right now, but they have a fundamental limitation. No matter if based on advanced heuristics or a small model that reads each turn and picks which LLM to use, a router will always be less capable than the model it’s choosing for. Replit Agent lets the model decide instead.
The main agent, or core loop, chooses its subagents’ tier and effort, and adjusts its own as the task unfolds. Given that freedom, GPT-6 Astra hands routine implementation to less costly subagents and decides for itself where its tokens are worth spending. On both DeepSWE and Terminal-Bench, Replit Agent is Pareto-efficient against Astra on its own: no published Astra baseline costs less and scores higher. It also beats a sidekick architecture, the same setup with one long-lived worker, by 11 and 16 points.

For the full post with animations and footnotes, check https://replit.com/blog/free-the-models (opens in new tab)
Why we scaffold less
Every model release invalidates assumptions baked into the harness.
As models become stronger at long-horizon tasks, they don’t need as much scaffolding at the harness layer. In practice, we’ve observed them lean more towards delegation on their own: using subagents for context management and parallelism. Recent breakthroughs, Navier–Stokes among them, came in part from coordinating swarms of agents powered by frontier models [1] (opens in new tab).
But the frontier is jagged. The strongest coding model is not necessarily the strongest at designing UIs or making slides, nor the best at writing emails. So we design our harness to let each model work its own way, with the guardrails it still needs and quality at minimum cost as the goal.
Each new model sends us back to re-test what we held firmly, and to experiment fast with techniques that build on emergent behaviors. Freeing the model, then, means letting it decide how hard to think, when to hand work off, and who to hand it to.

Composable primitives for delegation
When we started experimenting with GPT-6 Astra [2] (opens in new tab), we found that the model delegates well. The GPT-6 family is also the first from OpenAI to support effort changes mid-turn without breaking the cache.
To use these capabilities, we refined four harness primitives. They give the core loop a small set of choices at each step: what kind of subagent to dispatch, at what size and effort, whether to return to one it has already briefed, and how hard to think:
Domain-aware subagents. Alongside a general worker, the harness offers specialists: read-only explorers, browser testers, reviewers, and a design subagent for slides and UI, each with its own model and tooling. For now, the harness still decides which specialists exist; the core loop decides when and how to use them.
Subagent tiers and effort. Small, standard, and large, each a step up in cost and capability, and an effort level within the tier. Both apply to every subagent, and the core loop picks them at each dispatch. For example, a mechanical rename goes to small at low effort, while generating hypotheses for a stubborn bug goes to large at high effort.
Reusable subagents. The core loop can return to a subagent it has already briefed instead of starting over. There is no single sidekick kept alive for the session: any number of subagents stay warm across kinds and tiers, and it picks which to wake. A longer cache lifetime on OpenAI’s newer models keeps the cost of doing so down.
Dynamic effort tuning. Now that changing effort mid-turn preserves the cache on some models, we trained an escalation system that checks the trajectory at each step and matches effort to task difficulty. Unlike a router, it acts mid-turn on the work in progress, not once on the request.
The code quality of Astra and Fable 5.1 [3] (opens in new tab) also let us use our code-review subagent less, with no drop in our eval scores. We’ve not seen this level of engineering quality from any model before.
Newer models delegate on their own
Frontier models like Astra and Fable cost more per token, which makes them look uneconomical next to smaller ones. We’ve observed them naturally delegate to less costly subagents, keeping their own tokens for the decisions that need them.
Replit Agent never forces the core loop to spawn subagents. Table 1 shows how three models handle that decision in production:

All three models delegate, but each in its own way. At medium effort, Fable models rarely hand work to a general worker: they send out read-only explorers and reviewers and keep the implementation for themselves. Astra is the first model we’ve seen routinely delegate to general workers without being told to, and once it has briefed one it tends to go back to it rather than start over. This return rate has risen with every model generation.

Results
We evaluated Replit Agent in Max mode, our highest-quality setting with Astra as the core loop, on two software engineering benchmarks: DeepSWE and Terminal-Bench. We compare against two baselines: Astra on its own in mini-swe-agent, as published on each leaderboard, and a sidekick architecture, the same configuration with one change: its subagent primitives replaced by a single long-lived worker. Each chart plots score against cost per task, so the most efficient configurations sit toward the top left.
On DeepSWE v1.1 [4] (opens in new tab), which tests long-horizon changes to active open-source repositories, Replit Agent scores 72% at $2.11 per task. Astra in mini-swe-agent at low effort scores 67% at $1.60, and at xhigh effort 74% at $4.43; the sidekick architecture scores 61% at $1.34. Terminal-Bench 4.0 [5] (opens in new tab) tests multi-step work done entirely from a shell. Replit Agent reaches 49% at $2.53 per task, against 42% at $2.25 for Astra at low effort and 60% at $5.86 at xhigh. The sidekick architecture manages 33% at $1.84.

Replit Agent beats the sidekick architecture on both benchmarks, by 11 and 16 points. The sidekick costs less, and gives up a sixth to a third of the score for it. Astra on its own scores higher only by spending more: its best settings sit 2 and 11 points above Replit Agent at more than twice the cost. Neither baseline wins on both cost and score. We ran Replit Agent exactly as it ships to users, with no changes to the prompting or harness.
The bitter lesson of harness design
We read these results as an instance of Sutton’s bitter lesson [6] (opens in new tab). Baking human knowledge into an agent helps in the short term, plateaus in the long run, and is eventually overtaken by general methods that scale with computation. A rigid harness forces the model into one way of working; a composable one lets it choose. The smarter models get, the less the harness should decide for them.
Compared with a more prescribed architecture, this approach buys us three things:
It bets on model scaling laws. Delegation that relies on the taste of the model improves with every release. Early previews of next-generation models continue the trend.
It fits the task. The model spawns nothing for a small task, one explorer for a search, and a team when a build breaks into independent pieces.
It reuses without persisting. A subagent keeps its context in case the model wants it back, and nothing persists unless it does.
In Sutton’s terms, the harness should let the model discover how to execute the work, not prescribe how we would have done it. Free the models.
Acknowledgements
Written by Daniel Furman, Jacky Zhao, Vaibhav Kumar, Ed Sioufi, and Michele Catasta. Thanks to James Austin, Toby Ho, Preeya Kirani, Zhen Li, Robin Newhouse, Devanshu Sen Pandey, Ibrahim Sheikh, Samuel Spitz, Peter Zhong, and the rest of the AI team at Replit for their contributions to this work. If you want to work on AI at Replit, my team is hiring; reach out to pirroh@repl.it.
References
Amjad Masad endorsed “Free the models: harness design at the frontier,” shared via @pirroh’s post and written by Daniel Furman, Jacky Zhao, Vaibhav Kumar, Ed Sioufi, and Michele Catasta; the full post is here. Masad framed it as guidance for building a harness that reaches frontier performance at a fraction of the cost; the article argues that a composable harness should let models choose how to execute work rather than impose a rigid approach.