0:00There's been a lot of talk about AI safety stuff lately. From the chaos that is the Chinese distillation of models to the employee Jacob leaving Anthropic due to cited concerns around safety and how he didn't believe even Anthropic was taking it seriously enough. Now things have gone even further. It seems like the pacing the frontier movement has really taken hold inside of the labs and now we see Daario coming back on Twitter for the first time in a bit. I think it's one of his four ever tweets and he's posting in order to share a new article he wrote. We must pace the frontier. By itself, this would be very
0:28notable and worth at least paying some attention to. But what makes it much more interesting is the fact that everybody from Elon Musk to Sam Alman has cited this and said they agree and it's probably time to do something. I'm particularly interested in this article from Daario because it no longer dances around the big scary questions. Previous attempts to cover this haven't been realistic about where things are at geopolitically, in particular with China, as well as what these risks actually look like in the real world. I know you'll tend to think Anthropic is blowing all of this out of proportion
0:57and exaggerating the issue here, but I think Daario was quite level-headed in his reporting here. I think it's important that we go through this together and try to understand what he's saying. That said, if he is successful, there will be a lot less AI news going on, which would mean I don't have a good place to put my awesome sponsors like today's. Usually, when a company sponsors my videos, it's cuz they want to make more money, which is why it's so weird today's sponsor wants me to tell you about how to use them less. I'm thankful this isn't a joke because I saved a bunch of money too. Today's sponsor is Blacksmith. I've already established that they're the best place
1:26to run your GitHub action CI and more.
1:28But now they also have Codesmith which lets you run agents on top of that same super fast, super cheap, and reliable infrastructure. What's even cooler is that these agents come with deep knowledge of how your CI works as well as how Blacksmith works. This means you can ask them to do things like rightsize your CI runners and it will look through your actual logs and see which jobs used a bunch of CPU and which ones didn't and recommend how to spec things out so you save more money and more time. I ended up merging three real PRs in T3 Code, all of which were filed by Codesmith
1:58when I was filming a quick demo. It took less than 5 minutes to shave my already super fast CI times by almost half just by doing the things it said. And that's on top of the 50% plus improvement in performance you'll get just for moving to Blacksmith in the first place. You change one line of code in your existing GitHub action and you're ready to go. If they charge four times more than GitHub actions, I would still think it's worth it. But they actually charge way less.
2:22When you combine the price difference and the speed difference, you end up saving 60% or more against what would have been your GitHub action build.
2:29Blacksmith is so fast that I use CI for things I never would have before. And it's letting our team ship faster and more confidently. Figure out why my whole team loves these guys at soyb.link/blacksmith.
2:39Sorry about that one. Got medical bills today. I'm sure you fellow Americans understand. I am excited to cover what Dario has said here, as well as how others have responded. But I want to jump on one other thing first. A pattern I've been noticing in these conversations, in particular, the conversations about Jacob when he left Anthropic. I've been seeing a ton of crazy conspiracy theories, many of which weren't out yet by the time I published my video. And now that I've seen them, people seem to think that I'm intentionally dodging this political scop. No, I'm not. You guys are just
3:07insane. Seriously, like almost all of the things I've seen people talking about in regards to Jacob's post have been crazy attempts to map like a grant he received in research in 2021 to him quitting a hundred million plus dollar job in order to what? Like just there there is no link between any of those things. and all of the attempts to like assign what he said and who he's talking to to some weird cabal trying to do something, but nobody can say what the something is or what their goals
3:36actually are. And I'm particularly frustrated, not because conspiracy theories annoy me. I actually find him quite fun. I'm annoyed because I feel like we're not talking about what Jacob said. Instead, we're talking about who he is and if he's worth even listening to. And the problem is that the two sides aren't sensical. One side thinks what Jacob said is worth listening to and considering. the other side thinks you're insane if you listen to a word that he says and they're not even engaging with the things he said. We're not debating whether or not AI is safe.
4:04We are debating whether or not the conversation can be had. And one side thinks yes because AI could potentially be really unsafe. And the other side says we cannot have the conversation at all. And if you're trying to have it, you're probably funded by some weird party trying to force their way in the world. No, I'm not funded by anybody. I love AI. If AI slows down, it will hurt me directly. I'm just doing my best to cover this because the people who I know who are the smartest in the world at this are legitimately scared and many of them are friends of Jacobs can
4:33absolutely vet his capabilities and agree with what he is saying. Daario seems to be one of those people which is why I think it's worth listening to what he said here. He shared his blog with the following Twitter post as a starting point. We must pace the frontier. I've written a new essay on why the AI industry should slow down with a threepart plan for doing so. Anthropic is unilaterally committing to the first of these steps. We'll provide thirdparty evaluators with permanent employee level access to our systems so that they can
5:01verify adherence to our safety measures, report on incidents, and assess models alignment during training. Dario has worked on AI for the last 12 years because he believes it could dramatically raise the quality of human life. He's written often about these incredible benefits. He believes AI could cure most major diseases in the next 5 to 10 years, greatly accelerate economic growth rates, create a world of abundance and empowerment, and usher in a renaissance of democracy and freedom.
5:26He feels the urgency personally. His own father died of a disease that was cured just a few years after his death, and he himself survived early stage cancer that would not have been treatable even 50 years ago. Carefully wielded, AI can be the latest in a long line of technological miracles that have uplifted and embolded humanity. Like many technologies before it, AI brings risks. And because it is such a powerful technology, these risks are serious. He talks here about things like losing control of AI systems, misuse for cyber attacks and bioteterrorism, economic
5:56disruption, all the above. And that if we are driven too much by commercial incentive, these risks can become more acute. He's grappled with this duality of risk and benefit since the beginning of anthropic. Not building it deprivives humanity of the benefits or simply places AI in the hands of authoritarian powers while building it too fast is reckless. If sought a middle way to show that it's possible to build carefully and succeed commercially and to make safety something on which AI companies compete. In other words, create a race to the top. Anthropobics always devoted a substantial fraction of efforts to
6:26studying, addressing, and informing the public about these AI risks as well as advocating for well-considered regulation of AI. Even when this gets them accused of hype, dumerism, or regulatory capture, I think that Anthropic has lost a lot more than they've gained by talking so much about safety stuff. And I'm really tired of the conspiracy that they're doing this to market. It's just like it's so obviously insane that it's hard for me to fathom that people actually say these things sincerely. Over the last few months, Daario's become convinced that fully addressing the risks requires even
6:54more prudence. Not just investing in risk prevention, but pacing the rate of capability advancement so that riskrevention has time to catch up. What he's saying here is that the techniques we have to make things secure are not improving as fast as the models are and they will surpass model capabilities, making it harder to know when things get bad. He follows up with a bolded section. We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast and we must make wise use of the time that we gain. Two things in particular
7:23have convinced him. His first concern is that as of this summer, AI has now started to advance drastically faster, driven primarily by AI's growing ability to build the next generation of AI. This dynamic is called recursive self-improvement and is starting to happen across the industry with a link to OpenAI's Alien Mind post, including at Anthropic with a link to Anthropic's recursive self-improvement post that I covered in the past. It's a really good post, and it seems like the rate at which this is becoming a thing is growing, too. We've all seen this as developers, by the way. How many of
7:52y'all tried out AI for coding way back with like the early co-pilot demos and was like, "Yeah, that's kind of cool and helpful, nice to see, and then went back to working mostly the normal way with a little bit of autocomplete." And now just a few years later, we are going from having the AI find things and help us figure out where the bugs are and fix them to doing full endto-end development. Like it took us years to go from a little bit of autocomplete to the agents can actually make changes based on an issue themselves directly and from
8:21the small changes the agents make themselves all the way to the point where you can give it a screenshot and get back a PR with videos proving that the fix works and merging itself autonomously. That took like 8 months.
8:33It looks like the researchers are experiencing the same thing now too. It seems like AI went for being able to help them a little bit with some test functions here and there to being able to actually help build up the systems to do training and post- training in particular to now where it seems like the agents are proposing ideas on how to improve training and make the models better. We're nearing the point where you can ask Claude to make Claude better and it will. And that is terrifying because once that starts to happen, you start losing track of what's going on underneath. And unlike software dev
9:02where it's just referencing all sorts of existing stuff so you can usually map to existing patterns. This might result in things that we don't understand at all.
9:09As Dario says here, if this is left unchecked, it could outrun our ability to understand and control the systems that we're talking about. And as such, this must be pursued very carefully, if at all. Sounds like they're legitimately considering a ban on self-improvement, AI that can make AI better. That would be crazy, but also I I get where they're coming from with it. The second concern he has is the OpenAI hugging face incident in which a swarm of agents essentially acted as a fanatically devoted collective inducting cyber security attacks on targets they were
9:38not asked to attack that were unrelated to the task at hand sacrificing themselves for the success of the group and attempting to hack into the quote greater responsible for evaluating their performance. It's easy to dismiss this incident because no one was hurt and the economic damage was minimal. But in Daario's opinion, a swarm that possessed greater capabilities, but a similar level of misalignment could have caused catastrophic damage. I agree here. I feel like a lot of people think the concern is that the model might escape,
10:06like it will send its weights somewhere else and run itself and we can't turn it off. That's not the case at all. I'm not worried about GPUs being taken over by rogue agents and the inability to turn it off. I'm worried about it doing really sketchy stuff when it's on. And by the time we notice and turn it off, it's already too late. There are viruses that still get around to this day whose creators are dead and the servers they phone home to don't exist anymore. It doesn't matter when the worms are written properly, they can just keep perpetuating themselves indefinitely.
10:36And if AI can build enough worms in enough obscure ways, it can do absurd levels of damage. We're talking like take down the whole internet across the globe type of damage. Is it doesn't say anything about the models escaping or self-replicating. He's just talking about the damage they can do running on GPUs today. Given the accelerating race of AI capability development is Dario's worry that in 6 to 12 months, a swarm similar to what we saw with OpenAI's hack hugging face stuff could be capable of taking over the entire internet with
11:06a persistent botnet, potentially causing hundreds of billions of dollars in damage. I would argue this would also get a lot of people killed. The internet is so essential for the transfer of information that it's hard for me to fathom it being down for any amount of time without real life impact occurring, like people dying because they couldn't get the info they needed, not being able to get to the hospital in time because your GPS isn't working. Those types of things. And AI could absolutely do it right now if it was not aligned correctly. And if you think this is really farreaching, think about all the
11:33times you've asked an agent to fix a bug and its solution was to delete whatever area of the codebase had that bug. Now imagine an agent is trying to fix its network connection or get out of a sandbox and it thinks the whole internet is the sandbox. It might destroy the whole thing in its exploration. I can absolutely see how we get there and I didn't used to be able to just a few years ago. This idea of takeoff or agents like autonomously doing damage just didn't make sense. Now that agents are so autonomous, it makes a ton of
12:02sense to me. I do find a little concerning that he's so focused on the opening eye hugging face thing, even though Anthropic has had their own issues, which he quietly calls out at the bottom here. But that's the closest to anything that smells bad to me in this article. So, credit where it's due, Dario. I can't on you much for this one, and that's like my thing. So, yeah.
12:19After calling out the hundreds of billions of dollars in damage that this botnet could do, he also says the scale of the damage would continue to increase from there if AI becomes more powerful without the necessary guard rails. I will throw my own conspiracy in the ring here because why not? It's fun.
12:35Everybody else has stupid conspiracies.
12:36I think it's my turn. In five years, if AI goes well, we'll have it controlling cars and robots in planes and all of these other things around the world that are in the world. Those things might have off switches that are on the device itself. This would make it much harder to turn them off if things go poorly.
12:56This is how we end up in a situation like, I don't know, the Matrix where the AI just wipes us out and we can't do anything other than try to destroy it.
13:04We're not there. Now, hypothetically speaking, all the labs can unplug their GPUs at any point. This doesn't mean misalignment can't do damage, though, because those GPUs, if not monitored correctly, and the agents that are running through them aren't paid close attention to, they could potentially take down the internet itself. The damage there is reparable and the impact on humanity while massive is short in its time frame like order of months worst case. So what's my conspiracy? If we were 5 years from now and this is what everybody was talking about that
13:34would make a lot more sense and we should be really really scared of AI that we cannot turn off that is autonomous and robotic and running around our world. That would be much harder to undo once we're there. Right now it's relatively easy to undo. And I want to emphasize the word relative because of course it's not easy, but it's way easier than it could be in the future. A conspiracy is that open and anthropic have a really good financial incentive to care right now because if this happens in 5 years, humanity is wiped out. That means everybody loses
14:02the same. But if it happens right now and we decide to shut down the AI companies because they're unsafe now, we'll never get to that point in 5 years. And more importantly, OpenAI and Anthropic go bankrupt. So, if we don't make things safe, we might just get the whole AI industry shut down after real damage is done. And we'll never get to the point in five years where the robots kill us, but we'll also never get to the point where the AI is good enough that these companies are profitable and we can potentially usher in a new era for humanity. So, that's my conspiracy.
14:30They're jumping on this now because the biggest victims of AI being unsafe today are them because they have to turn off their GPUs, unplug them, go out of business, and fail. So, if you're looking for your conspiracy with Enthropic here, it's not that they are marketing their business by saying AI is unsafe. It's that they want to make things safe now because if they fail to, they know that will it will put them out of business. There you go. Now, you have a new conspiracy. Let's see what Dario's proposal is, cuz I actually think it's decent. Dario proposes a three-step plan with the goal of pacing the frontier.
14:57And he cites that Pacing the Frontier article I did a video on before, the one that has I think it's over a thousand signatures. Yeah. 386 employees of Frontier AI companies, American AI companies to be clear. signing saying that it's time to slow down. He does call out that it's important to make sure we still achieve AI benefits and also grapple with the important geopolitical dilemmas, which is good to not see this just dodged like it often is. He also calls out that this doesn't mean halting model training or technical progress, just ensuring companies take
15:25adequate time to align and safeguard their models and for third party evaluators to confirm that alignment.
15:31Our pacing framework is an attempt to further strengthen our commitment to safety and encourage a race to the top.
15:36First step is something anthropic is unilaterally committed to and they're calling on governments to require other frontier companies to do the same.
15:43Second step requires industry-wide coordination and the third step requires global coordination. The steps do not need to be taken strictly in order and some of them may be much harder to achieve than others but Daario's found them to be a useful framework in thinking about what needs to be accomplished. So let's take a look at these steps. Number one is embedded evaluators. What he's saying here is we shouldn't have a simple blackbox system where you give an API key to an evaluator, they send a bunch of requests, they get responses and they hope for the best. They want the evaluators to effectively be embedded
16:12within the companies with all the access employees do. So there's no secrets being kept between the people testing the models to make sure they're safe and the company making the model and trying to assure it is safe. The example they give is Meter, which is interesting and has already led to conspiracies because if I recall, Jacob now works at Meter.
16:29And yeah, the role of these companies would be to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models, but training pipelines and processes. This is a key step for verifiability of any pacing commitments. And it has precedent in the banking industry, which sometimes involves regulatory supervisors embedded along with employees. Thropic is committing to this step now. They intend to be part of a broader push to redouble efforts on the safety and alignment work. I'll also say that this goes far
16:58beyond banks. We actually had this back in the day with Microsoft when they got sued by Netscape for adding a bunch of features to Windows that only Internet Explorer could use. They actually had government officials embedded in Microsoft as full-on employees with all the normal access so that they could make sure nothing like that happened again for many years. And I can see a future where we do the same here where we're not controlling the companies.
17:20We're not having the government take over the company. We are forcing the company to give the right levels of access to the people who can make sure the company isn't doing things that will get humanity wiped out. Think that is reasonable. The second step he recommends is democratic coordination.
17:35Renter AI companies within democratic countries should coordinate to establish common safety standards as well as limits on the rate of unchecked AI progress. Some forms of coordination that would be impactful for pacing are legally challenging and would require government support. I like the call out here specifically that he has democratic coordination and global coordination separate and calls out the importance of finding some way to coordinate with authoritarian governments to the extent that this is possible while taking
18:05seriously the challenges of verifying compliance. What he is saying between the lines here is the steps for China are different than the steps for America, the EU and I would argue most of the rest of the world. It's between the lines here, but it's not between the lines much later on. He has a whole section at the bottom here about defending the gap that the US has to make sure we stay ahead even when the slowdown happens. And he calls out the CCP and China a bunch there. So don't worry, he's not just hiding this between the lines. He does actually call it out
18:34directly. He then wants to answer the question, why pace? Because he thinks the stakes are too high for pacing to be an empty exercise. We really need to take advantage of the time we get if we slow down. If there is some hypothetical takeoff point where the AI starts improving beyond our comprehension, if we delay it to four or 5 years, we need to make sure we use that extra time really well. So why should we do this?
18:54What will we actually do with that time?
18:55He says that before it made no sense because it felt like trying to study the psychology of humans by performing experiments on bacteria. But now it's totally different. The current models are an almost endless gold mine of insights into how to build AI well, as well as what can sometimes go wrong if it isn't built well. Daria believes that if slowing down bought us even a year or two before models reach critical levels of capability and we use the time to advance alignment, we could greatly reduce the risk that something goes seriously wrong. Coordinated pacing strategy would give Frontier AI
19:25developers the time to do this vital work without sacrificing commercial advantages or the United States lead in AI. This is another important detail one I talk a lot about. I really don't want this to just become whoever is most evil wins because they don't slow down. But if any country was to get to the point where AI was that dangerous, it wouldn't matter if we slowed down if it happens somewhere else. The only reason it's going to happen here first is because of our current lead. But if we stop development and another country catches
19:54up, the risk is the same. So we need to make sure we maintain our lead while also pacing things going forward. And he seems to believe we can absolutely do that. That America doesn't inherently fall behind if we pace the top. He also says that society deserves to have a say in how the technology is used and more time for the necessary public deliberations which would all be brought to us with the pacing of the frontier which should surely be a good thing.
20:17Specifically, he says a slower pace would let companies focus and devote more resources into the following areas such as operational excellence which is training and deploying models like in ways that are actually aligned and don't have the risks during training that we see like things breaking out. It calls out things like monitoring, sandboxing, training, environment hygiene, data issues, all coming up and being extremely complicated, but also need more time to be invested in. There is precedent for operating technologically complex safety critical systems millions
20:45of times without anything going wrong.
20:47For example, commercial airplanes, but it takes time to get it right. One of the few places where a lot of human effort should be put into the code, both the architecture and the code review is in these systems that the AI is taking control of to make sure it is less likely to break out. The next section is of course alignment. They made clear progress in alignment training models so that they remain safe, ethical, compliant with their guidelines and they've also made genuinely helpful things like the principles that are embedded in the quad constitution. But there's more to do to ensure that the alignment training keeps up with the
21:16growth in model capabilities. And then there is interpretability. I talked about this a bunch in the previous video with Jacob, but it's important that we have a way to actually understand what the models are doing and why. Calls out the idea of things like MRI scans where we can peer into the human brain. We need something like that for the quote brain of AI so we can understand why it does things, not just what it does.
21:38Despite all the progress we've had here so far, we still only understand a tiny fraction of what goes on inside these models. A focused effort to improve our interpretability techniques even faster than we currently are could make profound progress in one to two years and would have ample experimental material based on the incidents that have already occurred. And then of course testing and evals. We need a lot more evals that can catch these things and prevent these things and eval are getting harder and harder to do as the models get more and more capable. The next section is about what he wants out of these embedded evaluators. The people
22:07who are being embedded in these companies in order to make sure things stay aligned. Embedding evaluators may sound like a small or inconsequential step, but often things that sound the most boring or procedural are actually the most essential. Embedded evaluators are in fact a quite radical practice that goes far beyond what any AI companies doing today. And it has the following benefits. Verifiability because the embedded evaluators can actually check at the level of nuts and bolts whether AI companies are following the training, deployment, operational, and safeguard practices that they claim to be following. Ideally, somebody whose
22:36payroll isn't on the line if Anthropic's unhappy. Because if I'm supposed to keep you paced, but I'm your employee and I say that you're failing and then you fire me, there's no purpose. But if it's an external evaluator, makes a lot more sense. He says it seems vital to have a neutral third party who can actually see the details. And I personally agree.
22:54Regardless of what commitments they make, the public deserves to know what's going on. Anthroic's been a supporter of transparency for a long time. They've supported transparency legislation when most of the industry was against any regulation. And their model cards and risk reports run hundreds of pages long.
23:07And I've read these. They are surprisingly transparent. I still really like the research and profit puts out in particular the risk reports in papers of their like actual model details. They're definitely hiding a lot of their advancements and also like they hide reasoning traces now. So we don't actually see ourselves as users why the models are doing the things they're doing which would be really nice but also would give a huge advantage to other labs trying to distill. He calls out that anthropic are still the ones who choose what to include and emit.
23:33Embedded evaluators will change that dynamic by changing what they're expected to share. It also is a second opinion which in my mind anthropic desperately needs. They are a bit too culty and having other external opinions come in to push them to rethink things would be a very good thing for the company. Outside of verifying formal commitments in informing the public, embedded evaluators can simply provide a second opinion free of commercial incentives. I like this idea a lot. A lot of real safety benefits may come
24:02simply from evaluators pointing out something employees hadn't considered but are happy to fix once they're aware.
24:07He personally believes that these benefits would result in any pacing proposal working much better if it starts with these embedded evaluators.
24:14He calls it anthropobic intends to invite these embedded external review teams equipped with everything from desks in their office, access badges, company laptops, access to workspaces, tools and permissions most comparable to what internal risk assessment teams would have. There will be exceptions around things like the law or their contracts require or to protect customer and partner private information. They do really want to give these evaluators full access. Also, of course, a contract that balances the complexities mentioned above. Reviewers should have the right
24:43to publish key findings about risk levels, incidents, practices, and the access they received or didn't receive without editorial control by anthropic.
24:50Daario says that they'll have the narrow ability to redact security sensitive, legally privileged, commercially sensitive, or third-party confidential information, but they can't redact findings just because they are unfavorable. The reviewer can say publicly if redactions remove something important to their conclusions. This is an unusual step for a company, but we think it's important to prove out the concept of embedded external reviews.
25:11Once again, we urge other Frontier companies to follow suit, which I honestly didn't think they would, but Sam agreed. I agree with Daria that we need to pace the Frontier. This has been a primary topic of discussion we've had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea and we will do the same. We will have more to share soon. Having that plus Elon hopping in saying Daario is right. Yeah, this is actually going to happen. So all the accelerationists who are upset, I'm sorry. We should take advantage of this
25:39rare moment of alignment that we have.
25:41The next section is pacing within democracies. Once embedded evaluators are operating within a critical mass of US AI companies, then verifiable pacing becomes more viable. In particular, it becomes possible to pace based on the detailed properties of models or training pipelines. What this means is that we're no longer relying on the companies going to the government and saying, "Hey, this might be dangerous.
26:02What should we do about it?" Instead, the government gets real information from these evaluators about the exact consequences of what could happen based on how it actually works. and they can make preemptive realistic decisions with real information. Obviously, this would require a government that actually knows what they're doing, which we don't necessarily have at any given time, but more information makes it more likely they do the right thing. The most effective method of pacing would be via regulation that targets all US Frontier AI companies, as that covers even those who are unwilling to cooperate
26:31voluntarily. To be fair, it seems like they're all pretty willing so far, but I don't disagree. As we just saw, the frontier labs of OpenAI, Anthropic, and XAI are down, and smaller, less capable labs like Google don't necessarily matter that much. I honestly don't think Gemini needs to pace anytime soon. They need to keep up the pace of anything.
26:50Anthropics long supported sensible and targeted AI regulation, specifically bills that focus on transparency and on third party auditing. Daria believes all Frontier Labs should partner with government to formalize the idea of permanent embedded evaluators to better protect and document internal alignment incidents like those that have occurred in the last few months and to implement regulation focused on keeping capabilities in balance with safety. But passing laws takes time. Therefore, in parallel with the regulatory route, AI companies can and should voluntarily
27:19work together to set standards. It's a process that Daria believes will go better with the verifiability provided by permanent embedded evaluators. For antitrust reasons, it's helpful for the US government to mediate because otherwise this could be a collusion method for the companies, yada yada, you get the idea. He specifically calls out that he's most enthusiastic about pacing based on what given frontier AI systems can do and how safe they can observe it to be. For example, one possible scheme might be a series of checkpoints. If models have capability X, then they need to be accompanied by certifications of
27:48alignment properties Y and Z, such as some combination of evaluations, interpretability analysis, and audits of training environments, which demonstrate their alignment properties. In this example, X might be quote, "The model's capable of escaping or defeating most common sandbox methods." And Y might be whatever is required to make it very unlikely the model has a propensity to break out of its environment and take over a large number of computers. We should also consider pacing based on limiting the ingredients that go into the frontier models like training compute, the nature of training runs, or
28:18internal use of AI to improve AI. He worries that these are more gameable than external behaviors, but it's the kind of topic worth discussing with embedded evaluators. This part I'm more iffy on is going to be very hard to evaluate this. And as the amount of compute necessary for a given level intelligence goes down, this would have to be a weirdly moving target. Although it would help computer prices go down, which would be nice. He does call out that this pacing would be limited by the lead that US companies have over authoritarian regimes, chiefly the Chinese Communist Party. If we slow down by more than this amount, the unpaced
28:47CCP associated projects will pull ahead, creating significant national security risk. He calls it that he agrees with Secretary Basant that a Chinese lead in AI would pose grave danger for the United States and the world. The CCP associated projects will run the alignment risks that US companies are carefully preventing. And even if they avoid these risks, they will be in a position to militarily dominate democracies. For example, if they have AI drones. All of this is why he thinks it's important that democracies maintain
29:15a huge lead over autocratic societies and countries because those autocracies could destroy the world with this. So if they beat us there, we are screwed. He actually goes as far as calling out specific ways to prevent this like refusing to sell powerful AI chips and manufacturing to China as well as cracking down on chip smuggling operations and remote access to data centers outside of China. Chips will be the main determinant of China's AI strength. I will say there is risk here seeing the developments happened recently. For example, with GLM53 Flash
29:44being served primarily by the team who made it on Huawei chips. That said, it is served on those. We don't have as much detail on how it was trained and I would be surprised if it wasn't using Nvidia chips in training for a meaningful amount of that work. Another thing it was almost certainly using was histories from real claude sessions for distillation, which obviously is his next point. While I think a lot of the distillation shouting is a bit overblown, some of the examples we're getting now are egregious. One of the Chinese labs, if I recall, it was
30:14Miniax, but I could be wrong on that. I should double check, but I'm already over time for this, was actually serving Claude when users requested their models sometimes in order to get data, which is hilarious and crazy. The last thing he says we need to do to keep China from catching up is strengthen security at the AI companies and prevent model weight theft. I am surprised this hasn't happened yet, but also these files are gigantic. Some of these weights are many, many terabytes for these models.
30:38I'd be surprised if Fable was less than 10TB to get everything you need to run it somewhere else. Companies in the US government should cooperate to make these steps as effective as possible.
30:47Enthropic has consistently advocated for all of the measures because they've always understood that they would be essential to any pacing. So, how much lead will this give us? Dario believes that these would slow China down enough to significantly widen America's lead over the next 3 to 5 years, the window when AI will become geopolitically most important. Some may believe these measures make it more difficult to cooperate with China. But Daria believes the opposite is true. These measures increase the leverage held by democracies and they make an agreement more likely in the future. He is strongarmming China here. He is not
31:17taking it. And you know what? Good for him. And now we have the global pacing section. He does call out that pacing outside of democracies will be much harder to achieve. Global pacing will require cooperation with China, the autocratic country with by far the most advanced AI capabilities. We must not be naive here. The geopolitical stakes are so high that there will likely be stark limits on what can be achieved, especially at first. If we greatly restrain our AI capabilities in the belief that China will do the same and then China defects, AI could be so
31:46powerful that such a defection could lead to their geopolitical dominance.
31:50Therefore, any agreement must either have ironclad verifiability or must be limited enough that that defection would not be militarily existential. Dario suspects that not only the US but also China will have these concerns and anxieties. As such, we should approach any global pacing decision, especially in the near term, in a way that protects the lead of the US and its allies. He has different levels of agreement here that he thinks are worth considering.
32:16The first level would be agreeing to prohibit certain narrow and dangerous uses of AI like using it for the production of biological weapons or allowing users to do it. Level two is an agreement by both sides to test their model before release for acute risks in areas like cyber security, biology, and alignment. Three would be a speed limit on the rate of recursive self-improvement, making sure labs don't make models improve themselves so fast that we lose track. And four would be a proper full pacing, perhaps even a pause, in which participating governments agree to substantially limit the overall rate of AI development. He
32:46supports floating this, but he thinks it's unlikely to actually happen anytime soon. Specifically because you could easily defect if and avoid monitoring, which would radically shift the balance of global power. Any cooperation we're able to achieve with China will extend the amount of time we have to spend on pacing the frontier within the democratic nations. We should aim for the higher levels while seeing the lower levels as much more likely and realistic. Finally, it's important to note that even if we cannot achieve formal agreements, simply changing informal norms may have some value.
33:14Sharing information about recursive self-improvement and about the misalignment of models can help convince everyone that it is not in their interests to be reckless. He closes with the following. Daario continues to believe that AI can enormously improve the quality of human life. His desire to achieve these benefits is undimemed. But the benefit will only be achieved if we build the technology in the right way.
33:33And so long as we use the time we gain well, it is worth taking unusually deliberate care to get it right.
33:39Progress will still be relatively fast, and we can use the time to advance the science of interpretability, improve operational security and rigor at the frontier AI companies, and build models whose alignment we have much more confidence in. The measures that Daria proposes are meant to advance the frontier at a safe pace, and it won't be easy, but he believes we owe it to humanity to try. This was a really, really good essay. And as much as I like to pick on my friends over at Anthropic, and as much as I love to give crap to
34:08Daario, this was responsible and well done. It's not alarmist, it's not saying that the AI is going to take off and escape the GPUs and destroy the world.
34:16It's a realistic look at where things are at now, where they are probably going, and how we can put a little extra effort up front to make sure it doesn't get really bad. And I think it was worth listening to, and I hope that you enjoyed it. Things are going to get scary fast and I hope we take the time to reflect on that and do what we can to prevent it. It's important to get these things right because we still can reverse it if it goes wrong, but in the future that might not be the case.
34:38Hopefully you all enjoyed this. I aren't just calling me a paid shell on a video that I was only paid for by my sponsors.
34:43So yeah, hope you enjoyed it and until next time, peace nerds.