0:00I think we are like laser focused right now at the frontier. We're seeing all these early signs of recursive self-improvement. I think the other labs are seeing this as well. And so I think it's like underscoring the value of like today's conversation is with Logan Kilpatrick. He's a member of the technical staff at Google DeepMine. In this conversation, we talk about the AI industry, the model lab wars, what's going on at Google. Are they actually committed to the frontier? How are they investing capital internally? What are the specific products and the decision-m that's going on inside of Google
0:28DeepMind? How are the AI efforts at some of the other labs actually affecting their decision-making? And then how should you as an individual think about benchmarking, coding agents, different types of bots, and many other aspects of the AI industry that everyone's talking about. Logan is somebody who has worked at a number of different companies in the industry. He's very well verssed not only what Google's doing, but how the industry is developing. And I think you'll find this conversation very valuable. Here's my conversation with Logan Kilpatrick. All right, Logan.
0:53Everyone thinks that Google is behind an AI race. Uh what do you think? What is your response to the critics who believe that Google maybe is not near where they should be?
1:02Yeah, it's a good question and I think it is uh I think my reflection of the last two and a half years is like I think it's a fair criticism um because people expect a lot of Google. This is what I try to remind myself. It's like you know it's not people trying to be rude saying that Google is behind. It's like Google's an incredible company have such a storied legacy. Um people expect high things of us.
1:24I think [clears throat] the tension point for us is you look at this portfolio of stuff that we're doing. We actually talking off camera about this.
1:31Everything from like, you know, genomics work to weather and science to, you know, new frontier models with Gemini 4 to we just released a bunch of new audio models, um, etc., etc. Like I I have this firm conviction that like Google and Deep Mind, we have the world's best portfolio of stuff. um the tension is like the portfolio spikes in different ways and sort of this is a very natural thing. Um I don't think that's like an excuse to not be at the frontier and so I think I think we are like laser
2:00focused right now at the frontier. We're seeing all these early signs of recursive self-improvement. I think the other labs are seeing this as well. And so I think it's like underscoring the value of like being at the frontier and so hopefully we'll see that with with Gemini 4. But an immense amount of progress and I think you've seen this with the Gemini uh 3.5 3.6 6 3.7 3.8 lineup of like in literally like 3 to four week increments. Uh sometimes less, sometimes a little bit more. We're seeing like very reasonable progress.
2:28And again, this is like the early signs of this recursive self-improvement loop. Um so hopefully we'll see that sort of like translate over in the same way to Gemini 4. And it'll be our it's our largest most ambitious pre-training run so far. Um, so I think you'll you'll it'll it'll sort of get us back in in contention with some of the the frontier labs.
2:47Talk a little bit more about this like commitment to the frontier, right? I think that is one of the things that people have always wondered is um you can go after general intelligence, you can go after specialized workflows. You guys have a business to run. You got a lot of cash, but also you're making a lot of bets on, you know, this kind of being the future of the company. But I do think that when I speak to people at Google, maybe there's more of a commitment to the frontier than people realize kind of outside of the company.
3:12Yeah, it it's part of why I had this conversation with Corey and he was sort of like making this verbal commitment to the frontier. Corora is our our uh SVP of Deep Mind leads uh the organization now um and is incredible. I love working with him. And he he was sort of making this verbal commitment. It's obviously been everyone's perspective internally for a long time. So he was sort of doing it to answer a lot of these questions of like do we even care about it? And I think this stems from like folks asking the question about you know why haven't
3:41we landed the pro model? Why are we pushing on flash? Do we only care about flash models now because we're shipping flash? And the reality is like we were just seeing a lot of progress. There was a lot of juice to squeeze. The recipe that we had was like working extremely well with the flash models. And so there was this like really tight iteration loop. people were getting excited about the progress and they're like, "Hey, let's focus and make that happen." Um, but I do think like the one of the things that makes me most excited about Google and the reason that I'm here and
4:09the reason that I'm doing this work is because the front having a frontier model and the commitment to the frontier so deeply permeates across Google's business. You look at like every single thing that we're doing as Google, the way that you know our 13 plus billion user products are sort of like touching the world in different ways from workspace to everything else we're doing across search. Having a frontier model is core and fundamental and a direct accelerant of every single component of that business. Not to mention the things
4:38that are like more exploratory like what is having I mean we're seeing signs of this from open AI and others like what does having a frontier model mean for frontier drug discovery um and being able to cure cancer and like obviously there's a correlation between those things. It's it's like a not immediate correlation uh as you know having a bunch of products with billions of users but like there is going to be some correlation. I think that correlation is going to increase over time. So it is like the business is set up to like
5:08fundamentally depend on having a frontier model and so you go ask everyone I think this is to your comment like it is the most important thing there's nothing else we're not thinking about like hey let's go make small cheap models that are great for everyone else like we do that because there's use cases in Google that support it but the aim is to be at the frontier and being at the frontier will enable us to do all the other things as you go through that decision-m it is kind of an interesting thing most businesses they would look at like what's the problem we face? Okay, let's go build technology to solve that
5:37problem. When you're committed to the frontier though, there's some version of just like let's go build the smartest model possible and then we'll figure out how to apply it. And then once you get into the application of that, you know, kind of general intelligence, it becomes okay, do we do it in a general way where people all throughout the company can just, you know, ping it or or customers can do that or do we go build like specialized workflows and and kind of go through that path? maybe just walk me through like the decision tree like bring us in the room as you guys are thinking through some of this stuff like how do you guys decide where to put resources and like maybe what the
6:07sequence of events is at least you know aspirationally for you.
6:10Yeah, I'll make a general comment which is um you know Sam Alman has that famous interview where he's like being interviewed by people a long time ago and they're like uh so what do you what's the plan to make money off this thing and he's like I don't know we're going to make the smartest thing possible and then we're going to ask it how to make money or something like that. Um, and you know, a little bit of a tongue-in-cheek answer, but um, it I think that's actually like less of what we've been trying to do at Google from the sense of like we actually know where the models will create tons of value for our customers in the world. We have all these products, we have all these
6:39services, we have the fastest growing cloud business in the world, etc., etc.
6:42So, like it's very obvious what the commercial application is. You don't have to think too deeply about that. Um I think to answer the question specifically about like what are the set of trade-offs and how's the resource allocation being made I think that's where this like early sign of um of sort of like recursive self-improvement is coming from um and I think there's a lot of resources being focused on like how do we actually make the models better at coding and research and science and sort of the core work needed to make like
7:11further breakthroughs and accelerations of of model progress. Um, and I think folks are very interested in that because there's just like such a large economic opportunity and then having a great model will then enable us to do all the other things that we want to do.
7:24Um, so there's a huge I think um, everybody is very codep I think we were like a little it's not that we were late to the game cuz folks I think knew it was important but I think you um, it's like yeah in hindsight everything is much more clear. Like I think in hindsight now it's obvious like we should have you know from a order of magnitude of resource allocation probably put more into coding sooner and like that makes sense and we had a bunch of other stuff that we were doing which
7:54those things actually turned out quite well like you know a good example of this is um you know nano banana like a great incredible image model that sort of took the world by storm had this massive impact for our consumer products and a bunch of other parts of the business and like you know to make that model it took research and compute and time and sort of and had this huge huge impact and like was it in hindsight right to do that versus doing something on coding like I don't know that there's actually a clear right answer but like those are the types of trade-offs that
8:22are actually a lot easier to analyze in hindsight and it's harder to know in the moment whether you're making that right decision but we do that reflection and sort of introspection which is important if you look at some of the companies um I think open AAI has been very inquisitive in trying to go after some of these verticals uh obviously uh SpaceX AI uh recently went bought cursor and then they launched Grock bot and they've got a lot of the coding stuff.
8:45Um, it does feel like there's different strategies. There's a little bit of chess getting played. Some people say, "Hey, look, we're going to focus on certain things. We're going to build it from scratch and and that's kind of our ethos and DNA." Google in other areas outside of AI has done both. They've built some things, but they've also acquired things. How are you guys thinking about, you know, if you feel like you're behind maybe in coding or other areas, like will you guys go buy stuff or is the focus internally on like let's go build, we know how to do it, we've got the resources, it's just a focus and kind of strategy thing. Yeah, we we've definitely done both in this
9:14context. And so we had um I think the Gemini CLI launched early um maybe like almost like a year and a half or two years ago um and sort of had some early traction. I think a few million users using that product and then we also went and um did the sort of acu hire of the wind surf team uh which is now part of cognition uh whom we were talking about off camera but um yeah and so we have a bunch of those folks. they've sort of been building this anti-gravity product internally uh both for our internal engineers but also for external users
9:43and have seen actually I think like one of the biggest impacts has been the internal acceleration and so they're really really deeply focused on like how do we build a great product for for engineers and and actually even non-engineers now inside of Google make it work really well um and then that sort of will translate to a great product externally eventually and so it's been cool to see um us do both things I think is one of the interesting strategic advantages of Google is we get to take many shots on goal and so um
10:12yeah I think having a bunch of different coding products and it's also like I think what's most interesting on this thread is the ecosystem has evolved I think about this all the time like the coding product that you would go to market with today or the sort of like developer or whatever like knowled even more generally knowledge work product that you're going to market with today looks so different than what it was even 12 months ago and actually like Grockbot's a perfect example of this like it's not obvious to me that like the Grobox style product would have
10:41worked 12 months ago. I think it works now because the models are good enough, but like 12 months ago you actually like needed all of this like developer UI scaffolding of like you know give me all these extra features and buttons and things because like the model's not really that smart and I need to wield it and turn it and sort of critique it in these very very specific ways. Um, and so I think one of the my my observation of this and my point of this is like I think there's going to be many of these
11:08opportunities to take shots on goal as the model progress continues it like unlocks a new paradigm and it feels like you know Grockbot and uh and instinct and Muse and a bunch of these new products that are coming to market right now are like the evidence of like the models have crossed another chasm where the product experience you can now build is fundamentally different than the one you were building 12 months ago. And this this is true for developers, it's true in a bunch of other verticals as well.
11:34With Gemini 4, you guys have talked about this large pre-training run that you're doing. Um, what should we take away from that? Like are there specific things that you can share in terms of what that's going to look like?
11:45Yeah, I think the thing to take away from this is like the commitment to the frontier. Uh, doing large pre-training runs is extremely expensive. It is like a large order of magnitude of investment. It's like I don't I don't uh the numbers are are quite large. Um, you want to tell us?
12:01Yeah, I I don't I actually don't even know what the numbers are off the top of my head just, but I can I can do some of the math uh in my head and like it's a lot. Um, and I think it's important. I think the other point of this actually is like pre-training has been a significant strength for Deep Mind in the past. Um, and so I think this is like one of the areas where I think we have like some of the best talent in the world. Um we have like actually this is like the infrastructure scale of Google um is an advantage in this context like the TPU fleet is an advantage in this
12:30context. Um a bunch of the data infrastructure stuff we have is an advantage. And so I think the point of telling people about this pre-training running is to tell people that like it's not like we're rolling over and playing dead. Like we are pushing the frontier.
12:44Everybody's working as hard as humanly possible. We'll hopefully see a bunch of incredible results from this new pre-training run. Um, and actually interestingly like you do the model comparison today and like all of our, you know, I think this pre-training run will very specifically like get us to the category that we need to be at to be competitive with where our competitors are at. Um, and so I think the proof will be in the pudding when we hopefully launch this model and customers get their hands on it. Um, but I think that's the expectation and that's the hope right now. When you say
13:14competitors, I think most people will think about open AI, Anthropic, uh, you know, Grock or or SpaceX, etc. Um, do you guys worry at all about like the Chinese open source openweight uh, type model players? Do you worry about maybe other competitors that aren't one of those three companies?
13:31Yeah, I think what's so interesting right now is like the space feels incredibly dynamic. And so I think I mean I think all I I personally have never understood all these memes of like I don't think about the competitor. I'm like I think about our competitors because they're all incredible companies and like I think we would be wrong to not be thinking about what they're doing and you know examining are they are they making the right decisions? Are the things that we could be doing differently still sort of like knowing what our core focus is and so spend a lot of time looking at like what are the things that folks are doing. Um and
14:00obviously the Chinese model labs have done an incredible job so far. It's like there's there's clearly a bunch of like question marks as far as IP stuff, uh, model trading stuff, but like with the set of constraints they have, all things considered, they've seemingly done a pretty solid job and like there's clearly research innovation they're doing as well. It's not like they're just copying what everyone else is doing. There's like actual frontier research happening. Um, so don't want to discount them as a competitor. It's also clear that like startups and companies
14:29want to use models that they can host themselves. Like I think there's like a there's like a philosophical question of like oh how do you you know what's the what's the feeling about these labs in China open sourcing these models and doing the thing they're doing and then there's like a practical business question of like customers want these type of models. They want to like I talk to startups all the time and startups want to be able to take the weight of the models and customize them for the use cases that they care about. Um and so there's a huge market there. there's a huge opportunity and so I think the
14:57question is like will we see um like US open- source labs um and like Nvidia's you know spinning up these types of efforts uh we have some of this on on the smaller ondevice model side with Gemma um I think we'll see like reflection AI a bunch of other folks like take shots at like can you actually produce frontier openweight models um but I think the cool actually the cool thing for all of us is that how just how competitive it is. Like the fact that
15:26those labs are able to like stand in a similar regard at all to these large companies in the US um is actually I think a good thing for all of us right now. And so there's a there's a huge amount of competition that's pushing everyone to be better. Today's episode is brought to you by Token 2049. The largest conference in crypto is back.
15:45Token 2049 will host 25,000 people, 300 speakers, and a thousand plus side events in Singapore on October 7th and 8th at Marina Bay Sands. The speaker list is absolutely stacked. Shane Copeland from Poly Market, Jeff Yan from Hyperlquid, Adena Freriedman from NASDAQ, Arthur Hayes, Balagi, and Eric Trump. [music] Crypto and traditional finance in the same building, which tells you a lot about where this is going. And the conference runs right into F1 weekend, so the whole thing turns into one giant week. If you're headed to Singapore, use code POMP 10
16:14for 10% off your ticket. Token 2049, October 7th and 8th in Singapore. Go check them out in the link in the description.
16:22When you think about uh playing chess, you definitely got to understand what your opponent is doing. I I agree with you that, you know, kind of uh only focusing on your pieces does not help you win the game. Um with that said though, it does feel like there is a lot of question marks about how some of this stuff is getting done. You know, if you think of some of the math problems that have recently been solved. It was, hey, did the models train on, you know, other people's questions. Was there uh some peing at, you know, data that people thought was private? There's some questions now about was, you know, Kimmy
16:51actually passing some of their queries just to Claude to answer versus Kimmy doing it themselves. And it's very difficult, you know, at least for me, but I think many of people to understand like what is real and what is just like Twitter fodder or X fodder where people just say, you know, they like they like the drama. It's almost like the TMZ of the AI industry, right? like what what's the new thing that we could all you know grab hold of for the day? How do you personally think through, you know, where to spend your time in terms of like your attention? Because it it's happening so fast. There's so many
17:19different things, you know, even if you we just think over the last week or so, you've got everyone from Paul Tutor Jones putting out opeds. You've got, you know, Jensen talking about AGI, you've got Astra, you've got like all these components. Unless you figured out how to get more than 24 hours in a day, you know, you don't have as much time. So, like what is your process to do that?
17:39Yeah, I think this is actually an interesting point that you're making, which is and I think this has I think the po the the trend has changed over time in like the level of signal to noise. Um I think there's actually just a lot more noise these days. And so I do think it is like a it is a muscle that you have to build to sort of filter as much of this stuff as possible. And for me it definitely passively cons
18:13better models. Um there's a bunch of stuff to stay on top of and make sure that like we're reacting to the right things that are happening and being proactive where it's needed. But like I think it's really easy to to get caught up in all the crap that's happening in the world right now. And like um I think my my advice to people is like filter out as much as possible, be more intentional about how you spend your time because like there's a lot of noise and it's not always clear to me that like the noise is actually translating
18:40to any amount of signal. Um and so yeah, it's like there's yeah there's very specific cases where this is true. Like you know the hugging face open AI situation is like good example of like lots of noise. There's definitely signal there. There's something to be learned.
18:54There's something to understand. There's a lot of cases though where like this is not the case. Um and so yeah, trying to try and be intentional about the places where there's actual signal.
19:03If you almost take that same issue or challenge and flip it, the other side of that is like there's a lot of opportunity cost given that the cost or barrier to build things has come down so much you now have access to superhuman intelligence. You can vibe code things.
19:18You know, you can have the bots go and build companies or, you know, kind of run parts of your business. like that it does feel like not only are there more distractions but if you get distracted the opportunity cost is higher than ever and so how do you think about you know your role internally you've worked inside of open AI you've worked at Google like maybe like what are some of the things you've picked up and how you're navigating the productivity side of this as well yeah it's so true and I think actually the thing that I struggle most with now
19:47is like um it's this like level of ambition problem which is like I used to be able to be like, "Oh, I'll just go like do this thing. It's going to be small and concise and well scoped and and now it's like I actually like if I go do this, like this could be a billion dollar opportunity for us and like so I have to take it like quite seriously and like that like weighs on me and I'm like, you know, having to spend more time to be thoughtful about like is this the opportunity that like we really want to go after as a team
20:16because there's so much opportunity everywhere." Um, I think for me this goes back to like I'm I try to be extremely principled about like what are the things that Google is well positioned to compete in. Um, and there's a lot of things that we're not well positioned to compete in. There's definitely some that we are well positioned to compete in and like we have structural advantages with, you know, Google Workspace and Google Cloud and distribution and things like that.
20:38Um, and so that's sort of my filtering mechanism on the on the product side when we think about like what are the opportunities to go after like I don't want to go after everything. Um, I want to go after things in which there's like a natural uh lift because we have other assets inside of Google that will actually contribute to the success of these things. And so this is what we've done in AI Studio. we have all these deep integrations with Google Cloud and all this stuff that like no other product team in the world can actually do because they're not inside of Google
21:07building this product. Um, and so it means that we're um it means that we're competing in some of these categories, but we're running a playbook that only we can actually run. And so we'll see in the fullness of time of was that the right playbook? Does it actually make sense? Maybe we should have just been doing the things everybody else were doing. Um, but I think it's I'm I'm trying to keep that filtering mechanism um very top of mind as we're making the decisions. Uh, and there's just so many cool things that Google has that make this like fun. And so I'm I feel like
21:35I'm not limited by like um by this at the moment.
21:39One of the aspects of the AI industry that is just intellectually, you know, stimulating I think for you, me, many other people is there's a level of strategy that is being played out. So it's not just like can you get the hardware and the software to do certain things you know kind of create magic or or turn sand into intelligence like that that is obviously very difficult in plenty of uh challenges there but the strategy side you know we were talking previously that um most of the frontier models are pursuing general intelligence some form or fashion right but then
22:09there's a bunch of startups that are saying well what if I take a specialized workflow approach and um if you think of like what we've been building with Sylvia you know this idea of well if we go and we build a bunch of proprietary technology that from model routers to harnesses to you know data pipelines and uh our own models etc. It does feel like there's almost a point on each application of AI where you kind of have to decide do we go after general intelligence or do we go after the specialized workflows and Harvey Sylvia
22:38you there's many players I think that are seeing a lot of traction in specialized workflows but what I find fascinating about Google is you guys have multiple applications where you have to make that decision over and over again like do we go and build the specialized workflows or can we just use the general you know purpose model how are you navigating that like for each one of these use cases cases. It's almost like you guys may be making that decision more than anyone else in the world.
23:01Yeah. Actually, I I've got two points for you on this. Like one of the thought exercises that I am continually proposing to our team internally um is in 5 years do we expect and maybe 5 years is the wrong time horizon but like in 5 to 10 years do we expect Google to have 10,000 products or two or three products? Um and I think this gets to this like vertical workflows versus sort of like general intelligence. And so I think there's like clear signal in the
23:31market that customers want vertical applications. Um they don't like you know you think about like why apps are so successful and like that's sort of there's this um this user behavior pattern which is like hey I I think as a user in terms of like using a particular application or a tool in real life. I want to go swat a fly. I go get a fly swatter. I want to go drink water, I get a cup. I don't like go to this like Oracle all-encompassing tool that can
23:59morph to do like it's not it's not something that like we intrinsically have grown up and sort of evolved as humans to understand. Um and so I do think there's this like really deep rooted muscle memory and I think the tension point will be given that extremely deep rooted muscle memory. Um, does it uh is that enough of a uh a sticking point that will like keep these vertical applications alive in a world where alive and thriving in a world
24:28where like the general purpose thing can actually do the same stuff. And so I think that will be the most interesting and this is where I think this like um you know what's the intersection of like AI and new hardware consumer hardware devices I think is going to be really interesting like as people change the way that they work with software and with technology like you imagine you will want like new form factors because the form factors we have right now are sort of a little bit more of these like verticalized experiences. Um I do think
24:56there's a separate edge of this which is and this is true in all these vertical domains like the vertical domains are successful also because somebody is focused. Um, and I I like deeply believe this like you know startups are always worried about like oh is the big company going to come after me? And it's like you can always do a better job than the big company with like very few exceptions because you're focused and more deep on some vertical that your problem that your customers have that like no one else is going after. Um, and
25:25so I do think it's like an edge to have that verticalness. Um, which is really interesting. And I think the other point that I wanted to make and I'm curious actually what you think about this. I think there's all this conversation of like general intelligence and something that's been like uh that's been very top of mind is I think the labs have historically described general intelligence as if like they would build the general intelligence themselves. Um and that like seemingly the general intelligence would then be powered by
25:55like end to end almost the models that one of these labs creates. I think there's something really interesting about this future where like if we really had general intelligence like you wouldn't expect that the general intelligence that Google creates is only using Google product services and models. You'd expect like hey if that's generally intelligent enough to know that like this other thing that some other company created can do something that our thing can't do or can do it better. Humans are generally intelligent enough to figure that out. And so I
26:24think it actually it's going to add a lot of I think um on this path to general intelligence as like the model labs go and continue down this direction. I think it's going to add a lot of like confusion to even like understand what that really ends up becoming because I think it's going to look a lot Yeah. It's going to look a lot more chaotic I think in how this like these general intelligence systems play out than I think this like beautiful vision of something that can just like do anything and everything for you.
26:50It's interesting you talk about this. So before we talk about the general Antonio, let's talk about like a microcosm of this. And you know the problem that I've been thinking the most about for the last year and a half is Sylvia. But um for those that don't know the product, you basically come in, you attach your financial accounts, you put up your private investments and you start talking to Sylvia. But we have chosen to go specialized workflows and we have done a whole bunch of you know very innovative things I think in terms of the AI harness the uh memory and file system the model routers that you know uh fine-tuning etc. But if you take like
27:19the model router, one of the perceived advantages of not being a frontier lab is that you should be able to route queries to any of the model lab, you know, models, right? And so what are the odds that Google is going to route to open athropic? It's not zero, but it's not, you know, 90% either, right? And so same thing I think with each one of the labs is like what is the incentive for them to keep the queries within their family of models versus the ability to
27:48act more as like a third party and actually route across. I don't think we've really seen how everyone's going to play that. And so as a third party you're like well I don't really care right. I just want the best level of intelligence at the lowest cost that answers the question for you know our uh our user. And so I do think there's some of those things also where people are trying to figure out not just like, you know, if you then extrapolate this out to like general intelligence. The user doesn't care if it's a Google product or not. Maybe there's some like ethical or moral things that maybe they align more
28:17with, but for the most part, they just use a product cuz it's the best one. And is that going to actually be built by one company or is it going to be, you know, kind of a a bundling? I mean, you know, what's the saying is like the world is just bundling and unbundling over and over and over again.
28:33And so, you know, it is a very um difficult thing to predict because I think we're so early in this journey that like every week's always got a new model that seemed to leapfrog everybody else and then you're trying to predict how consumers are going to interface with this stuff. And maybe like the last example I'll give is in my own life, I have been using for the last couple of weeks Grockbot professionally and instinct personally.
29:00They actually do a lot of the same stuff. But to your point about like you get the cup for water and you know you get the fly swatter to to swat the fly like I just kind of have in my head you know okay instinct when I got a personal question and grock when I'm doing something professionally it's probably pretty dumb you know like from like if they do the same thing like why don't you just use the same product but it then goes to like okay well now you're starting to see four or five others come to market and as somebody who likes to be an early adopter like well I should try those but then what about the
29:30context what about the memory how do I import that over and and there's like this, you know, kind of like user journey we're all learning of is it worth the time to go try the new thing if I don't have some kind of shared memory or, you know, especially at network effects where like you're a wife and you have shared memory somewhere, then how do you interface with something? And so like it it's almost like more questions than answers right now. And I think that's probably why, you know, you and I and so many other people are so excited about this, right?
29:57Yeah. Yeah. No, I think you're right.
29:58And I love this like the bundling and unbundling analogy because I think there's another version of this is like what's old is new again uh or what's new is old whatever the expression is. And like actually you see this with what what's so interesting about Crockpot and Instinct is like um this like form factors like the what's old from a form factor is now like message like chat was one of the original ones. you had all these like meh chat b thoughts like in the 80s and 90s or whatever it was and then like that came back and then boom
30:26all of a sudden what's old is new again.
30:28Um and then the same thing is now true for messaging. It's like we all use all these messaging apps and then it's like now all of a sudden the hottest form factor for AI is like the messaging. And so I think it's an interesting exercise of like in actually in all of these domains. I think the reason people feel this way is because like they're accustomed to this experience. Getting some customer to like adapt to some futuristic new thing is like actually extremely difficult to do uh and takes a really long time. You want people to go to some form factor they're familiar
30:58with. Um, and so I think about this all the time. Um, for like how do you actually get consumers or users to go and adopt new technology is like you want to make it feel familiar. And I have this hypothesis that like messaging like in actual like chat apps like has it it seems so unlikely that that's not going to be the dominant form factor um a few years from now. Like it's surprising to me it hasn't been more dominant. And I think it's because actually the operating system like
31:25messaging app owners are like just at the cusp of this, but like you'd expect, you know, like Apple and Android and WhatsApp, etc. to like really lean into that form factor and they they already have where all the communication is happening and you throw some agents in there and like, you know, it it feels like it's a natural place.
31:45It it does feel um like there could be some platformers. Now, I think that the people developing these products are obviously partnering with and trying to prevent that, but you know, if you wake up and you've got a chatbot that's very popular and all of a sudden you're blocked on Apple's system, that would be a big problem, right? And so, you know, I I don't know if that's really in Apple's best interest to do that stuff, but I do think that there's some folks who are kind of thinking through that.
32:07The the other aspect though around um you know, the kind of the chat bots and the messaging, I do agree that it's an interface that we all are very comfortable with. you know, I use it on a daily basis uh with these bots, but I have not yet become a very big uh voice user and I have a lot of friends that are like voice pill, you know, they're walking around, they're like whispering in their in their microphones or whatever, sitting at their desk. Uh do you use voice or like what is like maybe the adoption if you had to predict it inside of like the AI team at Google in
32:36terms of people who are you know fat fingers on a keyboard versus using voice? That's a good question actually. And I I I eb and flow between this.
32:44There's like what actually what I found is like the best use case for me for voice is when I'm like doing some sort of demo in front of other people and that way I don't have to like fumble typing things and I I can't spell and all that stuff. And so [laughter] just voice like straight in is like way faster. It makes the point it's like much much more succinct. Um but there is something about uh I think it's a and I'm sure there's like good you know neuroscience research out there that explains this like I think it is like a people manifest thoughts in different
33:14ways and like the physical manifestation of thoughts either coming audibly or like through tactile typing or even writing like to me I feel like I have like different thoughts depending on the sort of expression form and so like the way that I I spend a lot of time talking and doing stuff at work and I spend less time writing sometimes and so it's like actually quite helpful for me to like pulls me into a different mode of thinking when I start to write just because of the form factor. Um and so I think we'll see actually more of that as well and that's I think the you know
33:43obviously people like audibly speaking and it's a helpful way to think through things as well. Um but I think we'll see the sort of buckets of these different types of thinking uh manifest from like how you interact with AI as well.
33:57Yeah, it it is interesting. I think the science shows um the single best way to remember something is to physically write it down like with your hand. Next would be typing, right? And third is just kind of hear it and don't do anything. Um but maybe there is something about not just the memory but also the ideiation right you know we definitely know uh from science that walking outside kind of the the act of moving you know forward um does a lot of uh ideiation showering right you know there's many uh kind of examples throughout history of people who just
34:26went in the shower and thought of things shower twice a day for this reason it's not for hygiene it's it's just for I mean not tongue and cheek but uh sometimes honestly cuz like you do just have head space it's Right. Yeah. And and look, part of it is like are we just so all terminally online that just like the shower is the only place that the phone doesn't go or is it like there is something about, you know, even 50 years ago before people had, you know, supercomputers in their pocket. The shower did lead to new ideas, right?
34:52I think the shower was just cold 50 years ago and so people were just being shocked and uh [laughter] and you know, having new ideas probably.
35:00I love it. Um let's talk about DeepMind more specifically. you know, the work there obviously um has uh has been very broad for a very long time. Um we mentioned a little bit about the genome uh kind of project and the work that's being done there. I'm pretty surprised at just how large it is, but it still doesn't get maybe the respect that it deserves. Can you talk a little bit about some of what's going on there?
35:23Yeah. No, 100%. I think and I'm not an expert on all the science stuff that we're doing, but it's incredible to see the progress. I think across um the way that I would frame this is deep mine is split up in a couple of different ways.
35:35There's sort of like foundational Gemini and there's like we want to make the best frontier model and a bunch of different sizes of that model and sort of all of the different modalities that work in mainline Gemini. Um and then there's a there's a whole science unit and inside the science unit there's everything from u the genome project a bunch of the alphafold stuff um there's a bunch of like science things related to like biology there's of a bunch of other science stuff related to like
36:02weather and mathematics um and that whole portfolio is like also at the frontier of doing all these really interesting problems that no one else is doing. Um and then has all these very unique collaborations actually with folks like Isomorphic Labs which is the um our sort of like drug discovery company inside of Google and DeepMind that Deis is the is the CEO of um and you know they do all these deep collaborations and so it's a really interesting way for them to like not
36:32only solve the problem and make progress on the problem from a foundational research perspective but then actually have like the applied side of it as well. And so I think it's this like unique flywheel that exists inside of Deep Mind itself where like we're creating a bunch of the frontier innovation. Um we're doing all this interesting science work and then it actually has an application. It's not like we're just doing it for the sake of doing it. And I think this was actually the lesson from um from Alphafold which was like hey we did all this really
37:00really interesting work. It was super interesting but we were actually doing it like to solve a scientific grand challenge less because like we had somewhere where it was an immediate commercial application of um but it's like it became very clear like hey there's all these commercial applications we're opening this up the scientists are all using it um and so I think the general philosophy is like do this frontier science work um across genome across weather etc and then actually have a place to apply it to inside of Google and then actually most interestingly take a bunch of the
37:30lessons and learning and data and other things and upstream those back into the mainline Gemini model because ultimately like the mainline Gemini model is going to become better at a bunch of those things than those individual domain specific models and so you need to make sure the flywheel also goes back to there and so that's why having it under like a single roof actually makes sense and we see the cross-pollination between these things we've seen historically like all of these interesting like alpha proof with mathematics um trickle back
37:58to directly increasing to reasoning ing capabilities of the model for mathematics in the mainline Gemini model. We've seen this for cyber now.
38:06It's it's not an alpha project, but it's a similar domain where like cyber capabilities as we push the frontier on cyber directly correlate to like models having better coding capabilities. Um, and so there's all these other examples where like this flywheel spins and I'll make one comment which is I think people talk about the flywheels stuff like this as if it's like a magical thing that just like works. Um there is an immense amount of effort and energy that is required to actually this the flywheel does not like you don't spin it and then
38:35it's like a hamster wheel. It's like you are manually pulling it and like forcing it to work because like you know that the outcome is going to be great but like I have to remind myself this and our teams this internally because you think of this this magical thing that's always spinning and you just throw things into it. That's not how it works.
38:53Um it's a lot of effort and energy to make the thing actually move. Today's episode is brought to you by Arch Public. Arch Public has just expanded its agentic trading platform beyond crypto. So, pay attention. This is a big one. Now, they are automating strategies across stocks, commodities, and ETFs.
39:08And I think that this is going to be huge. You can now automatically take profits when one market hits new all-time highs and rotate that capital into other markets, showing more opportunity. Whether you're rotating capital into AI stocks, gold, if you're investing in the S&P 500, or you're accumulating Bitcoin, Arch Public brings real discipline and automation to your investment strategy. Additionally, they've launched a powerful new tax loss harvesting tool with crypto being so volatile and its exemption from the wash sale rule. Arch public can offset gains with losses without compromising your
39:37long-term positions. It's exactly what every serious investor does. Institutional grade automation that works across every major asset class. There's no more emotional trading, no more missing tax opportunities, just smarter, hands-free execution of your preferred strategies. Go to archpub.com right now. Connect with our team, set up a time, bring your accountant if you'd like. Then you can learn what automated trading can do for you. archpub.com.
40:02Today's episode is brought to you by Simple Mining. Bitcoin mining has a reputation for being complicated, risky, and hard to evaluate as a real investment. If you're considering mining in 2026, what actually matters isn't headline profitability. It's uptime, repairs, and whether the operation [music] is run like a real business.
40:19That's why I've been using Simple Mining. They're based in Cedar Falls, Iowa, and they run a white glove hosting operation where you own your miners. You choose your own pool, and you have [music] Bitcoin sent directly to your wallet. They were featured on the Inc.
40:315000 list as the fastest growing company in Iowa with over 40,000 machines under management. What stands out [music] to me is execution. They have the number one rated ASIC repair center. And for the first 12 months, repairs are included. [music] If mining margins get tight, you can pause with no penalties.
40:47And if you want to resize or upgrade your fleet, there's a [music] marketplace to resell equipment instead of being stuck. To help people think it through whether mining actually makes sense right now, they put together a short resource called the [music] 2026 Bitcoin mining blueprint. It walks through the five mistakes investors make when allocating to mining, and they also explain how to avoid [music] them before deploying capital. If it sounds interesting to you, you can get it for free at simplemining.io/p.
41:11That's simplemining.io/p.
41:15Go check it out today and see if you should get into the mining game. It's funny like, uh, I thought AI was just going to solve all our problems. We would have no jobs. Um, you know, we just be like hanging out at the beach.
41:27But, uh, I have said it over and over again. Every single person I know is working harder today than they've ever worked in their career. And a lot of that I think is just they feel like it's a big moment. You got to kind of accelerate to be able to uh to capture uh kind of your piece of it. But but at the same time I think that people are inspired right that there is this element of imagine if you can be part of a team that accomplishes you know XYZ thing and you know obviously Deep Mind is is a big part of that. Before we um before we let you go uh let's talk about
41:55is it kegle or Kaggle? How do you actually pronounce this?
42:00Kaggle. All right. Well explain a little bit as to what it is. you're now running Kaggle um and maybe kind of like what your vision for uh for the product.
42:08Yeah, I think what the the sort of um the perspect actually historical context Kaggle was sort of a startup Google acquired in I think 2016 or 2017. It's done a bunch of interesting stuff inside of Google. We sort of brought the team over um into uh to be part of my team earlier this year. And the sort of the basic hypothesis for this is like model progress itself and like this is like such a important point to underscore model progress is gated by our ability
42:36to measure progress. Like you cannot make progress on something that you aren't able to measure. And so actually as you see one of the most interesting things in the last like three or four weeks is like you look at fable 5.1 you look at Astra you look at hopefully Gemini 4 as it lands in the market. like these models are saturating all of the available benchmarks. And so now you sort of sit there and you're like, okay, well, where where do we go? Like we don't there isn't a bunch of problems that are difficult that sort of that we can actually measure and like
43:06scientifically continue to hill climb.
43:08And so the mission for the for the Kaggle team and for building this platform is like we want to build the most open benchmark and evaluation platform in the world so that people can come together and collaborate on um on like all of these extremely difficult frontier benchmarks and challenges and competitions so that we can actually see difficult problems that models can't yet have not yet saturated and we know what we can actually measure. Um, and I think there's like a bunch of nuance bits of this, like
43:37the the everyday, and I I say the everyday person in quotes because like I'm sure the everyday person is not going to, but like people who care about this technology being able to show up on a platform and have a voice in like and a say in, you know, how are we measuring progress towards AGI? What are the things that we should care about? What are the types of tasks and benchmarks that like actually prove these things?
43:58What are the for my for a company? What are the things that you as a company building Sylvia actually care about?
44:04What are the capabilities you wish you had in a model that would unlock entirely new sectors, entirely new geographies, entirely new use cases for your customers. Um, and having a place where you can articulate that in a way that this is my I had this epiphany a year and a half ago where I sat in years of like customer conversations where sort of you would take a customer and they'd say, "Hey, I wish the models could do this and here's an example of
44:33that." And then you'd see a researcher sort of, you know, with a blank stare because like it's the the the work to translate sort of one anecdotal example into something that like an AI researcher can actually take action on is like it's on two ends of the spectrum. It's impossible. It's not it's not capable. And so you have all this great feedback coming in from customers and you can't take action on it from a model perspective. And so the exercise is like how do you get people to speak the same language? the language that the
45:03researchers speak, the language of model improvement is in the form of benchmarks. You need to be able to measure something to make like scientifically rigorous progress on it.
45:13Um, and I think the world is like slowly starting to wake up to this fact. Um, and I want to help accelerate this because like progress is bounded by our ability to measure progress. Um, and so yeah, excited. We're like definitely in the early stages of this. We'll have lots more stuff to share soon, but uh trying to get the world building more benchmarks uh so that we can make progress for the stuff that like real people actually care about, not like a bunch of academic stuff that people don't care about, but like use cases that like you personally and your
45:41company have and every other startup and company has um is really important. I think this is like one of the problems of the of the decade. We we were talking earlier about, you know, one of the things that we've been talking internally quite a bit about that uh I still don't have an answer for. I don't know if anyone does, but when you look at these benchmarks, you know, uh you may see on a scale of 1 to 100, uh somebody comes in at 82 and somebody comes in at 78. And you're like, "All right, well, I know 82 is a higher number than 78, so like they're quote unquote better, but to the naked
46:10eye, does that actually a difference that the human user can even tell? You know, what does that mean? The four percentage points, it's kind of like a benchmark we invented, right? And like, is it real? Is it not? Is it noticeable?
46:23does it improve accuracy or or like whatever the thing is and to me you know the work you guys are doing there but but just more broadly as an industry 10 years from now we'll probably have an excellent answer we'll be like pump I think the nuance to this and this is this is why the building the platform and the transparency matter so much because all of the detail is in exactly what are the four tasks that are different between those things and so here's a great example of this like if you haven't spent any time looking at
46:53benchmarks before. Like you the more time you spend, the more you realize the things that we're measuring quality on and the things that people are talking about are crazy. Like none of it makes sense. Like for example, and here's like one specific example. Um there's some of these new coding benchmarks, and I won't name names because I it's I said that's crazy and I don't want to uh disparage these folks because I think they're doing a reasonable job. But like some of the new coding benchmarks have like 6% of tasks
47:22on like a programming language called Zigg uh Zig. Nobody's ever heard of Zigg before. This is not a programming language that any engineer at any company is actually using. I'm sure some people are using it, but like the 6% difference could be like the quality on Zigg and like maybe some model happened to get access to some data and whatever this language is, but like that doesn't matter for 99.9% of startups. Nobody cares about this thing. And I'm sure they have a good reason for including
47:50that data. But like being able to do this like introspection of like not just there's five percentage points difference between these two models, but like why is there a five percentage point? Does that actually matter for me as a business, for me as a developer, for me as a user of this model? I think it's the whole game. And like the products and services and like even the benchros themselves don't do this right now. They sort of show as in this this empirical thing that you know your 75% should be the same way that I perceive
48:1975% which is completely not true. Um and so I think it's like a fundamental problem with the way that things are set up right now.
48:25It does feel like on one hand uh personalized benchmarks are going to become a thing. Yeah. I don't know how somebody smarter than me will figure that out, but like that obviously is going to be, you know, important, especially for businesses that are kind of like, hey, here's my specific, uh, ramifications. Um, the second thing though is, uh, we came out, um, at Sylvia and we showed that a lot of the harness work and and, uh, things that we had done made the Sylvia product more accurate than the Frontier Labs at answering tax related questions. And to me, I'm like, you know, more
48:55businessminded, not as technical as as the engineering team. I'm like great you know we we we are higher we are better we are more accurate you know etc and immediately the engineering team was like we better publish the eval the you know the rubric we better publish the user parameters like there was this entire effort as to like how much can we publish o and open source without actually giving away things that would be considered you know kind of very important IP related you know type things and we went through a strategic
49:24you know kind of debate internally as to like there was definitely some stuff that we published that we could have not published and it would have maybe given us a little bit more of an advantage.
49:33But it was like if you're not a frontier model and you come out and you say that you're, you know, more accurate on something, you almost have to like open source more to let people validate it themselves. And so I do think that um if you're a frontier model like you kind of don't care if people believe you or not because you're just like, you know, here's the evals, whatever, but the users care, right? Like the kind of what you're talking about, I think, is a is a different situation. And so it's less the academic application for you know who's got the best model and it's more about like I am a business or I'm a user
50:03and I'm trying to evaluate which one of these things I should use for my specific use case. I mean the benchmarking industry is going to be you know significantly bigger than it is today. And you know obviously you guys have kind of a lead there in in what you're doing.
50:15Yeah. Same thing with the data industry.
50:17That's what's most interesting is all this this data moment is and I we don't need to talk deeply about it but like it's having this like crazy I'm sure you're seeing this on the startup side. this just like absolutely ridiculous.
50:28And I think the framing of this is like 2023 the question was like does the recipe work? Do we have the recipe to get to general intelligence to get to this sort of like AGI thing in the future? Um and I think we derisk the recipe and we've made a few tweaks over the last few years but like generally derisk the recipe. Then it was obvious like oh there's not enough comput in the world. Let's blast hundreds of billions of dollars into getting compute online. Um that is the going to be the blocker and and sort of like to keep scaling up. we need more compute etc etc
50:58all the all the labs have now done that we now have enough compute it's coming online there'll be further investment but like generally people know that that's something that needs to be solved it's now all data bound like the data to make progress on the model does not exist in the world like it is data that has to actually be created net new that doesn't exist or is coming from like even startups uh I think there's like a huge like wave of like startups that are going into these like exclusive data licensing agreements with like data providers or model labs and like the PE,
51:27you know, if you want to make progress, it's all data bound. Um, and so it's like this like data business is very tied to this benchmark ecosystem is like very tied to ultimately model progress at the end of the day. Um, and so it's very interesting to see like how quickly these things are like spiking uh up and to the right. I um am very biased, but uh I'm an investor in Micro One. I think they've done a fantastic job on the data side. Um, but I'm also an investor in a
51:56company called Sunset. Um, and Sunset uh, they started out as a company to help other companies shut down. So, if you have a startup, it doesn't work, it's a pain in the ass, right? Like you got to get the lawyers involved. You got to figure out how do I save as much money as possible to get back to investors, but I also have like, you know, kind of a responsible way to wind down. And that's where they started.
52:15They I don't even know if they could spell AI at the time. I love them, but you know, that was not their focus.
52:20Well, all of a sudden they realize like there is a unmonetized asset that these companies have which is like all the Slack messages in Google Drive, you know, just like the the corporate data and could they basically at the point of shutdown buy that data from the company which creates a new asset that then can help them get more money back for their investors. They have to clean it and structure it and you know kind of do all these things to make it a usable form but then they can turn around and they can then sell it to the model labs. And so you almost have this like beautiful
52:48thing where like a normal company would never want to sell that data because they were worried about all the competitive, you know, components and all the stuff. But if you're shutting your business down, you're like looking under the couch cushions for a couple, you know, pennies, right? You're like, "Hey, wherever we can find Oh, you want to buy our Slack messages? We were just going to delete them. So like here, knock yourself out, right?" And so to your point, like I do think that this has happened. I've seen a ton of startups in all blue collar work, you know, medical, etc. They're just trying to figure out how do we go and find data
53:16sets that no one else has has yet and then turn around and let's go and use it for robotics, you know, model training, whatever. I don't know how big that thing can be, but it feels like we haven't even scratched the surface of what that whole industry is going to look like. It's going to be massive. And I think actually the the most difficult part of this, and this is the part that's still like a dark art, is like having data is not necessarily the problem at the moment. The problem is like getting the data into a format that
53:45the model labs can actually use or that the data vendors can actually some of the data vendors are now doing a bunch of this stuff though like it's the most difficult part because like the raw Slack messages like there's like an immense amount of work and labor that's involved in like taking that and finding some way to take that data and like make it actually usable from a model improvement perspective and like rigorously can go and like increase quality in some dimension. Um, so there's like a I think there's even just like that business of like helping
54:14companies understand how they can actually make what's the value of their data and all that stuff I think is a is a really really difficult problem that it feels like we're still early in trying to solve. So um and is like fundamentally correlated with like if we can do that we'll see more model progress 100%. I think my takeaway from this conversation Google the frontier commitment is real. You guys got a lot of stuff going on. I think you guys are doing a great job. Um, if people want to connect with you or or find you online, where should you send them?
54:42Uh, X. I'll see you on X. Ping me. I'm also on uh LinkedIn if you want to ask uh more boring questions. So, [laughter] all right, my friend. Thank you very much for doing this. We'll do it again in the future.
54:53I love it. Thank you for having me. If I'm car reception