0:00Oh, like I love cloud code. I love codeex. I love pi. Like no, I love oh my pi. I love prime agent. To be honest, I think all of them are the same.
0:09You know, your your models actually are capable a lot more if you try harder. So this is a skill issue.
0:14Yeah. Yeah. Basically, I I I mean I I I think there's one thing I want to see. I I appreciate that there's a big focus on like jagged intelligence because it paints a big picture of like we can do this if we really set our minds on it.
0:30But I kind of wish and and maybe someone in academia should do this. Like really just sit down and think about like if I took Astra even the current frontier models are not good enough at like doing a particular job over let's say the span of a month consistently and well. It's uh 10,000 agents in 88 hours. Um 130 billion public tokens which is estimated to be about $40 million in public pricing.
0:56May maybe you'll have to pay like $40 million to get a result. And it's like well but it's exciting. I I I will say like it is it's very exciting that we even have the option to point $40 million at a problem and solve it.
1:11Before we get into today's episode, I just have a small message for listeners. Thank you. We would not be able to bring you the AI engineering, science, and entertainment content that you so clearly want if you didn't choose to also click in and tune into our content.
1:24We've been approached by sponsors on an almost daily basis. But fortunately, enough of you actually subscribe to us to keep all this sustainable without ads, and we want to keep it that way.
1:34But I just have one favor to ask all of you. The single most powerful, completely free thing you can do is to click that subscribe button. It's the only thing I'll ever ask of you. And it means absolutely everything to me and my team that works so hard to bring the Inspace to you each and every week. If you do it, I promise you we'll never stop working to make the show even better. Now, let's get into it.
1:59All right, we're here in the studio with Alex Jang. Uh, I guess most famously of RLMs, uh, but you you have a lot a few other affiliations. Welcome to the show.
2:08Yeah, thank you for having me. Yeah, I guess GPU mode as well.
2:11Yes, GP mode as well.
2:12Uh, you were shephered in by Mark Sarafin. Not Not everyone gets that kind of welcome.
2:16Yep. Yeah. Yeah. I'm very very close to all the people in GP mode. So, yeah, we often end up working together in various capacities like even beyond just GP mode itself. So, yeah. Can we explain uh so people who are not that close don't know about this. It's a it's a discord. It used to be focused on I guess CUDA mode and then then generalized a little bit. Uh it was started by Mark. Yep. It was basically like to me it's like the hiring pipeline of the PyTorch team and then you left PyTorch.
2:46So it used to be I think uh it started actually around when I was in college um in like 2023. Uh I think it was started by Mark Andreas and Jeremy Howard. The original premise was just like it was a GPU or it was a discord dedicated to learning how to write GPU kernels and they had like lectures. That was basically the extent of it. Um, and I got interested in it because I was writing GPU kernels. It was actually out
3:14of like pure chance. I was interning at Snapchat at the time and Rexus.
3:19Yeah. I was very bored with Rexus. So, um, they had a project where like they were interested in writing. It was this paper called Infinite Attention. It was like a Google paper.
3:31Yes. We've covered it on paper club.
3:33Yes. Yeah. So I was interested in in whether or not you could write specialized kernels for it at Snapchat.
3:39Um it didn't nothing really came of it but I joined GPU mode at the time. It was called CUDA mode. Uh I think for like legal reasons or something they changed the name but I met Mark uh I met Mate I met a bunch of other people that were very involved in the community. And then Mark had pitched this idea called Popcorn which was now what you see as the leaderboard today. But the general idea was like I think all of us had this like intuition that GPU programming is
4:07like very similar to if you guys have done like competitive programming. It's it's a not I I don't mean to say like they're transferable skills. You code golf a little bit.
4:16Yeah. Yeah. Yeah. And there's like there's actually a surprisingly small space of optimizations that people do.
4:21Uh and there's actually not that many kernels per se that people are interested in optimizing. And so we kind of had this thought that like if you had enough data like in the same way that code force is there's millions of problems. If you could do this with GPU code like you could scale and automate kind of GPU kernel development which for researchers is a huge deal cuz I think one of the bigger bottlenecks like if you look like mamba for example like they release the paper with kernels because otherwise like you can't really
4:50use it in any meaningful way you know and not everyone has like a tree on their team. So we're very interested in this. Uh kernel bench kind of spawned from that too of like can we get LLMs to automate um GPU kernel code and I think that was like a it was a very very fun time. It was like between college and my PhD and yeah I I had a really pleasant time doing stuff with GPU mode. Now I kind of just help with the lectures sometimes. Um I'm not as involved and I think in general like we don't have as
5:18many competitions as we used to but um yeah I still keep in touch a lot with with everyone there. Is there a friendly rivalry because like I think the previous community that used to do this like MLS ML Perf thing. Is there a friendly riv rivalry? Is this like just new generation MLF or what's going on?
5:34The nice thing about GPU mode is that it is also a community in the sense that like a lot of the lectures are very easy enough for a beginner to follow and ask questions and things like that and like the competitions are like somewhat not secondary but like you can participate in them to learn. I think with a lot of like MLS, ML Perf kind of benchmarks like for the most part like only serious labs and companies participate like like seriously in them at least. That's that was my understanding of it. Um I I could
6:04be wrong, but I I think also beyond GPU mode now, one thing that has been really exciting is there's a lot more websites and like people that work on hosting competitions. Like I think there's this I think there's this website called like leak GPU or something and it's like leak code for GPU problems. Um there's like other ones too that I like we've we've seen like there's many that have kind of spawned and like talked on GPU mode and like um it's very exciting in the sense
6:34that I think GPU programming used to be super niche like when I was interested in it and the only reason I got interested in it was Trio gave a talk at Princeton because he was applying for faculty I mean he is faculty there now but I listened to his talk on flash tension in like 2023 and I was like wow this is like the coolest thing ever. Um and I was like this is like this is what everyone should be working on. I guess like VLM and stuff had come out too and it was like oh you know we should be writing kernels but now it's like you know everyone writes kernels like
7:02everyone it's it's I think it's actually almost saturated in some sense uh as a field.
7:08Any any interesting takes for people that want to get into it? So I think one of the biggest news is GPT 5.6 wrote more efficient kernels so Terra and uh Luna could be 80% cheaper. Yeah. And then we've seen other competitions where people are like setting records and they're like we're doing some auto research loop and these are people that don't have a background in any kernel writing. Right.
7:33Yeah. So even on the GPU mode leaderboard uh if you look at like a lot of the recent problems almost all the solutions are AI generated. However, you'll notice on the leaderboard so there's this guy named Gaurst who is like a very very like regular member of GPU mode. We've always known for a long time that he's like a super super corrected like GPU kernel writer. One thing we discovered on this leaderboard is like almost he also used AI to help him with these solutions, but for the
8:01most part like he helped prompt and move it in certain directions. Uh we found that like his kernel was like basically the only one in like the top 10 that was actually stable in like actual like endtoend systems. And it does bring into question like it's not like I mean GPU kernels have a verification problem.
8:21Like we've kind of known this. It's been a problem since kernel bench was released. Like there's a lot of reward hacking that goes on. But you Yeah. You also notice like the lines of code is a lot a lot smaller.
8:30Is that noticeable or is it just Yeah. No, it's it's it's definitely like very important. And I think like it's it's really interesting that still there's a lot of alpha in being good at.
8:41There definitely is. Yeah. I I think like and this applies to a lot of AI systems as well like I think you know even with the most recent like math proofs and stuff like it doesn't necessarily mean mathematicians are obsolete I mean these companies still hire mathematicians like to do whether it be like data labeling work or or even just like steering the models to solve problems like there is still a lot of alpha and being like knowledgeable in these things. So is it just knowledge or is it also
9:11there's just more planning and is there uh an emergent style of planning that works better?
9:18I think it's it's a mix of maybe this is what you mean like intuition for how to solve the problem.
9:25Like for example I always diagram my code.
9:28Right. And then like if there's a part of the diagram I don't understand I work until I understand it. Otherwise I it's not allowed.
9:34Yeah. Yeah. So I think it's like it's a mix of those things of like the people who work like the people who know how to look at these problems and how to solve them like also know how to use AI to do them because like you're acting as a very strong verifier like if you are knowledge or if you know what to do and you're also like I I think the thing that we've kind of discovered with all these agent swarms and things like this is like when you throw enough compute at a problem you like can sufficiently
10:02explore solutions to that problem. But often times like you know maybe you can burn like a hundred billion or a trillion tokens on something but if you bring in someone who knows something about the problem um they can uncover something for the model that would like erase that one trillion token spent. I mean it's not it's not super clear like what exactly the trends are here. But I think like there are so many problems in the wild still right now that we want to solve and like we can't afford to just always you know throw as much comput as
10:31possible at it. like there is still an efficiency aspect of of all of these things that is super super important.
10:38Is there like a theoretical right answer that you can just calculate based on physics and then you just get close to the physics limit?
10:45Yes. So for GPU kernels you can compute it's actually not that easy to compute sometimes like depending on how complex the problem is like for for matrix multiplications it's very easy to compute um this like speed of light kind of uh estimate of what the fastest kernel can be and like this is also assuming like you know maybe all of your all your data starts on the CPU or maybe it starts in DRAM on the GPU etc like this changes these these numbers slightly
11:14the transfers and all these things.
11:16Yeah. But I I I will say like it's not clear though like in a lot of cases if it's even possible to hit this theoretical number if that makes sense. Like this is assuming like perfect overlapping and transfer of data and and like there's maybe some bottleneck that you can't get around but often the kernels are not even close like that we write are not nearly close enough to this number to be like meaningful at all.
11:40Yeah. And is it speed that matters? Do you also care about uh obviously memory which which feeds into speed. Do you care about power consumption? So one of my uh one of our top pods of the year was Jeff Dean who was like actually I just checked the microJoules or like the nanojles pico jewels.
11:55Yeah, pico it's optional pico.
11:58Uh do you care about that?
11:59So I don't I guess Vivie I'm not I'm not as but everything here is speed right? Like nobody's counting ples. Um but I there's a caveat here which is I think like there is speed in the context of a single kernel and there's speed in the context of a larger problem like maybe the N10 model because like one thing to consider and this is why it's important to talk about what what speed of light is referring to because in these cases for the kernels like we always start
12:28with everything in in uh like HBM for example right but you can imagine that like an end toend kernel like an end to-end model. What you might want to do between two layers is like you might sacrifice the speed of the first operation to keep things in the cache for the second operation. And like these are things that like you can't really get out of in isolation with like these kinds of kernels. And you know people call this like the fusion pro or like
12:57the fusion problem. Yeah. Or like mega kernel stuff. And it generally only applies like when you are like memory bound in in in most cases. But this is something that like also there there is this question of like as these models get better like should we just be generating like mega kernels? Is that like what we want? Um what's your take? I think that this is really difficult because you need the data to do this. And I
13:26think like I have yet to see an example in the wild of like we bootstrap the ability to solve a very difficult class of problems without any examples. Um, and I think like the other reason why I think maybe this isn't that interesting is that at the level of an individual kernel, a like they're not that they're not as complex, but b you're somewhat confident that there's not as much structure in a single kernel. But like
13:55in a mega kernel, like I would be more inclined to believe that like a compiler would be better here. like some compiler over like higher level ops makes sense because in general like actually I think mega kernels are very like the pieces are very composable of like the individual kernels. There's some areas where you might want to do like weird fusions and everything but in general I think these are cases that like a compiler can probably handle and there is a company that's working on this from from what I understand that has given some talks on GPU mode as well.
14:25Yeah, I want to basically cluster all the GPU mode discussions here because obviously there's other parts that we need to move on to. I think there is something to plug. You guys do host a lot of really good lectures. They're all on YouTube. People can follow along. And you lead quite a bit of it. You're still quite I used to uh sometimes I still do. Uh I think they're mostly Mark. Mark is the one who usually does them. Mate does sometimes as well, but um yeah, I highly recommend them. They are extremely good resources. Like I think it's kind of
14:52crazy how much people share on there. So yeah, because like if you're there like you're very very like you're exactly the right audience, you know? Yes, exactly. This isn't going to reach the main and there's a lot of like introductory material as well um that we've put on um that I think is useful for people.
15:07Colonel Bench was kind of influential. I just want to see like you know that was last year. Uh what other ongoing work do you want to shout out that people should pay attention to because obviously you're involved in this field. Yeah, I will give maybe the in the background story of like I am actually involved in a lot of benchmarks or I used to be uh maybe prior to my PhD. It started because I was at Princeton uh I worked with the Sweet Bench team there. John John Carlos here they're all great like I love them. Yeah.
15:34There's basically this I think people don't understand how many how much benchmarks come from the same group.
15:39At Princeton Do you know Shun?
15:43We had him on the pod before. Now now he's like running 10.
15:45Yeah. Now he's like he's like a superstar. Uh when I when I I met him so he uh he was advising my friend Michael Michael Tang who is now at anthropic but they work together a lot. We were like the two undergrads in Caric's lab. I and then some others joined later as well.
16:03But yeah Shinu is great. I I did not know he was like such a superstar until like later on like after I left. But yeah, I mean like like okay, so there are very few PhD students like yours yours is like the next one like once a year we feature someone like who is like basically entire PhD uh has been like on target.
16:21There's not that many of them. Shil was like clearly one of them and uh you know Jack Jack Morris is another one. And like you know we talked before the show we talked about research taste right like somehow some grad students just have a very blessed career where like yep mostly like yep yep yep yep yep this is like going to stick around relevant everyone should know this and then others just nothing.
16:40I think this is also true of like even people within like industry labs as well. I think it's just like grad students are a lot more visible. So you just see like you see like you know there are some people who really like get lucky in like or I mean it's a mix of being lucky and also being very smart and things like that. I think like with with research tastes as well like I think it it gets it gets developed through opportunities at least in in my case like I got I was very fortunate to have like taken the path
17:10that I took like working at Princeton and then like later like finding my ad like Omar at MIT like he's a fantastic adviser. I will say though I I find that the most successful research from grad students or like in academia comes when people care about problems that maybe like most people in industry are not looking at. I think this is the issue that like a lot of grad students work on things that benefit like that look good to an industry lab. Like for example,
17:39they'll work on some like some benchmark. It's really popular now. I mean, I think benchmarks in 2023 were a very different story than benchmarks now. There are a lot of people that work on like harnesses and like meta harnesses and like specific harnesses for XYZ task. And when you really think about it, the reason someone would work on this is like maybe there's like a clear goal shaped around the models that we have today of like this is what I
18:08want to see. But like I'll give I'll give the the like RLM like the recursive language model paper as an example because I think like it's a super super simple idea. I think when it came out as well like there were a lot of people that were like when they see something like that they're like what is even the purpose of this or like yeah like why like what this is just sub agents or something right? Uh, and I think it's like when you get a reaction like that, it's almost like a good sign in the sense that like it's clear that
18:35people aren't thinking about what the purpose of this is. And I I will give another example of like Sweetbench. When Sweetbench came out, I mean, OIR loves to tell this story when it came out like nobody cared. Like everybody was like this is an impossible task. Like why would we ever even consider this as a benchmark? And it wasn't until Devon came out that everyone was like wo like this is something we want to hill climb.
18:57And I think this is this this rings true for you tend to see that a lot of ideas I I think like the my favorite I guess example of this is Eric Zelikman's work uh with like star and and like quiet star like I think when you read the paper at least when I first read the paper I was like is this not like an obvious idea or maybe not I don't know I mean I was like oh this seems really simple or like chain of thought same thing or like shouldn't use react it's like okay like yeah I mean Sure. But
19:27then like when you really think about it, it's like why what is the value of the paper? And I think it comes from like it tells a bit of a story as to like what you want the field to look like. And that is something that it's very hard to do this in academia because if you look at all these papers, QuietStar, React, RLMs, Swebench, none of these papers are it's not like a GPT6 Astro release, you know? It's not like everyone's like, "Oh my gosh, like I'm going to use this now. This is the best thing in the world." like academia just
19:56can't afford to do this at least right now. I mean that there's a whole slew of reasons why I think that should change, but I think it's like if you don't have as a PhD student, I think you're in such a unique position where you can work on literally whatever you want for the most part. If you're not taking advantage of that and working on things that like nobody cares about or like people see as some trivial thing like, "Oh, I thought about this, but like I don't use it." I just think like in the end the research is just never going to be that
20:24interesting because you kind of need to take big bets if you're going to be in academia because otherwise I think like just go to an industry lab like you know they have tons of resources, tons of talent. Why constrain yourself in an area where you don't have a lot of resources and like there's not even that many people around and I think it's just it literally just comes down to like big bets. um you just have to take big bets and like a lot of them will fail, you know, like that's just it's it's
20:50natural. But I think that is as a PhD student like that's the biggest advantage you have over any single person at another lab because you don't have to deal with bureaucracy and all these other things.
21:04Yeah, I ask a lot of people this question and usually they hand wave away. So I think I appreciate that you're actually think giving a thoughtful response on like no like uh this is your unfair advantage because everything else is biased against you basically.
21:16Yeah. Yeah. Exactly. And so like it's honestly I will bring up Jeb as an example because it's it's not an it's not an academic project. I want to bring this up because this also happened with RLMs and uh it happens with many other works like things get overhyped right to to an extent like something gets overhyped and then people are like why is this overhyped? like this is trivial, this is stupid. And I saw the same thing with Jev because I think the release was like I mean there's this whole thing about like academics or they're not academic group but like people have to do
21:46branding and they have to like kind of market their research and so like I understand but I think there was a lot of discourse about Jev just being like something we've known for years and I think it's kind of missing the point of like why is such a system so interesting. It's why is it not just some stupid MLP classifier that like we've we've been doing, you know, back in our intro ML classes or something. I think what's really interesting about
22:12Jev is that it kind of opens up this question of are language models correct?
22:19Like in the form that they're in, can we consider a different design space other than text to text? What what they're doing is they're basically saying like I will take advantage of this language model backbone like I know it captures a lot of information about language but I'm going to change the output space of the model to give you a trade-off which is I will do very very fast inference over like if you have some prior about this problem like let's say I only need
22:48to make a binary classification am I going to ask my language model to do this and pay like a 400x cost like no it's it's silly, right? Um, and I think for the longest time because the labs are the only places that control, you're never going to use something other than like GPT, GP6 or Fable because they're the best models. But because of that, like people have gotten kind of accustomed to this idea that a language model is just a auto reggressive
23:16decoder. Like we have accepted this. And I think when RLMs came out, it was the same thing. Like one of the comments I like a very frequent criticism I got was like this is not a language model or like when I look at this like I thought it was a new architecture but it's actually not and my response to that is like well a language model is just modeling language it doesn't have to be this transformer decoder you know uh and Jeb is really really interesting in that like we now have a new thing to tune
23:46which is like what is the output space and how does this affect inference latency And I think we can actually start this various parts of the language model itself.
23:59It's a similar idea of a lot of the attention around it was like this is a silly idea like why who cares about this? But it's like it is a simple idea but it's actually it opens up a whole new set of questions that I think like especially if you're a PhD student these are the things that you want to answer because I I think it's like we don't know for Jev for example we don't know how far we can take this uh for loop transformers we also don't know how far we can take this what if you loop only a
24:27subset of the model um what if you route to like only uh like you have some router to different parts of the model like can you mimic what you do in a harness inside of the model? What can you bridge between a harness choice and a model architecture choice? These are all questions I think that open up with works like this. And that's like really exciting. Jeb in particular, when I saw it, I was like, this is actually really useful for RLMs. Like I think it's it's it makes sense because the the biggest
24:55bottleneck in RLMs or swarms or or systems like these is they're slow. when you do multiple language model calls all the time, you're not distributing your compute correctly because like maybe there's something trivial that you just want a simple model to do, but you can't do it because your language model is just this this bulky thing, you know?
25:14So, I'm very excited. I think I think we will start to see new types of models emerge beyond just the bog standard frontier model. And that is like there's so many things that you can do with these like new trade-offs. Um I also shout out Thinky with their interaction model.
25:33Yeah. Yeah. Another great example.
25:35Yeah. So like basically try to break the paradigm from sequence to sequence decoder only and just literally do anything else. Yeah, I mean I think there's there's a level of if you're trying to compete, you're not going to compete with a frontier lab doing an auto reggressive decoder on like the amount of compute scaling resources they have. Even thinky will not, you know, okay, there may be one of the handful that can, but you're not really going to do much in that. I don't know too much about Thinky
26:04or I don't want to say anything either, but it's like if their strategy is just to replicate open AI or anthropic like that's a horrible strategy because well because like I mean they just don't have like you you kind of just have to think of it in terms of like what advantage do you have and if if you're going to use the same setup I I'm sure they're not but it's like if you're going to do the same setup like you're basically competing on the things you can control which is data and compute and obviously they cannot compete with the Frontier Labs on that. So yeah, it makes sense
26:32that like if you are a Neolab like actually I don't know if you would consider them to be a Neolab but I guess like I guess they're kind of a weird one. Yeah, they're in that list. They ship inkling like they count.
26:42Anything other than open anthropic maybe like Meta and and and GDM like you just you got to do something else, you know, like it's just it's the sad reality, but I think I I actually think it's a good thing. I'm very happy that like scale works and these companies will just keep doing it because it opens up like potentially new players like if they uncover something really really interesting because I sort of have my doubts that this is like seriously going
27:12on at Frontier Labs cuz it's like like why would you do that you know like why would you take the risk of allocating a large amount of compute to new bets when the old bet already works. And I think that's what spins off a lot of Neolabs, right? You have a side bet and you don't get compute and you're like, "Okay, I'll go I'll go and you know your example of the potential upside is something like Jev, which is X00 times cheaper, you know, comes out and maybe it is language model. In this case, it's just different."
27:39Yeah. Yeah. Yeah. Yeah. Um, so no speculation on on what Jev actually is.
27:44I guess I have some guesses for what it might be. I I have seen some people say like ah it's like a diffusion thing. Um I guess that which is the parallel decode.
27:54Yeah, parallel decode. Honestly, I think regardless of what it actually is cuz I think you can I've seen some like open source replications of it. What is really exciting to me about what they did is I'm not entirely sure what their optimization objective was and how they trained it. And I think like this is a thing for RLMs that we've also been thinking about which is like okay like RLMs are a very simple idea. If I come out with this paper like anyone can use
28:21it now. But what distinguishes the actual value of an RLM is whether or not you can train it properly and whether or not maybe you can mold some architecture around this system to make it really really good. And that's something that like I mean I'm actively working on I guess but I think for them like they figured out a way to train the system which is completely non-trivial.
28:43Like I I actually don't really know how they did it. And I've seen some comparisons online of you know some people are claiming they used Quen or they post train on top of Quen but every open source Quen that you use is going to be worse cuz whatever they did to train it clearly works very well. And so that's that's very exciting. There's one element of calibration which uh is a is a rare topic that I don't think people even knew about or understood. We covered it with um our conversation with
29:11Clementine Forio of Hugging Face and she used to run the evals um at Hugging Face, which is basically the idea that uh you know models are attuned to give you the most likely next token, but uh they're going to lie to you when you ask them how confident are you because they're just going to give you the most likely next answer instead of like actually like no let's calibrate like I am actually 50% sure, I'm 20% sure and like let's try to calibrate that. I would say like if anything I think that actually that's pretty easy to generate
29:39synthetic data around uh because you can sort of see the truth and then synthetically generate a bunch of answers have Jeff classify it and then compare with ground truth that that would be my reverse engineering I've actually I think calibration is probably the under like people are just using it as a very fast classifier but they're actually not even using the probability or calibration estimates estimates I I think it's also still just misunderstood to to reiterate when you ask a model, how confident are you it
30:07will spew out what 43%. Uh the the big delta is this is a grounded classification, right?
30:13Yeah. Yeah. I'm I'm very excited to see what people do with this model. I mean, is it going to solve everything? Like, no, of course not. But I think it solves a class of problems that we traditionally struggled with, which is low latency things. So, I I love the examples with games. That's actually like I mean I Yeah, the Doom example. Good.
30:32I have a benchmark on language models playing video games. I've always been fascinated by whether you can have an intelligent system play new games and things of of of this nature. And so I think it's really really cool that they have sort of a unique way to do this to capture language and understanding in like a fast a very very fast model.
30:55Well, while you're on the topic, anything you want to point out?
30:59Oh, yeah. I mean, you know, you did do a benchmark on many anyway.
31:03Um, I guess these numbers are very outdated. A lot of the models are are are very different. And I've seen actually people run there there are some folks out there that are running newer models on these games, which is really cool. I guess the general premise of this benchmark was we just want to see if vision language models are are good enough at just like plugging into games with the latency constraint included.
31:24Cuz this this actually I came out with this right after Claude Plays Pokémon came out. Um, so this was like 2 years ago, which I guess is like ancient now. But I think what's really really cool about this suite of tasks, it's very diverse in terms of what games they are.
31:42And also, I think most of the games are games that people know or like have seen before. I saw Yeah. Jev playing Doom. I will say I don't think I think they were just playing like really simple levels and stuff, but honestly like most models still can't really do or I don't actually think any models can solve these games very meaningfully. Like there are some games that they can. I think I've seen Astra be able to solve the Kirby game. And we also for this benchmark, we intentionally designed a
32:12really minimal harness. And I I will get to this point about harnesses because I think there's a whole conversation to be had about like what is the value of a harness? Like what is even the purpose of a harness for a model, but in general like I think it's uh yeah, I I hope to see very quickly or very soon like all of these games beaten by newer models.
32:31Yeah, it's interesting like the old old Cloudplace Pokemon, they like read state from RAM and saw what tiles are walkable and what not. We did a podcast with them a long long time ago.
32:41Gotcha. Gotcha. Yeah. Yeah, I mean it's similar like Jeff doesn't have vision so you have to kind of feed in uh these like uh game state and all these things. I mean let's go right into the harness stuff because you you brought it up. Language model harnesses a compositional generalizes struggle to read that explain.
32:58Yes. Okay. Okay. So I have been a little unsatisfied maybe with with how people think about harnesses because people compare like oh like I love cloud code I love codeex I love pi like no I love oh my pi I love prime agent to be honest I think all of them are the same most of the design decisions or like the design choices around these harnesses are the same maybe prime agent is a little bit different cuz it's like inherently an
33:26RLM but in general like I think we can be a lot more creative with harnesses.
33:30And what I mean by that is if we think about this from the perspective of what exactly is the harness doing for the model well basically when you're trying to solve a problem and you want to use a language model to solve it like a very difficult task. One thing that we have discovered is that next token prediction is a really really awkward form to do a lot of these tasks. So for example, take Swebench. When you're navigating a codebase, like are you going to be able
34:00to figure out how to do all of this with a single language model call? Like you just say solve code or like solve my query over this codebase? No. So we rely on a harness to help you do these kinds of things. And I think what what is interesting is like a harness is a very very opinionated program over how you want a language model to be form fit over a problem. And I bring this up
34:28because when we think about the what a harness is doing, we should really think about like what choices in the harness let the language model solve this task and can I actually just have a language model that just does this? Because a harness if if you think about it now that loop transformers are a thing I think what's really interesting about it is you can model a looped transformer in some ways like with a harness as well, right? you're just looping over the model. Now, you can say like, "Oh, I'm not decoding, so it's like a little bit
34:57different, but in general, we for whatever reason have stuck with the same model architecture choice forever." And there's many arguments for why, but clearly like we are now training models around harness rollouts. And so there is this very awkward way of doing training over harnesses which is that we train a language model to act within a harness but like now it's like a really long maybe like multiple agent roll out that we're doing and there's like really
35:27awkward hacky ways of doing this. So what this blog talks about is like well one way you can think about what is going on here is if the harness is basically helping the model solve a particular task. Can different harness design choices actually do something a little bit more meaningful beyond just here are some tool calls that will help you. Here is a way to GP through your codebase. And so this actually the idea
35:55for this blog came with the RLM idea as well. We just didn't package it that way. And I think this is actually true. There are many other ideas around RLMs that like we will be coming out with, but we're all there from the beginning.
36:08I mean these are all design decisions around um I think with what what is what I like about the RLM is that there were many many many iterations and versions of different abstractions that I was interested in doing and ultimately the RLM made the most sense. But there's a lot of reasons that aren't public as to why that's the case. Um I mean you you will see. So in this example, one of the things that we see with an RLM is that if you sufficiently offload context and write ask the model to write code over
36:37that context, you get this really really weird but useful property which is that when the model recognizes during training how to solve a task, it turns out that the solution to many tasks is very similar. across like tasks where you you don't even like it's not even that clear to you that the solutions are similar. So in this example, we have like a retrieval task and we have like a an aggregation task and they're very very different query. Like the domain is
37:05just completely different. And when you train a regular language model over these two tasks, the trajectories look very different. And so what you're relying when you like let's say you use pi or cloud code or something, which is not in this blog, but we do have these results. You'll find that like these harnesses distinguish too much between these problems even though the solutions are the same. And so one thing that we find when training RLMs is that like when it sees these problems as the same.
37:33And the reason it sees these problems as the same is the sub agent sees different problems, but the sub agent is solving an easier subtask. And so you're confident it's smart enough to do it. But for the base overall strategy, they end up looking the same. And so when you train on the left task for example, it can immediately solve the right task.
37:52And so if you go down to like the plots that we have, one thing you'll kind of observe is that the as you just naively train your RLM on these tasks, they naturally learn to generalize, for example, to longer tasks because the strategy is basically the same. You're just modifying like a length variable. And this actually also holds for tasks that are different. And it's not even they're not different across length.
38:16They're completely different tasks, math tasks versus writing tasks, but the solution the like meta highle solution is the same. And so when you train the RLM on one of them, it generalizes this behavior to the second one. And there's no magic here. I I guess maybe that's the thing that I want to kind of stress like how how would you kind of verbalize what they're learning? So I think in here you say you train on short tasks.
38:41They generalize to stuff 8 to 30 times longer. They are learning how to solve these type of problems or what's the what's the core thing they're actually learning?
38:49Yeah, they're learning how to solve these types of problems at a certain length and it turns out that when you take this strategy that they learned, it is directly transferable to the longer length. Like they're effectively the same program. And this is something that is really exciting because what this kind of implies is that when you have a corpus of data or environments that you train your model on, the hope is that like a you can train on less
39:18environments and generalize to more than what existing models can can do through na or harnesses through existing kind of like naive training. But b also you still want to use all the data that you have. So when you train on these tasks like hopefully it generalizes to a wider class of of problems and why this is also even more exciting at least in the context of RLMs or any recursively calling system is this argument holds inductively. So like I'll give you an example cuz I said I claimed that
39:46competitive programming and GPU optimization use very similar skill sets. the model can the harness potentially I mean you might have to nudge it a certain way but it can learn that like okay how I'm going to go about solving this GPU programming task is very similar to what I learned for competitive programming so I'm going to list out a set of solutions I I'll like spawn sub agents to list out promising solutions and then I'll like write this loop to go through and check these
40:14solutions maybe evolve them and like evolve them against a verifier and between these two tasks this looks the same, but what the sub aents are doing are maybe like unique and something something like different, but even what the sub agents are solving might actually also be of the same form, right? Cuz it's like a it's a recursive argument. And so what I'm trying to get at with this whole blog post is just that like we should rethink what the role of the harness is because harnesses
40:42can actually greatly increase the generalization capability of your model and the amount of data that it's given. And this is not exclusive to RLMs. I think there is a wide class of harnesses that are yet to be discovered that actually can yield similar properties.
40:58And to extend this argument a little bit further, I think what you can also extrapolate from this is like if I look at an RLM, what are the components of an RLM that are actually necessary? And can I actually just directly train a model to do this? Can I train a model to act as an RLM implicitly in its forward pass? It's a really weird thing to think about because like you know you might say like ah code is non-ifferiable blah blah blah blah but there are many
41:25approximations of of this behavior that we will start to uncover and I think like we will see beyond just like I'm going to design a new coding harness that uses a special form of compaction or something. I think we can be a lot more creative here. Like there's there's so much we can do with these language models that I think we are just not doing. And I'm like very excited about this because I think like I think we can get serious gains from very very
41:53opinionated and good harness design that lends itself better to scale. And what I mean by this is like you know the RLM for example is a very primitive inductive bias. Like there's nothing super special about the design other than the fact that it's very different than what we currently do. But this may potentially scale much better with like the data and the environments that we have available to us.
42:16I guess the you know opposite thing that people would probably ask is current harnesses are very generalized towards coding which people see works for a lot of domains. Cloud code is being used for design presentations everything. um Muse, Spark, Grockbot. These are very simple, non-opinionated harnesses that are good at code and that is also scaling out. Um what's the example of how we improve those I guess?
42:45Yeah. Um let me let me bring up another paper uh which came out very recently. It's like the harness tax paper. I think it's by Arena. I really like this paper because it puts forward a prior that I had, which is basically that like most it confirms a prior.
43:05Yeah. Like most harness choices don't matter because all of these harnesses are the same.
43:11But I I will say like Grockpot, for example, is actually quite different, I think, from my understanding than than how some of these other harnesses have been designed. And I like that a lot.
43:20Um, and I think it's clear from here at least that like I'm pretty sure Anthropic or OpenAI are exclusively training on their harnesses. They're probably not training on their competitors. I mean, I'd assume not because I I don't know why they would do that.
43:34But I mean, you know, this is a thing you see in open models, right? Like Quinn is really good at using open code. They need to train in harnesses. Old Gemas were notoriously bad at this.
43:44Models are good, but you need to train in a harness. I I think though as models get smarter or like as they get better this distinction becomes uh not that important in the sense that like if you take Astra and you put it inside of open code like it's not going to go crazy because I think it's like just sufficiently good. And so I say this because the only benefits between these different harnesses is just cost for the most part. And and I think like what I'm getting at here is if you plug these
44:14models into RLMs though, they're not that good still. They're okay. And I think it's mainly because the types like the class of harness that we are training around is this class of harness. This like pi loop this like I I like to call it trajectory as a prompt which just means like you keep the whole trajectory of the roll out as the context that that that your main model is is using. Even if you use sub agents, it's still like kind of this form. And I
44:42think we're going to if we want to explore new harnesses, there needs to be teams that are dedicated to actually running meaningful experiments over like scaling out new harnesses, like maybe post-train scaling out on different harness designs. Like I I think we can actually get very meaningful knowledge or gains from doing this kind of thing.
45:04Uh whether it's an RLM or whether it's something different. And that's exciting because I think for example if you train a lot on Fable for a long time was the best model for RLMs because they had dynamic workflows and it was pretty obvious that like this was a capability that was somewhat trained in even if the model was still like a little dumb like in the RLM harness it still worked a lot better than than other models did. Astra is now also like good enough at doing these things.
45:30But is this is this where we see that Fable is the best for RLMs or is there is there some other Oh no these are all internal results I guess. Yeah, I don't I don't have them public right now, but in general like I think you can you can very easily tell that we have not optimized for RLM like workflows yet. And I think if we get models that do this correctly, they will be a lot more efficient as well. I mean, you can kind of just think through like why this
45:58is the case, right? I think this is the point where you have to give the 10-second what are RL because there's a lot of listeners here that we're assuming a lot of knowledge also I think but you have set some context so you can like with everything we just said can we have a clean crisp definition of RLM? Yes. Okay. I want to go back to the blog the the compositional generalizers blog this one. Okay. This is like the best I wish I had this in the original
46:24paper. An RLM is basically just a harness design where the only tool in the harness is code. Uh which is this programmatic sub aent calling thing where it has the option to call itself as a tool and it has other tools but all of these things are functions in code and the context that it's dealing with is always stored in some memory inside of this code environment. So this could
46:54be a file system like this could be like I'll give you an example prime agent the trajectory of prime agent like the context even the when you compact and do all these things is stored on disk so the model can always reference its original context even if it's compacted and all of its tools are run inside of let's say like a python ripple or a bashrepple and so this it's like this very primitive abstraction and I would say like where most harnesses differ is
47:23a context offloading is not done that like that. I mean if you look at prime agent by the way prime agent does context offloading but not all the way.
47:32So like it still maintains the standard cloud codeex loop of like trajectory as a prompt where you compact but it has the additional kind of like the context is offloaded and it only has the unique point of of prime agent is that the only tool is is IPython. So this is like the very kind of generic abstraction around like RLMs concretely. What's the core thing RLMs are trying to solve? At one point I think when it first came out it was context and
48:02I'll pass the question.
48:03Yeah. So the original problem was harnesses have a really bad time dealing with long context. They typically were used to only really do them for like specific things like code. Like you know they could deal with your codebase because it was trained on it. But now it's more around what this blog is talking about which is compositionality and the fact that like I think harnesses we want to have language model systems that have much
48:33more control over the actions they make at every step. And what I mean by this is tool calls are very limited because you have to invoke them every turn. like you have to invoke tool A then tool B then tool C and there's no central context that you can kind of draw back from and RLMs are specifically like designed around composition and having like a central context that you can always draw from. Um and this context is
49:01like designed around the existing language models. Another like very similar example actually in design is like agent swarms for example the the hugging face incident like these agent swarms have like a message board that they that they learn to communicate over and this message board in some sense is the shared context that they like act over and RLMs basically say that like the best way to communicate through this is in code like you write the code to do this and it's it's because these models
49:30are so good at writing code um like we want to take advantage of of that And then for for this compositional thing up to and including generating your own harness you know specific for the task.
49:42I think we will start to see that if you go up to uh this figure. Okay. So we we talk about this idea of locally in distribution tasks for a harness and it is like an idea on top of like indistribution tasks when we think about language models. An in distribution task is just a task where like the prompt is something that the model has either seen before or like has seen some version of it. Most harnesses work like the one on the right which is they keep appending
50:12the trajectory as a prompt and so eventually unless you're anthropic or open AI and you train on like kind of these like user trajectories most of these things end up being out of distribution for the most part but locally in distribution is basically the compositional argument of if an RLM breaks down its computation into like a kind of a meta harness of sorts or like a program that involves sub agents that like look at a local problem every individual ual language model call over the course of this task is in
50:41distribution even if the entire task itself is out of distribution and this is a very desirable property I think for for obvious reasons like if every task is in distribution for each individual language model call you will probably get to the right answer um so I mean the the logical limit of RLM is rlms where like you not just you don't just write the harness you also train a custom model you collect data everything like it's like a fully
51:11automated AI researcher inside of your harness.
51:13Yeah, we will see where the training of RLMs goes. I will say as an academic I am not working on this at MIT or at least in the scaled sense because I I can't afford to. Uh but there are companies out there that are working on this. I think Prime and Select is very clearly working on this and it's very cool. Like I'm I'm very excited to see maybe we'll observe I don't know but maybe we'll observe better like post-training scaling walls with when you train around a smart harness. Maybe we'll even see smarter harnesses that
51:43come out and like they they work better around these kinds of principles.
51:46What is a smarter harness? Like that doesn't mean anything. You just said they're all the same.
51:50No, what more of what I mean is like claude code, codeex, pi, etc. are all the same in that like when you break down the the logic of the harness, it's like virtually the same thing.
52:03Yeah. Two calls in a loop, whatever.
52:05But with RLMs and and with other harness abstractions, it looks very different.
52:11And this is where I think you really distinguish. I mean it's it's in the same way that like I think with language model architecture choices a lot of architecture choices end up kind of looking the same when you like scale it out or like it the differences end up being like somewhat minor in terms of I mean for a lab it's not minor but you know maybe one model converges better than the other one like slightly but in general like if you were so for example pre-training scaling walls only hold
52:39because the architecture choices we have are somewhat stable right But if you were to completely change the architecture, pre-training scaling walls probably don't hold or like these kinds of this like power law is going to look very different. And it's like the same thing with harnesses. Like I think all the harnesses we have right now for the most part roughly look the same. But there are some exceptions to this. I think that are that are coming out. I was going to say actually one of the things that I've been more interested by you know talking about PhD students who
53:08take big risks is that people have been people also pursuing the other side which is pre-training scaling laws don't hold if you change data uh right now it's just raw and structured text uh corpus of internet what if you had a better data representation to train on then then yes your scaling law would change as well so there's architecture there's data and you know whatever else uh you can think about well so I just want to get back to to this it's all makes sense. It's very interesting how you sort of recurse up and down the stack from like very
53:37conceptual to like no like well this is this is where we are today but like yeah obviously uh you can can scale up and down I guess um I'm curious how did you start working with Prime? Is Prime taking on more work with this? Is this their answer to Hermes agents? Uh you mentioned Grockot is a little bit different. I just wanted to like name check all these guys and get your thoughts on each.
53:58Yeah. Yeah. Uh, I got involved with Prime uh, after they they released a blog post, by the way, not affiliated with me at all about like how they believed RLMs were kind of the the future. And I I had a friend that was working there, GP Mode Mate, like we got in touch and I think I agreed with a lot of the researchers there and like what they believed about harness design. Like I I I was I was very impressed. I think that like they understood the purpose of
54:28the RLM paper which is not necessarily just to say that like you know we're solving long context tasks but actually like we want more opinionated harness designs.
54:36Yeah. There's always like the result of the paper that you choose to highlight.
54:39Versus the actual point.
54:40Yes. As I I I would love to talk about like the incentives of academia and like the things around like why it's kind of flawed and all the issues and we'll get back to that. Yeah. So, so anyways, um you know, I I I I love the guys at Prime. So, we kind of had been after after we decided to work together, we decided to look into training an RLM and also build this kind of RLM harness and kind of see where we can take it. That
55:08is how like Prime Agent came about. And I think the reception for Prime Agent has been pretty good. Like the one thing I was worried about with Prime Agent is that none of them at least at the time when we were building it, none of the models were that good at doing Arlem stuff. So this was like prefable pre-astra.
55:24I guess I think to take a step back, can you explain what prime agent is, how it's different than a traditional, you know, cloud code, what people would expect harness.
55:32Yes. So prime agent, I think I I mentioned this a little bit earlier.
55:36Yeah, there's is basically uh it it is a a harness on top of pi like pi mono which is pi mono for context is like the like a minimalist yeah like harness. I I use pi as the reference for everything because I think all other harnesses are basically just pi. Um but it is pi except we explicitly restricty to be the only tool that's available to it. every other tool gets loaded in as like a
56:04Python module um or like a like a bash kind of script that it can it can run.
56:08So it uses the core RLM abstraction on top of pi and then it also has this continual harness thing which is uh Seth he's another PhD student. This is a thing that he used to get language model harnesses to play games. Like so he worked a lot with Joel uh who is the like the Gemini plays Pokémon guy. And continual harness is also by the way very simple. I I quite like it. It it basically is this design uh principle
56:37around like what parts of the harness can you let the harness itself modify?
56:42There are certain pieces that like you'll let it modify its own skills, the sub aents available to it. uh what the system prompt to the model is and continual harness is available basically as a tool inside of the IPython kernel and so that's what prime agent is like how what is designed around everything else in prime agent is like standard right yeah I think what is what I I I really liked about it and we got kind of lucky is that like a lot of the new frontier models actually work really
57:11well inside and actually even a lot of the open source models work really well at least some of the newer ones And there is another thing in prime agent I should highlight which is that like we have a very particular agentto agent communication system or like uh like framework which is because RLMs tend to spawn many sub agents. We want a way for sub agents to communicate with maybe the root or with each other. And so there are some design decisions around like
57:39what each sub aent is allowed to talk to. How it does it again everything is in code. So it writes the code to do this kind of communication which I think is really really cool. And then there's I guess persistent sub aents is another thing that was kind of added which is the sub aents. They can last beyond like the standard runtime of of the actual like original agent and you can go into that sub aent you can prompt it more um like you have more visibility and flexibility into what is kind of going on.
58:06This is my number one pain with codeex right now. They their sub agents are just very ephemeral and they they actively discourage you from using it for longunning things.
58:14Yeah. Yeah. Yeah. Which which I think it makes sense. Yeah. I So I mean the trick is just externalize to a file system, right? Like that's the trick. That is the the big trick.
58:23Yeah. And well and also like force everything to run through code uh you know trust the model that can write code and it's going to write its own harnesses itself. So is is Prime going to take on like training post training custom models for this? Is this a one-off collaboration between you guys?
58:37That's it. like what's yeah they are training a model uh internal I think they were they were pretty public about this actually um back in March I mean clearly it is their business yeah so uh yeah I mean they're they're showing that they can train it on their their kind of hosted training stack but no so for model training I'm I'm not involved with them on that the main reason is just I have other things in the PhD I want to work on I think like there are many other big bets to take um outside of just RLMs uh some some I I don't know how much I can share yet uh but in
59:06general Like I think I actually think one of the luxuries of being a PhD student genuinely is that there's so many big bets to take. Most of them will probably yield nothing. But it's a really really exciting time to be in research especially because I think most progress in the field has been a little bit boring. Like I'm not saying the the outcome is boring but the the process of of doing these things tends to be quite boring. Um and so there is kind of this question of like you know what what do
59:34we want to do next? Um but we can talk about that later.
59:37Yeah. Okay. I want to close out a little bit more of your research and then we can we can uh spread it and out uh since you release RLM a lot of excitement about it. Any secondary third party work that you want to shout out as like that you guys should take a look at this.
59:50Oh yeah yeah yeah. So, uh, Harvey, the legal AI company, released a blog post, not affiliated at all, but they postrained an RLM on their like legal work, which often involves a lot of like sifting through documents and kind of looking through like a variety of uh, specific information that maybe is not so easy to retrieve with like a pure retrieval system. And they show like really really good results. It's very exciting. I mean I I I was I was shocked
1:00:19that that that they worked on this. They did not tell me. So when this came out I was like oh that's awesome. So there's this one I think is super super cool and and what they're doing there. I think this is a collaboration with base 10 by the way as well.
1:00:30This is B Headlong which is law the Lud Institutees kind of uh it is their like persistently running harness. Um it's very cool.
1:00:39Oh they renamed it. They used to call it something else.
1:00:41It was like auto terminus ter Yeah. I I know they they they've they've gone through Yeah. Um, so this is Andy Kwinsk's big project. Um, it's super super cool. I love Andy. I don't want to downplay what they're doing because they're using the RLM abstraction, but they're doing something much cooler than the RLM. Uh, which is like they they have a system that kind of what they call like thinks persistently. So even when you don't query it, it has a way to
1:01:10um think through problems that it has in its context.
1:01:13Oh, is this just like always on type?
1:01:14It's like an always on thing, but it's like not that expensive. Like they they they control the the token cost um to make sure it's not like burning through all your all your credits. This is super cool. I'm trying to think there there are many actually if you go to the RLM uh GitHub page, there's a bunch of things I've linked uh below. There's a ton of really really cool kind of things that people have been doing. Axe is another really cool one that I think it's just by this one guy. It's like a
1:01:42harness or on DSP and RLMs. DS PI also has an RLM. Oh, the last thing I'll shout out is on ARC AGI 3 I believe there were a lot of harnesses on their like Kaggle competition like the official one, not not the like public prime select one that or like what people have evaluated on. they all like claim to use or they reference like like Tufa for example some form or some inspired form of the RLM abstraction in their harness which is really really
1:02:11cool. I think it's um this is where I mean this is exactly the setting where you would see a lot of benefits from composition and and and using code and combining like neurosymbolic systems with with AI and so yeah very very cool we love a good neuro symbolic reference.
1:02:26Yeah. Uh you also uh mentioning RKGI3, OpenAI comes out and says we're 99.9% on this. Uh they also say we we solve Nevia Stokes. We just throw a model at it.
1:02:36Uh there's some debate around whether or not it's just model.
1:02:41Are they using an RLM? Do you know? I I mean I I I would guess probably not unless you say like be I'll be careful here because you know people debate what is an RLM, what is not an RLM. It's somewhat clear that what they used is some kind of swarm of agents with a shared some shared context like some shared file system. And like this is very much in the spirit of RLM stuff, but I think there's a there's a lot of like, you know, more clever things that they did that's not maybe related to to
1:03:10the RLM itself. I agree a little bit with the idea that like a harness is not that necessary for what they did. The way that I would put this is that I think a model like a GPC6 Astra type thing is technically smart enough conditioned on the right information to come up with a proof for these very very difficult problems. Now, how you get to
1:03:38that information is a giant question mark. And in in their case, it probably came down to like a very long search over like many of these sub a or many of these like agents in the swarm and and maybe also like researchers. I'm I'm actually not sure about this part, but putting in like their kind of intuition as to like what you should explore and things like this. And ultimately like this produced some information that some agent was able to take to finish the
1:04:06proof. And so in that sense like I think was the harness that important? No. And I I think what this is pointing at is like these specific details of a harness do not really matter. And I think that's also what that uh what the harness tax paper is is pointing at which is that like beyond the user's feeling of the harness realistically all that matters is just like how are you composing these agents in a meaningful way to get to the final answer. And maybe that's what like
1:04:34swarms and all these things are are are really about. And so for my POV at least, if we if we start to think about like for user use cases, what do we want out of harnesses and and and things like this? Like we want to take the good parts out of these like the cloud codes, the codeexes, like the stream that people like to see, but like under the hood, whatever is running can be some really weird complicated swarm of agents that like ultimately come up with an answer. The user doesn't want to see that though obviously right like it's
1:05:04it's not legible information and so I you know this was another kind of thing in the spirit of RLM's like recursive language model it sounds like it's a language model and but it's not a language model architecture but the reason for this is like I think we will start to see in the future probably one day and I I I wrote a blog about this what we think of as a language model like the thing that we query might actually be like a swarm or like a scaffold or like some weird harness
1:05:33design that scales very well, but the user just doesn't see it. Ultimately, all the user sees is some front end version of this harness. And yeah, I I I I I think it's a relatively safe bet at least to make that this is what we will see.
1:05:45Um, and this maybe goes back to the limitations of the base transformer. Like, you know, obviously if you just took a base transformer and you said like solve Navier or Stokes or something, it's not going to do it.
1:05:56Like, yeah, we we all know this is not what's going to happen. But yeah, I think this is maybe the more interesting part uh and may maybe the claims around like you know did the hardness matter is is more around this of like just arbitrarily pointing models like or agents at a growing kind of context of information maybe is just enough to solve very difficult problems um that I can buy. Uh while there's a lot in there I do want to say uh the the amount that open spent is semi-public. It's 10,000
1:06:23agents in 88 hours. uh 130 billion tok out tokens which is estimated to be about $40 million in public pricing.
1:06:31Surprisingly actually like less than I thought.
1:06:33Yeah, not that much.
1:06:35130 billion output for the final but as you said there was a lot of context being passed around. It's more than double that in just the total agent messages being sent.
1:06:45I think uh one thing I was honored to bring up also was cursor as far as like swarm stuff is concerned. Uh this is slightly older uh meaning February which is ancient. Uh but if you scroll down all the way to the the final sort of multi- aent architecture that they arrived at, it was basically an org chart of a normal software team. One thing I'm thinking about because I I basically there's like this um uh gather all function that you have to do with with sub agents or you know it's very similar to GPU programming actually.
1:07:14Uh and and so uh that's a bottleneck.
1:07:17This is a bottleneck. there's one main guy uh it's coordinating all the sub guys then they have to like re you know gather again and then and then reocoordinate that's slow that's that's crappy uh what a true swarm should be is everyone is just their own person yeah yeah well I agree with this and I I think that there is a question to be had though let me give an analogy which is like when would you use compaction and when would you use an RLM and there are many settings where an RLM can solve
1:07:46maybe a more difficult task than compaction can, but you would prefer to use compaction in most cases because it's cheaper and it's quicker. And I think in in the in the context of agent swarms, there is a similar thing going on of like I'm fairly certain that like 95% of the swarm is entirely useless or like what it's exploring is entirely you're just burning tokens versus in this setup maybe not so much. I'm not sure. I mean, maybe it's also the case here.
1:08:14Everyone has a job. There's a Jira board. I think at some level that's how problems are framed, right? So like if you have a search problem and you're spanning out a bunch of sub agents to do search, there's going to be a lot of useless information, right? There's one retrieval answer that you're getting and you're spanning off. But that is consciously understood, right?
1:08:33Yes. But there is kind of this question of like what is appropriate to solve for which problems like what what design you know in theory open AAI can use can can can package up this API and they'll call it swarms and then they'll give it to you and they'll be like point this at any problem and we'll give you a solution but maybe it's called pro.
1:08:51Yeah. May maybe you'll have to pay like $40 million to get a result and it's like well but it's exciting. I I I will say like it is it's very exciting that we even have the option to point $40 million at a problem and solve it.
1:09:04But there is there is still kind of uh a lot of research to be done in this area around like you know what is necessary like what what do we want to do? What design do we want I mean we probably don't want everything to be a swarm but like where do we draw the line like can the agent design that or decide that for itself etc etc. Yeah.
1:09:23Uh, and then, uh, just a quick check in case you have opinions on this. Have you looked much into open-endedness as a general category of problems?
1:09:32Meaning no prompts, just go a little bit. Um, so I I was at Sakana for a summer uh, right after or I guess right before my PhD and and that's something that they work on a lot there. And I think there's a lot of people even at like recursive super intel. There's many of them now.
1:09:51Oh yes. Yeah. So, so like Richard's company and then and then and then also actually wait that might be it might be the same company. I I don't remember. Is is Tim Rockell also? Okay.
1:10:00He's the main co-founder. He's the head of open-endedness for Google.
1:10:04So, I think with open-endedness problems like I view them as somewhat similar to even like unsolved math problems. Maybe that's a weird way of of framing it, but like I think a lot of the techniques in terms of like how people like approach them are kind of the same. Like evolutionary search is like very similar to just launching swarms of agents and hoping that like they come up with like an interesting I mean this is what what like Alpha Evolve and some of these
1:10:33other works did uh like a year or two ago. But I think what maybe is is not uh and and I'm not sure if this is what you were alluding to, but I think what's not super clear in open-endedness style search problems is do we frame all of them as basically like an unsolved very difficult problem or like where the objective is clear. If that's not the case, I still don't know yet entirely
1:11:03what the value of it is. Maybe maybe you have other opinions like I I I don't have too many opinions on this but at least from my time at Sakcon like I I got the sense that like we ultimately still kind of want to approach things the way that like say OpenAI approached Navier or Stokes. We want we there's still a lot of nudging in certain directions that we want to have to like get to the point where we have something interesting.
1:11:26My version of it is like u maybe it's a split between basic science and applied science. Basic science you're researching for researching sake. you just want to understand things better.
1:11:35Uh I have no idea if like there will be any application at all whatsoever. But that I mean that u and then applied you have a goal. You're trying to minimize loss in some way, you know? Uh and so uh what I really uh think you know in terms of like the the big bets that people have uh what if there was no prompt like I see you know you just pick like you just like you just in this like swarm of things and you're like hey what's up guys? I like what you guys working on and like you just decide I see this is an interesting problem. The
1:12:03biggest issue that like and may maybe there's a clear path to this, but when I was there at least, the biggest issue that people had with open-endedness was like how do you ultimately choose at the end? How do you pick out the interesting stuff? Because when there is no like maybe the agent comes up with a goal but in a lot of cases like what they and they have something called fugu I think which is like a it's like a model router type thing that was like inspired at least by this idea of like let's pick a problem where maybe we can pick out the
1:12:33best solutions to something. In this case it's like pick the best model for this problem. Um you're the first person to connect motoral routing to open-endedness. No, no. Yeah. But so I I I bring this up because I think like with open-endedness, like just generally the issue is like when we have this giant corpus of like slop, like how do we sift through and find like the hidden gems and like the solution to open-endedness really just letting models run forever
1:13:02and like finding like just doing data gen super super super high throughput data generation. like a meaningful maybe that's cool.
1:13:18To me it's like very interesting as a counter to basically all of machine learning where you have a goal uh to have no goal.
1:13:26But or or like an illdefined goal that you're like well what what about this goal? And you're like well okay maybe. And then you like sort of research more and you find that that is an interesting goal cuz I think like finding your objective function like like you said like Jeb found an objective function that was interesting that no one was exploring.
1:13:40Y I think that is like you know similar to your message about grad students as well like you know you stay in school because you are you want to pursue open-endedness. If you want to you know uh profit max and like you know join the escape the permanent underclass then you join lab right?
1:13:58It's funny. I I feel like I I don't I don't hear this discourse a lot. I mean, I'm I'm in the East Coast, so it's like a very different type of But then when I whenever I come here, it's like that's always like the topic of discussion.
1:14:08You cannot pay rent without doing this.
1:14:12You're getting priced out, guys.
1:14:14Okay. So, so yeah, there's there's all that. Um I don't know if you want to if it's relevant enough to talk about the mismanaged geniuses, which you were you were pulling up.
1:14:21No, no, it's just one of your blogs, but I will poke on uh Yeah. Yeah. So they they did in their um blog post I I talked to them about this as well. So one of the cool results of Ultra is they basically just let it loose on automated data science research uh with little to no human intervention.
1:14:39It's just kind of making Yeah. So this is auto research which is a little bit more open-ended and there's degrees of open-endedness and I agree with that.
1:14:47Yeah. Separate than auto research with objective this is just do stuff. But you know it's cool. They're they're working on it for those that are interested.
1:14:55While you're bringing it up, actually uh what is your take on Sakana? Like what are they doing apart from being, you know, the Japan one?
1:15:01Yeah. Yeah. Yeah. Um I mean I actually love the people there. Like I I think they have a really really smart team. Uh and it makes sense. I mean it branched off from from like an earlier team at GDM which was also kind of I guess doing this kind of like open-ended evolutionary research style stuff. What I liked about my experience there at least was that they did have that like mishap back I forget at this point when but I think about scientists.
1:15:29Uh no the GPU kernel that went up.
1:15:32Yeah. People cannot forgive them for that.
1:15:34Yeah. Yeah. And I guess like the AI scientists like I mean there's there's some criticisms of it that I I mean I don't I don't work on that so I I have no no kind of take on it but I think in general like what I like about them at least is they're a little bit more of a researchy type lab. So like they don't operate in the same space as like open AI or anthropic like for sure like definitely no they do not I mean it's pretty obvious probably that at least when I was there they do not have a big competitor model or something that like you know that that everyone is is is
1:16:02using but I think they kind of operate in some ways as like a PhD lab which is cool like and I I I think like David Ha is like he's he's really really smart like I think he has a good sense of like also I think the market in Japan is also a little bit different for AI and like who they're targeting is is slightly different than maybe what we're used to here. But yeah, I like that they take kind of a lot of their research is kind of weird I think when when people view it, you know, and I I like that. Like I think it's
1:16:31you should have more weirdness. Yes.
1:16:33Uh and you say different market just is it like enterprise like the way it works there is is a bit different like how how deals happen and stuff like that. They do have a I mean I guess this page is originally in Japanese but they do have a model specialized for the Japanese.
1:16:47You didn't know that?
1:16:48I didn't know that you did. I also have personal friends and know the team. So there is a even from the sense of a way that you speak culturally uh you know responses are tuned towards that. This is not like it's frontier on benchmarks. It is a cultural appropriate model for them and they have like chat and all that. But uh to to mirror your point, you know, there's also like a how should education look like and someone wants to work on it and they're very PhD lab of do your
1:17:17thing, why not? We have money, go research.
1:17:19Oh, they say it's a Kimmy fine, too.
1:17:20Oh, there you go. Good for them.
1:17:23Um yeah, I mean, and you know, speaking of Kimmy, right, like another uh you know, just a a grad student that split out and like I just have this like Kimmy Delta attention that wants that I want to work on uh and like somehow managed to make one shot.
1:17:36Don't understand it still.
1:17:38Yeah. Yeah. Well, he's he's super cracked. Um, at least my my understanding. I I think in general like a lot of a lot of the Chinese labs have have done really really cool work.
1:17:48Yeah. Like any thoughts on Kimmy agent swarms?
1:17:51Yes. Um, one thing I will say is whatever OpenAI is doing with their agent swarm is like clearly the right thing to do. You have to kind of think about it this way, like nothing, especially without like a very very smart harness design, which I I don't I don't think anyone really has so far, nothing comes so easily for free. For example, like the agent swarm design is
1:18:20not something you can just take for granted. Like it's not like GPT6 Astra is just super smart and then it just got agent swarms running well. They clearly trained I mean the the hugging face incident was them training a system to be like a swarm. Uh and I think like clearly they've done something really well to the point where you can throw $40 million and solve an unsolved problem. And I I think with the Kimmy agent swarm thing like at least from when I read it just came off as like
1:18:50this is interesting but I don't actually know whether or not this can solve anything novel.
1:18:56Yeah. They just they were like it makes spreadsheets.
1:18:57They kind of were just like yeah like you know here is a here's a swarm that like kind of does stuff and it's cool.
1:19:04Um but and I I will say same thing with dynamic workflows. I actually kind of think the dynamic workflows release was sort of a flop. I I don't know how how you guys feel about it but I think it's like uh my understanding is like it's not used that often or like it's just very expensive.
1:19:18It's too expensive.
1:19:19It's ultra code. It's basically like like you know take over my bank. doesn't act. I've tried it and like it doesn't act in the way that like again I'm not I'm not the biggest open AI like you know stand or something but I think whatever they did was very impressive like they somehow managed to get a way for this this swarm to actuallywards towards a goal and yeah it's very difficult.
1:19:41I see. So so like efficiency of of the multi- aent swarm is the objective function here right like like how much of this is sloth like this is a lot of slop. Yeah, I think so. I think we we take for granted what it means for a swarm to converge to an answer. It's just like not not something we can take for granted.
1:19:59Yeah. Yeah. We've we've done one pod with Nome Brown and like his death that thing his whole thing was like okay like we've worked on a lot of like competitive agents. Uh we're working on collaborative agents and like that's now called a swarm.
1:20:11Yeah. I would do a quick poke. Do you have any thoughts on Gemini? They were also IMO gold. like, you know, there there was a time where they were getting agents to reason for a long time and um yeah, I mean, is it too old to think about?
1:20:25No, no, no. I I I think it's a little bit blown out of proportion like Gemini for again, so I should preface by saying I haven't worked at any of these places, so so take this with it with a grain of salt, right?
1:20:36I mean, you have strong opinions on agent harnesses and you know, this was But um let me just say this first. This work was really impressive. I think what what they showed here was like they took a a time when the models weren't that good. Yes. And they managed to be very very smart about like what the harness does. I remember for this at least like yeah like alpha geometry. I guess that was a year before this. But it was very cool. They took um they took it to the max and they like designed uh I don't
1:21:03know. It's I I I'm not too big on competitive math. I think like GDM it's sort of a shame like everyone I've talked to about GDM kind of has the same opinion which is that it's way too like bureaucratic whatever is go they have the talent and the resources to do almost anything but like I don't know until they figure that part out like nothing against anti-gravity for example but like I don't know anybody that uses anti-gravity so I've tried it once and it's I don't see a reason to switch to
1:21:32it um and I I think for whatever reason like they've been struggling with this so Yeah. Well, uh, you know, a lot of people dogged on meta for a long time until they started coming out and like I think Google's going through that phase right now. Yeah.
1:21:44Um, and you know, it's it's you got to stay alive and I want to focus back on just like your uh your thoughts just general. Uh, we can talk about speculative PCC uh mismanaged geniuses or just like throw away all this and just talk about whatever else.
1:22:00Okay, let's let's talk about mismanaged genius for a little bit. The only comment I'll say on speculative PTC is that it's a really simple idea. It's almost like obvious that this should be done and like there's not much more to talk about it. Like I think it's just like you should just use it for for like coding like anything with programmatic agent calling like RLMs or code act like yeah it's like it's like a no-brainer.
1:22:22What's the you know for for people that haven't read it, what's the oneliner? Uh the simple thing is when the model is uh writing its code or like even as like after it finishes writing the code a lot of tools tend to be like sequential or like you have to wait on them. So you should just launch them in advance like if if you're able to compile this code you can probably figure out like even though it's like kind of variables and stuff like you can figure out like statically analyze
1:22:52Yeah. There's there is this uh someone pointed to me some actually like academics have especially PL like programming languages people have some like very kind of cool ways of doing this and so like you know at some point maybe I'll I'll I'll like work on this mostly you have to change language because if you are in JavaScript python you can't do this yes yeah yeah so like uh hasll yes um um what's what's the uh what's the normal one that's that's not lisp uh lisp uh or camel
1:23:20oh camel any functional language you can actually pipeline this um so effect ts if you want to do typescript okay we can switch over to uh diagram which which is like your you guys' whole thesis right like that uh I mean to me this is like kind of like a restatement but maybe I'm missing something of like well work on better harnesses or like you know your your models actually are capable a lot more if you try harder so this is a skill issue.
1:23:45Yeah. Yeah. Yeah. Basically, I I I mean, I I I think there's one thing I want to see. I I appreciate that there's a big focus on like jagged intelligence because it paints a big picture of like we can do this if we really set our minds on it, but I kind of wish and and maybe someone in academia should do this. like really just sit down and think about like if I took Astra even the current frontier models are not good
1:24:13enough at like doing a particular job over let's say the span of a month consistently and well and I think this is like a stupid problem like I I I genuinely think we can solve this you don't need to be a Frontier Lab and like do all this like you know fancy stuff for your IPO like I think like these models are so smart that even if it's like a silly way I think that it genuinely is a skill issue of you can get a model to be as good as let's say
1:24:41like just some 18-year-old high school kid doing some job. I think it's like ridiculous that we can't do that. And it's part of the reason is like the the the format of a language model is is not really amendable to that. But I think you can shape a harness around it and do it. And I think like this in itself is like ignoring the RLM stuff, ignoring all the like you know what abstractions should we use like I just think someone can design a harness that can do this like I I I mean that's maybe you say do
1:25:12Do long running but simple tasks uh and and do them reliably.
1:25:17That's an example. So, is this different than like, you know, pick your favorite company, Harvey, for example, using LLMs to do legal work or what's the I mean, I I I guess it's kind of like that except if the bottleneck was not like certain legal knowledge or something like I I don't know. Let's say I mean, I guess like the examples, you know, people can take models and build pipelines or whatever and have an agent repeatedly do whatever tasks they want, right?
1:25:43Yeah. Yeah. So, so for example, like if I wanted a general system that I could kind of talk I can talk to it like I would talk to an intern and basically just ask it to do to explore some small thing. So maybe an example of this is like very silly auto research is is maybe an example of this of like not necessarily finding super super novel solutions but at least optimizing all of the easy parts of any problem. they
1:26:13often end up being overindexed for like ML training and and and things like that. But yeah, I don't know. May maybe maybe that's that's that's not like super super clear, but there is a lot of people's general workflows where you probably could just vibe code up some specific harness to help you do like automate this thing. Some examples are like automating finding like research papers and stuff like that. But usually
1:26:42people will design like a specialized agent to help them do this kind of thing or like like they'll vibe code a harness like and just run it or like their Slackbot or something. But I almost think there's just like a standard form like just a harness that you just plug in like you don't it doesn't need to be designed for finding papers or like you just kind of tell it find this for me and you like plug it into that setting.
1:27:04What I'm getting at is that I think there's a lot of easy things that can be automated and is this like a hypothesis or a point around like capability overhang like even if we paused there's still a lot of impact to be had with current state of models in some sense yes like I guess what I'm presenting is the easiest form of this but what this is kind of saying is that like we have jagged intelligence on a lot of things like for example models
1:27:34are like disproportionately good at coding and math. This is saying like we can translate those abilities to many other things. So like for example, if you took usually if you take someone who was an IMO gold or something and you kind of apply them to a lot of different problem solving domains, they can figure it out. I don't actually know if this applies to models. For example, like in GPU code optimization, one very interesting question is whe if you were
1:28:02to take out all of the GPU programming data like from a model, but you it was re it was like as good as Astra is now just without like with that taken out, would it be able to still optimize GPU kernels? uh like would it be able to learn in context roughly what it needs to learn and then like have some pipeline or like or come up with some solution to solving like optimization tasks and I think like there's like a mismatch between like if you took a human that was as smart or like knew as much as Astra there is a mismatch
1:28:31between what that human can do and what Astra can do maybe around a harness um and I think we can actually approximate the human a lot more to me it sounds very approximate to the continual learning problem uh I think you're making that's the best example Yes.
1:28:45Why didn't you just say that?
1:28:46Yeah. Yeah. I guess I I was like thinking like could I just blurt out some words like I'm I'm careful with that. But yes, I I you have a very like maybe I maybe like a bit of a I don't know like so I I think you're a very uh I don't know. I know your undergrad actually. Are you like a math personally? Yeah. I think study math.
1:29:08Yeah, that's what I wanted to do at least. like a category theory type of uh abstraction where you think in categories and then you have to like then translate down to the specific and but then you like actually really care more about the category.
1:29:21Uh and like that that's the communication error because like everyone's listening for the specific but actually trying to also you know convey the general.
1:29:31I don't really know. uh you you can you maybe use like a shorthand of like okay I'm at level two and then I'm going to go up to level three and come back to level two that we should have say some like epistemic like shorthand for like this kind of thing because yeah it it's hard like you're compressing a lot into word after like sequential word decoding um should we convert to neuroles you know is there like a a better form of output than English or
1:29:58uh python or javascript I don't know this is very kind of like a post but like people have speculated about like what is the native language that people want out uh that models want to output uh some people say binary that's a that's Mark and Jason's thing I don't know yeah ptx let's say a mix of English and Python and I only say this because the capability of a model is somewhat a reflection of what we train them on so
1:30:28we still want like like yeah I don't really buy the the binary argument. I guess I like I understand but it's like uh like yeah you you want to you want to model the world in some way. I think uh you know uh one thing I I'll bring up is always which I always do in this kind of conversation is superior warf uh which is you if you choose English you will have locked into however long English has been around which is let's say 500 years uh which is not that long like actually uh you know like what you the language you speak constrains how you think and if you learn a different
1:30:57language uh for example uh someone uh in Chinese we don't have tenses um I don't know if you I I actually didn't know that and I speak Chinese oh I I I did know that but my my Chinese is Not great.
1:31:09Yeah. Yeah. Or like uh in let's say in Japanese or I know h I forget what language it is like in in Korean everyone you speak to you have to like acknowledge social status but there's a different dimension than gender right like it just like it just influences everything you do. Um when I take ling 101 um apparently there's a there's a there's a language in Africa where like there's a vegetable gender.
1:31:32Uh yeah right like just like you have or like smo is no word for snow or whatever like anyway. So like the the language that you adopt affects your thinking and if you choose to output your train of thought in English you are biasing towards whatever English solves.
1:31:45I don't know what the the sort of prior of English is.
1:31:48That's interesting. I I did not think of it that way.
1:31:51I mean at some level it's interesting right? So you're right on language at a lot of model chain of thought also fluctuates language. Uh the obvious example is Chinese models speaking in English might still reason in Chinese.
1:32:05uh but at the same level most models are very capable multilingually and that adaptation we can see you know you can add in languages you don't get that much for maning a whole language but they'll reason interchange yeah so we're all all regressive I mean but also like let's say German like you know uh subject object uh agreements uh you have to put the the verb at the end which is very super annoying like very famously uh yeah right you don't know what you're doing until the end where you're like oh that mess of nouns and
1:32:33then the verb Well, the the the most classic one that most people be have heard of is arrival where um they have the heptopods where they they think uh the time is like flat to them. So they think in they they they they output entire sentences at one shot. Um so it's it's closest to like the difference between auto regression and diffusion.
1:32:52We talk in auto regression. What if you could talk in diffusion?
1:32:55Uh where things just resolve over time.
1:32:58But the whole idea shows up at once.
1:33:00Ah I see. So that is a drastically different uh language, but it is a language.
1:33:07I see. Oh, that's really interesting.
1:33:08That's really really interesting.
1:33:09Uh- which like machines could speak that we, you know, probably would never speak, but like yeah, machines don't care.
1:33:15Maybe this is a huge tangent, but are there are there not like things inherently that are reasoning chains that are inherently auto reggressive?
1:33:26Sometimes like Yeah. or like Yeah. Yeah.
1:33:29Yeah. Something happens first then something like Yeah. Anything in code for example like has to be causal in in in some or usually at least has to be causal. Um well no uh so so then you have to then you're not exploring enough uh function programming language theory uh where uh everything is like pure functional and like completely relational and and u you sort of abstract away the solver that translates the relationships that is are always true into code. So I I yeah I I feel like this is maybe a little bit too out of my depth.
1:33:58Uh but I love languages uh whether it's coding or human and I do think a lot about how that affects reasoning and and the boundaries of what we can do.
1:34:08I don't need to go too much beyond that. I don't know if you have any other thoughts. Uh my closing question was going to be uh you have all these research uh directions that you want to do. you're you you know you had a GPU mode phase, you had a uh RLM's phase.
1:34:22Presumably, you have other stuff planned which is why you're not um doubling down on that. By the way, I noticed that it is interesting how you guys do start with the GPU side and then you migrate towards the zero gradient side, which is what uh Shenu called it. Doesn't that feel less legit than messing with GPUs?
1:34:41Yeah, I guess in in the sense that like so you did bring up that like I I like to think about things in like a math oriented and it's like very uncomfortable sometimes to be working on like harnesses and agents because it's so because you think all harnesses are the same like two new ideas in harnesses.
1:34:58Yeah. It's it's it's also just like like empirically it's hard to like verify a lot of findings at least with with the compute that we have available to us. But the reason why I think I've moved on to a lot of these problems is I think actually this is where most of like the innovation is yet to happen to me like the GPU level is a means to exploring other ideas like you you you you want to for example like get good at writing
1:35:27kernels or like even automate writing kernels for the sake of a broader goal of like I want to explore ideas where I'm not bottlenecked by systems challenges in that sense like you know I guess a lot of what's written there is all harness stuff, but I am also interested in things at the model level as well. Um, but I'll just leave it at that.
1:35:45Okay, that's a good hint. Um, anything any if people want to reach out to you, uh, what are you looking for help on?
1:35:51What do you want collaborators on? Um, any sort of calls to action? Yeah. So, um I guess there's nothing I have in particular where I feel like I need to work with someone on unless it's like unless it's with a company before like compute or like with to talk with other people about it. But I will say I'm not I'm never opposed to working on ideas with other people. I get reached out to a lot by often undergrads or or even like other students. Podcasters.
1:36:22Um, and usually I feel like I get an email that's something along the lines of like, I really like RLM, like I want to work together. Uh, and I I feel like I Yeah, that's a bad reach out, right? The worst is like, can I pick your brain?
1:36:33I'm like on what like read my paper, dude. Like like they'll be like, I read your paper in quotes like recursive language models or like prime agent like a self-improving RLM harness or something.
1:36:43Uh, and it's kind of like I I I like I really like people that are opinionated even if we disagree. I think if you have strong opinions and are are able to like think through why you think those opinions are are right or wrong cuz usually it's it's hard to to to to actually tell but like you have strong convictions about certain problems like I'm I'm always happy to like chat and even like potentially work on something together. I have like no limit to who or like what I would like to work on. So um no limit.
1:37:10Yeah. Yeah. Yeah. I I you know in in the era of agents I think there's a lot more work you can do you know like like bandwidth wise so I yeah I I think in general like I am not hard to impress but I think it just takes a little bit of effort to kind of yeah know know what you want it's very clear I mean and like when you see a new thing come out well executed good simple idea uh then then get that immediately gets your attention right like it's actually like not that hard to
1:37:40get the the same attention all the Frontier Lab guys because they are looking for you. Uh you just have to put yourself out there, right?
1:37:46But yeah, it's true. I I will say I think I think human attention is very scarce right now and I do struggle with like the number of projects I have going on and uh I don't know how to manage it. I don't think agents are helping at all. Like I would just prompt it and I'll prompt a thing and then never look at it.
1:38:00Right. Like which is very common.
1:38:01Yeah. And that sucks.
1:38:02I guess maybe the one of the smaller differences in like I mean actually maybe you were doing research. I'm not sure. But for for me at least, like I I'll have maybe like 10 or 15 different ideas that I want to do, but the thing is like most of them are bad.
1:38:16And and and this also maybe is true even for someone that reaches out to me. Like maybe the idea is actually bad, but it it looks interesting to me. And so like we can spend like some time looking into it and if like we feel like there's actually something there, like then we should we should take the next few weeks and just really pursue it. And like this is my my style with this is why I love the PhD by the way because there are times when I'm just thinking about problems like maybe on a run or just like playing tennis or something like I'm not I'm not working I guess but it's like those are the most fun times and
1:38:45then when I like really am convicted about something I'll just like drop everything and just do it like just spend like all my time thinking and working on that problem and then I you know once you get to the point where like you can just run experiments then it's it's it's kind of easy coasting again. So yeah sorry this is like the the fourth last question which is like uh I I think a lot of people are also thinking about science as the next frontier like uh like physical sciences, bio, math even um
1:39:13how do you separate like I guess let's say your choice of projects that is applicable for industry and then maybe choice project that's just like science.
1:39:24I actually worked on like AI for bio stuff before I started my PhD. The field has changed a lot since then. I I I should say because it's it's like it used to be a theoretical like of course what what do you mean like you know I have one path and then I chose that but now a lot of people are crossing over and like so we have started a science pod to just cover those things because a lot of engineers are like well actually there's tractable problems there.
1:39:46Yeah. You know, I will preface by saying my my understanding of a lot of these topics is probably pretty limited, but I think like if I find out either because someone reaches out or like I look at a problem and I'm like, hey, like some design principles that we use or that we're thinking about right now actually make a lot of sense for this problem. I get excited about those as well. But I think I think it's harder. I I I don't know. I think with I think science especially like empirical or like
1:40:12applied science has very very long like what is it called like feedback loops or whatever. Um yeah you convert this to a robotics question.
1:40:22Yeah. Yeah. I mean to me like also this aspect of like what is worth spending and betting my time on now because like maybe I spend a lot of time on this problem and then like in 6 months like a different solution kind like maybe a new model comes out and it's like ah it's way better for this and so I do have to be careful I like you know you have to be conscious about like where you think things might be going. Um so publish publish cycle.
1:40:46Yeah. Yeah. Yeah. So I can just Oh 3 it's saturated. We we did it.
1:40:51Yeah. I mean, ArcJ 3 got saturated in less than a year. So, it's kind of, you know, like it's I don't know like if if if you were a lab picking that problem like you're probably kind of sad now cuz Exactly. That's why knowledge work, gaming, all these things are saturated. Now, actually the frontier is science.
1:41:06Knowledge work is saturated.
1:41:09Yeah. GDP is like 80 something 90 something. Um, I mean like you know there's there's 90 to 100% that's obviously going to take the next 10 years but like uh well the next lowhanging fruit is going to be like you know the other stuff.
1:41:23Yeah, maybe I'll think about that more.
1:41:24I actually I I haven't given too much I'm just trying to guess your next direction. I I will say because I think I think especially at MIT like it's it's there's a lot of really talented scientists there like in in the natural sciences and I think it's it's a little bit like sacrilegious almost to be like I'm going to figure out like your problem like no that's no I know yeah so I interviewed who did the IMO thing he's never I've never been to IMO I don't even know what it is model dude yeah yeah which is like very
1:41:54disrespectful but like whatever at some point like you have to respect ect like okay you know the the progress is being made like number is getting output right true yeah I mean that is very bitter lesson I it's very interesting like cuz the mathematicians are responding this way to nar right now like Terry turns to is like no like let's not let's not use AI I'm like well yeah I mean I I think that that whole thing is kind of weird because I feel like I feel like they would have had a stronger case if a lot of them
1:42:23didn't work with openai before like all this happens no that's ad hominemum uh and they they're really trying to stay away from that then so what right so what like I don't know like like so what they got you know they they they've collaborated I'll collaborate with people that I don't agree with or whatever or like I did a thing and then now I now I regret that I changed my mind or whatever right so I'll defend their right to say that but like uh yeah a lot of people are reasonably disagreeing with uh with them um okay cool uh thanks for you're joining us congrats on uh your success
1:42:52so far I can't believe you're still not done with your PhD year two you know can't believe we did this podcast without going through the RL paper.
1:43:01He he had a definition.
1:43:03I think the paper is more about like empirical results. Like the actual idea is quite simple.
1:43:07Yeah. And you've talked about it many times.
1:43:09Yeah. Yeah. Yeah. At this point, I think there's there's more interesting things to to to look over now. So cool. Well, we're excited to see what you do next.
1:43:16Thank you so much.