0:00effectively the the the end game of something like interface world models is you have a fully neural operating system. So I think uh uh Andre Karopathic has written about that quite quite quite a while back but it's uh you know I think to me it's it's a bit um it's a bit odd that you know we have for example with an interaction with with an LM of today you have this LM that can basically talk talk to you
0:28about anything it can you can take the conversation any direction you can it can it's very general so it can solve all those different tasks but you interact with it through a very rigid interface.
0:38And so to me it's just a matter of time before the interface itself becomes learnable and becomes you know part of the the the whole loop of like you're not just delivering you're delivering an application end to end and that means you're delivering the language model but you're also delivering the the render and the pixels and and that's also a learnable component. Before we get into today's episode I just have a small message for listeners.
1:03Thank you. We would not be able to bring you the AI engineering, science, and entertainment content that you so clearly want if you didn't choose to also click in and tune into our content. We've been approached by sponsors on an almost daily basis. But fortunately, enough of you actually subscribe to us to keep all this sustainable without ads, and we want to keep it that way.
1:22But I just have one favor to ask all of you. The single most powerful, completely free thing you can do is to click that subscribe button. It's the only thing I'll ever ask of you. And it means absolutely everything to me and my team that works so hard to bring the Inspace to you each and every week. If you do it, I promise you we'll never stop working to make the show even better. Now, let's get into it.
1:47Okay, we're here with Anastasis from Runway with me and Vivu in the studio.
1:53Congrats on all your success and progress with Runway. You're opening offices all over the world. Did you envision this when you first started out?
2:00Uh not quite. I think I think even when we started with we had this idea that you know it was more a matter of when not if uh we were seeing the early generative models of 2016 2017 and just extrapolating assuming you know we resolution quality increases predictably over time there's going to be a point where most of content will be generated.
2:23Um and that was maybe the initial the initial thesis of runway was we will need as a result of those generative models rethink how creative tools are made. Um and as we built out the research behind um our generative models it then became clear that they were useful far beyond that as well and it is more obvious now with like the real world stuff and the world models that we'll talk about later. I I'm just kind of curious how you go from a background in like ZDOC and you know
2:50computer vision into runway like take us back to that early conversations with Chris and you know whoever else is on your founding team. Uh I was always split into those two worlds. One was the you know I had my own art practice. I was making a lot of interactive art. um I think for a long time uh and and then on the other side um I was working in startups and I was working as ML engineer as a backend engineer at different different companies. I've always been interested in uh coding and
3:20computation and especially interested in simulation and brought it back into my early artwork as well and at the same time I was interested personal site has has a few right um is there is there one we should pull up just just in case there's something that's like I just like to go down memory lane okay what is this so this was a project that I made I think back in 2015 where I I I built this software that would give uh voice instructions to people in a gallery space. So it would basically coordinate
3:48interactions between people and so it will first give you an identity like you're an architect, you're 30 years old and uh you like sports. Um and then it would match you with another person. You have this completely generated interaction. Obviously language models were not quite there at the time. Um and so it was made it was a mix of some templates and some like some mark of chain generated text. uh and it would just completely simulate this small talk
4:16conversations between uh everyone in the in the gallery space. Um so was always very fascinated on on the one end with uh generative models and like the early machine learning work that was paying out that time but at the same time there was this separate thread of simulation and what it means like what can we learn about humans by creating those very simple models of their interactions and their behavior. Did you generate the prompts or you know the 30-year-old whatever was it you generating them?
4:46How' you like Exactly. So, so the program would just generate those uh from Yeah. A lot of it would be kind of m lib style of just you know you have lists of different professions, lists of different uh personality types, list of different uh ages, things like that. And then he would just combine those things together. And then maybe the next project we go is Anani Valley, Anani Road, which was one of the first projects that uh we built with uh with
5:16uh one of my two co-founders, Chris. This was taking um Pix HD, which was one of the early image to image models that Nvidia released back in 2016 or 2017.
5:29uh and it was a model that would take a semantic map of a scene and then generate a photorealistic let's call it uh output obviously very very early days so it was not very high fidelity outputs but it I think was the first image generation model that would generate at uh at 1K resolution and it was all trained on self-driving data sets so the semantic categories it would support were only you know things you would encounter on the road so it would be
5:58pedestrians, traffic signs, like uh bikes, stop lights. And so that was that was actually one of our first indications that we built this and people were making all this like very surreal imagery of uh a million a million pedestrians or a million traffic signs or like gigantic humans. And it was a indication that you could take a model that was trained on this very boring data set essentially of like not that many interesting things happen when you're on the road and then you can
6:28repurpose it and go very out of distribution and make something that was artistically compelling. And that was it's a summary of the thesis of runway in some ways like you can take the same generative models and if you look at them from another direction if you build interesting tools around them and you give them to artists they're going to do things that you don't expect. Very cool. I like the, you know, UX of it.
6:50Basically, you're just given an empty canvas, drag whatever, do whatever. And in the other one, like you see everyone with wired headphones, you know, like that's that's a sign that it's it's very Apple ads. Yeah. Yeah. Yeah.
7:04Take us to today. You've been doing this for seven years at Runway. How have we got to this, you know, like how do we go from driving simulator data to all this?
7:14and you you kind of cover the whole stack of generative media, you know.
7:18I mean, interestingly, we're almost back and uh you know, we're full circle.
7:21We're we're now applying our models and kind of beyond creative tools into real world scenarios. But it was a it was a long journey. It was very early on we realized that yeah, the first version of runway was a way to easily use the all the open source model of the of the day, things like pixtopics to and give them to artists. That was that was initial idea is those models are too difficult to use if you're not a machine learning engineer like what happens when you give them to artists. Very quickly we
7:50realized we needed to build a research org uh inside of runway and that happened maybe on year one and uh a lot of the mandate there was the image generation models of the time the video generation models most of the time or there were barely any video generations all of the time but they were not quite there where they could be productionized and brought into tools that would be part of creative workflows. Um so we need to push the frontier of the research and so maybe the the first four
8:18years of runway research was almost happening on the background until there was a moment in 2022 uh with light diffusion with uh deli 2 where you know there was that step function change and you guys maybe remember I started l space because of basically lat diffusion and stable diffusion because I was like wow this is not only like feasible. It is actually doable on consumer hardware.
8:46I think the delta is also huge. Like I learned pixtopics like this was intro to ML the TensorFlow like Jupiter Google Collab notebooks are like this and then you have a sudden step function change you know with diffusion whatnot. Any other ones sense that like there were clear examples of what early diffusion were to get to here. any other changes in key technology uh research between uh 2018 when we started in 2022.
9:14Um so one of the early work that we did in runway was solving segmentation image and video segmentation. It was a very important problem because most VFX involves essentially separating subjects. Yeah. Rotoscoping extremely manual process. Nobody nobody enjoys doing that. Uh and so a lot of the early days of runway was building this tool that was called green screen and it was for a long time the main thing that people were using runway for. It ended up being used in uh everything
9:42everywhere all at once and a bunch of other kind of high visibility films and series but that was essentially runway for a long time was a post-production tool until late in diffusion generated gen one gen two um happened.
9:56Cool. I mean let's let's go past that moment. you've come a long way that you started releasing your own models. Maybe describe that journey as well.
10:05Yeah, so we got at a point yeah in kind of mid 2022 when it became clear that we're doing research at a fairly uh fairly small scale of of compute and it became clear that like scaling laws would apply to um image and video gen in the same way that we're applying to language generation. So we made a big bet and I think at so at the time we signed this deal to build a cluster of a thousand A100s which at the time we were
10:32a series B startup that was a almost you know slightly irrational decision maybe but we really believe that if we trained a video model at the large scale uh we would we would get like a great model at the end and at the time the goal you know the goal or we set the goal around fall of 2022 of what is what is the latent efusion stable diffusion moment looked like for video and at the time the the best model of the time was
11:01called cog video. Uh it was one of the early kind of video models was very 256 x 256 resolution very not not very high quality. Uh and so we decided we're going to build out this cluster and we're going to just invest in like in in building out our our own video model. Uh it became clear as as we're training Gen One that it was it was difficult to get to fully uh we wanted to build text to video. Uh but it it became clear to us
11:30that an easier starting point would be to to start from video to video because when you have a stronger conditioning it's it's basically an easier problem to restyize an existing video versus generate a video from scratch. And so we released Gen One first back in it was January of 2023.
11:48It's just a fun visual podcast.
11:50Honestly, like we can see February 2023.
11:53What was the state of stuff?
11:54I mean, it's so interesting cuz at the time when you see those results, you think this is this is so incredible.
12:01This is like it's almost like image generation or video generation is solved. and then you look back a few years after and it's like it's obviously it's just like you get used to results very quickly uh with those models. But at the time when we started seeing those results it was it was you know it felt quite quite incredible uh and the level of like quality you could get and um so the gen one was a depth conditioned
12:30video model. So it would turn it would it would take an input video uh it would predict uh it would you first convert it into the depth map and then we would generate uh pixels with a with a light and diffusion model.
12:45Yeah, very effective.
12:46I didn't realize how distracting the blog post would be. Sorry. Yeah, but uh one of my favorite examples actually of the on those on gen one was both if if you go up to mode three or mode two, there was this storyboard use case where people would make would basically um you can mess around with make a city out of books or out of boxes and then they would they would kind of shoot a video with their phone and then translate it into a for the realistic output. there was all these ways in
13:15which those models were starting to be used for storyboarding and also for really uh and then if if you go to mode 4 like of taking kind of untextured 3D scenes and then turning them into photorealistic output. So, we saw a lot of use cases early on where people that were familiar, you know, where Power VFX editors would just take a a Blender uh render and then they would get translated in with Gen One or create a
13:43scene in Unity and then take a capture a video of it and then and then translate into, you know, restylize it. So, I still I still think video to video is powerful. I think we we had a recent video to video model as well and it's one of my favorite ways of using using those models is essentially using them to to use ground truth video as like the initial inspiration and then translate into into different styles or different outputs. But I think we're going to go into like the rest of runway and catch people up to speed today. I did want to
14:12cover the let's call it this the stable diffusion controversy or you know what happened with stability AI whatever you know I think there was like sort of two sides of the story I think there's part of that is a normal thing of like people you know join and leave companies but what is the you know retrospective now that you know there's been some years behind it yeah it's a very it's a very long story to go to go into I think which I remember you actually wrote a really long post about we probably cover
14:41the whole hour to to to go into it in in more detail, but uh essentially you know there was the latent diffusion paper that came in um I think that was at the at the end of uh 2021 and then Patrick Esser who was one one of the researchers behind uh latent diffusion and he worked at runway at the time he built latusion collaboration with Robin Rumbach and a few other folks back in the at convis uh which was a a
15:10lab, the research group. Yeah.
15:12And uh after releasing the early latent fusion model, they um essentially they were you know the goal was keep working on versions of the model going to scale it up uh incorporate new data, incorporate new tasks and stable diffusion was basically the same model but trained on more compute and then with a few more tricks like classifier free guidance paper came at some point I think in the early 2022 which like was a big prompting improvement. Yeah, you know that
15:40improved results. I was trained on better data. So like the aesthetic subset of uh of lion, but it was effectively you know the same uh the same underlying architecture and there was that big training run uh that uh happened on stabilities cluster uh stability kind of financed uh finance that run and looking back at that story I think it was the work to build and train that model was was done. It was a it was a research project. It was done
16:09as part of like continuation of the latent diffusion work. it then I think the model became very successful and it um I think there were the and and I think as as a result of its uh it success other companies tried to uh figure out the commercialization path for it but for us it it was very important that we try to you know we we make sure that we it was meant to be an open source research project and so the
16:37we decided that we should continue releasing versions of it uh since that was kind of the original the original goal of stable diffusion and that led to releasing stable diffusion 1.5. There was maybe a day of a bit of uh miscommunication there but ultimately that was resolved very quickly within within hours. Um so yeah that was nice.
16:57I just wanted to you know you have the you are actually one of the main players in in that sort of journey and so it's nice to hear from the source of like what happened. Yeah.
17:07Yeah. I mean uh I I think it's all it's all in the past now I would say. Um, and like both companies, you know, stability took its own path. Runway took its own path.
17:16There's still I mean James Cameron is backing the new stability whatever they're doing with the the Hollywood studios. I don't know what they are doing. I think one thing that that impresses me and I'm happy to move on is that back in the that time let's say like 2021 2022 there was this community of people that you were involved in that was researching all this stuff right and like from everyone I talked to who was active then it seemed like it was fairly obvious that somebody would do the hero
17:45training run that would produce stable diffusion. So I guess the question is you know like you had the you you were you had made investments you you you had the foresight. Is it accurate to say like that is reflective of like what people were thinking at the time or was it still very much like well we'll use it as like a post-production tool or something. I don't know you know like where where in the sentiment were we that maybe you can sort of think back to like what what the community was like
18:11back then. I reminisce and I um I I think of very fondly those early years from like 2018 to 2022 because it was a very small community that as you said were very convinced that this was going to be big thing and at the time you know anyone who because it was such a small circle and everyone who would like be part of that circle and like make projects with it would you know immediately kind of get you know go viral. Uh so like I remember you don't even know who they are, right?
18:40They're just there some name on the you know GitHub dollar hugging face somewhere.
18:44Yeah. So uh so I remember one of the first big viral moments of creative AI was uh there was the neural style transfer paper um that the something dreaming uh I think it was called neural style transfer. There was also deep dream the puppy slide which was uh also also really cool. Uh but uh there was this project that uh Jin Kogan who was an early adviser of Runway and one of those uh marketing guys big uh creative uh
19:12creative AI uh folks. He literally just like showed a video of himself taking the New York subway kind of and going over the Williams Rook Bridge and then Salized it with um I think in in the Salivan Go or like one one painter and and that was like at the at the time that was like so so cool and it went viral and it was completely revelation to people that you could do those with generative models and that was only you know it was less than it was maybe 10
19:41years ago. just like as an indication of like how how how quick like things have have gone.
19:46It's pretty crazy like even since then you've kind of got people at every level of the stack. You've got devs, creatives, artists, hobbyists, you got everyone using it. And for people that tried stuff early, they'll remember how hard it was to use regular diffusion, right? Like nowadays you can use your favorite chatten or whatever, give a sentence, get a beautiful output. But diffusion was like you know the whole ultra HD 4K high resolution like prompting these things was very different. Uh anything you learned on
20:15the tooling side like from the offerings you guys have now so like creatives devs uh you really took the research and brought it to everyone to use. Um anything interesting there to share?
20:28We had to build the entire model serving infrastructure for video diffusion models. there was no nothing else uh already like cuz we had gen 2 was the first tech video model I think out in the in the market. So many things that we learned over time. I think the I think the biggest one was like we you know it was very clear early on that text to video was not going to be the the answer like you like people wanted a lot more control than that and so we invested in like control building on top
20:57of those models very quickly. you know how how do you use the camera trajectory as control? How do you use an initial input frame as control? So, so that was a very early learning for us with text to video was um like Gen 2 was an amazing you know step function improvement in the quality of video models but it was used much more in an exploratory way because there was nothing to ground it to. There was no reference that you could bring into it.
21:24there was no you couldn't really control the the camera motion, you couldn't control the object motion. And so the first year in 2023 was really all about what are all the interesting ways in which we can condition those models. And it was a lot of just post training runs on top of the base model to figure out like what how do people actually want to control them? And so there was like this quick succession of the we uh it was called motion brush which was you could like you could basically draw arrows and dictate where things should move in the
21:53scene. There was camera control that was you could just describe like how you want the camera to move in the scene.
21:59And because we work with filmmakers from kind of the most of the history of runway, we immediately got this feedback and got this, you know, decided that this was worth investing in. And so controllability became a big theme I think very early on as we're building as we're building those models. Something fun that I haven't really talked about too much was just how Gen 2 came to be out of Gen One. So, so it was a bit strange because we announced gen two
22:25months after gen one and um it was it was before gen one was even generally available.
22:32But gen one was a depth to video model.
22:35So it would take a depth map and it would convert it into RGB. Um and we couldn't get you know text or image to video to work directly and that's why we started from depth to video. Um but uh you know and we had discussions of like okay we need to spend the next 6 months actually investing in tax video maybe increasing the compute scale or the model scale like train a larger model and I I had this weekend project idea
23:02which was what if I take a model that uh starts from text input and converts to depth maps and then and then use gen one to convert the depth maps into RGB. Uh so and so Gen two was basically that and it's a hackathon hackathon pipeline but it looks good and it worked pretty well. I mean there there were you know if you with a knowledge that it has this like two-stage pipeline. You can tell in some
23:30cases that the structure of the video looks a bit off because you had to generate the depth first before you go into do the output video. But it worked and it allowed us to bring bring this to to our to users very quickly. But it it's actually now it's interesting because like people are coming back to this almost two-stage approach. Like if you look at the the Reeve text to image model that came a few months ago. It had this um this planner model that would
24:00generate bounding boxes before it fed that into the diffusion transformer.
24:04Yeah. Ideogram also the same day. I I remember that was very strange that both of them came out the same day with the same exact innovation. It's a small community. I think I'm like this is like this is completely coincidental, right?
24:16People talk. Um so so yeah, there's there's definitely something into this approach and obviously now like every single video like video generation model in production uses a a complex prompt completion pipeline under the hood. I think that's no secret that there is humans are terrible at prompting. I think across the board, but yeah, I think like the original Sora one blog post even told you that what happens after your input is re rewriting your prompt, it's much more descriptive about
24:46what you would want.
24:47Exactly. Yeah. Um and there was the Delhi 3 paper beforehand that uh was the kind of the first public uh description of the fact that synthetic captions and really detailed captions work really well and then Sora kind of built uh built on that. Yeah. Yeah. So, it was 2023. We were kind of releasing all these updates to Gen 2 like the camera control motion brush. And there was actually something very interesting about camera control cuz it it was the first time that you felt that instead of
25:16like you were creating video, you were creating a short video, you were actually navigating inside the world.
25:21And I think camera control was maybe the seed of some of the ideas that we had around world models and really opening up that research direction. we realized, you know, it was this era and this series of, you know, gen one and gen two models really proved to ourselves. Yeah, this is this is the uh so this is not the original camera control. This was the update to camera control on top of Gen 3. But yeah, I think it made those
25:49models usable to to filmmakers. uh I would say the so camera camera control was very uh was was very popular and so we realized you know there's one way of seeing those models which is you know you're just as a as content creation machines and there is the other way which is you're as you're predicting video in order to predict video well you need to simulate the world in an increasing and increasing capacity and if scaling laws apply on video just like
26:18they apply on language models then as we scale the computer that we put in those models, then they're going to be able to simulate physics. They're going to be able to simulate human actions and dynamics increasingly well and predictably well. That was the thesis about around our efforts on world models. And we spin up this research group to just focus on on world models and how how do we turn the video generation models that we're building into something broader and something that would be useful beyond uh also
26:47content creation as well.
26:49And that was roughly when?
26:50Yeah. So that was in in late 2023.
26:53Interesting. You know, like I think uh a lot of people have been saying a lot of video gen model companies have all pivoted to world models these days, but like you know 2023 you're posting it. Uh I mean it's it's debatable whether it's a pivot like arguably that's what you always had to do anyway, right?
27:11It's in a way an expansion of the applications of the models as as they become more capable. the early signs. It seems like the original models you guys guys had, people would say it's very not bitter lesson, right? You're adding rewriting prompts. You're having all these one-off things, but that's just the state of the tech as it was versus the future of, as you said, you can scale it up is, you know, we can scale up to world models.
27:35Yeah. So it just became and and and if you looked at the outputs of Gen 2, it was not I think it was not obvious to people that this would scale to become a general simulator of the world like you had very limited movement, you had you know very low fidelity resolution like obvious mistakes in human anatomy like all kinds of limitations but it was just you know uh the idea was that's just GP2 and GPT2 you know you can barely
28:02generate like coherent sentences similar Gen two can barely create coherent video, but if you scale it up, you're going to there is no reason why it shouldn't work in a way. It's uh and I think that was that's that's always the mindset of kind of of runway is like this extrapolation of like if you know like even when we started in 2018 and you looked at the results of the day you need to look more at the trend of like where we were in 2018 versus when we were at the you know when the first gun
28:32came out in 2020 uh for 2014 or 20 uh 15. And you start from like 32x 32 images of faces and then by the time in 2018 you could generate you know street images at a 1k resolution and it was kind of the same with world models very early signs of something much bigger.
28:52Yeah. I was almost think like it's kind of diffusing into focus like if you look at our visible output from year to year to year it looks like a diffusion process itself.
29:01Yeah. Especially watching the early like old old blog post, you can really see the choppiness, the details like human civilization starting from random noise and then noising into Yeah. Yeah. Just run it.
29:15Yeah. That's how you know you're on track. You know, you're still noising, right?
29:18Yeah. I like the way that you guys phrased it when you announced it in June, which is, oh, you had a sort of video essay. The human mind is no longer the center of AI. Our world is, right?
29:27which is uh you know let's let's call it the the past 5 years of of LLM based AI is very much like trying to emulate human preferences and human speech but now that's like mostly solved or I think that's like some of the context of your essay which you also wrote around the time and now it's like the focus on modeling the world accurately.
29:47Exactly. Yeah. So the the the way we see it is uh there is that um that uh initial mission statement of deep mind which is uh uh solve intelligence and then use it to solve everything else but I I think it's starting from everything else uh could be valuable of like starting from you know there is just so much complexity uh and detail in the world that in order to that it's it's hard to learn directly from just human descriptions of the world like we're
30:17assuming saying that, you know, like language models learn from everything that humans have written about the world, like our own understanding as of, you know, the 2020s. And there's just so much that we don't know and so much that's not captured by existing text uh about both the, you know, the low-level dynamics of the world like we're not describing in detail. you know, if if I tell you to describe like how do you tie your shoes, that's a very difficult thing to to describe in words, but it's
30:47very obvious thing to demonstrate. And so I think there's been and there's, you know, more of paradox like we're constantly underestimating all the complexity that goes into very like things that we do subconsciously as humans and we don't even necessarily always have the words to describe them.
31:06And so in my mind the simulating the world and simulating um physics, simulating the dynamics of the world has always been kind of underestimated uh uh compared to uh we place too much emphasis on the things that are easy to talk about. Uh but there's just all this complexity and kind of richness of the world that if we just try and train train directly on that observational data instead of training on how people
31:35describe the world, we would learn something new that we wouldn't otherwise know. You think that the present architectural paradigm is fine. You don't need like another layer like Japa, you know, like another famous uh New York AI leader would say.
31:49You know, we're a very pragmatic research lab. If uh we have evidence that an approach works better than the approach that we're taking, then we have no qualms to taking it, we just have seen no indication that video prediction itself doesn't scale. And even if you look now, you know, not just our work, but the work of others, you're seeing in robotics, some of the most promising work um starts from video prediction models and then you adapt them to also be action models, for example. Um, so there is very little evidence that you
32:19need something else and that your time is better spent on a novel architectural change compared to improving data and improving the and scaling the current the current approach. And so we don't have any indication that you know the the there is that counterargument that uh I think there was a a tweet by Yan Leon a few days ago that you know uh understanding the dynamics of the world is very different than uh generating uh
32:48cute videos and your answer is no. They're the same thing.
32:52My cat videos are the same as understanding physics, right? Cuz if you want to generate, you know, obviously video models can cheat and like they could you could give like successive uh shots of the scene in a way that doesn't require you to actually simulate difficult physics. There's always like all these different ways in which you can hide the deficiencies of the model. And it's important not to be kind of too tricked by the performance of the current video models. it's easy to, you know, cherrypick examples and
33:21and think that video models are further advanced than they actually are. So, there is a lot a lot more work that we need to do to improve those models. But in my mind, very similar to language and like we've, you know, you go from barely coherent sentences to something that, you know, could hold a conversation with a human to something that could can operate autonomously for for a day and like create entire code bases. And the main difference there's obviously some architecture improvements along the way
33:51but the main thing is scale and so it's the same bet for video and we have no indications that this is saturating like we have benchmarks that we use for measuring the physics of those models and we see those predictably improve as we scale those models. So there is if if you want to pull up uh physics IQ uh is one of those benchmarks that measures how well does the model perform at all mechanics or fluid dynamics or optics.
34:16I'm curious if you've seen any emergence any scaling law around this.
34:21Yeah he's saying there is a scaling law right.
34:23Exactly. Yeah. So, so the way those those those models those benchmarks work is you you know the researchers have gone and like captured uh a few videos that are representative of different physical phenomena and then you can take the first frame and then pass it through an image to video model and then generate generate kind of a roll out that shows what should happen next. So you have a a ball hanging from the ceiling and then you use that as input uh and then you you the model predicts
34:51how the the the ball should fall on the ground. Uh and this measures you know we have an intuitive understanding of physics. I know you know you can imagine what will happen next if I drop this uh this bottle. So it's measuring that same intuitive physics understanding of those models and we've measured that at different model scales and we see uh compute scales and we see that the score and physics IQ predictably improves.
35:15There's other you know tricks and techniques that you can make to improve the score even further but even scale alone helps uh in in the in the model learning better physics. My main sympathy with Yan Lakun is the uh Plato's cave allegory, right? Like you're like learning on the output of a thing, not the internal process of a thing. And it's very very noisy. And you know, if only you could observe the internals of a thing. It's hard to observe the internals of a human mind.
35:42But you can very much observe or at least we have a whole bunch of science and physics that we're ignoring on how to model physics and and movement and uh you know gravity and you know other interactions uh and we're just like throwing away all that and just saying just just scale data which is very much the lesson of unsupervised learning but it feels wrong. I think that's like the main idea.
36:05I think the history of machine learning is large it feels wrong.
36:09Yeah. It's a bitter lesson right now. is is the simple answer to that I guess how much can you scale so like even on let's say the video generation side like there's one side of video understanding video generation are we still going to have tools where it's like I want to generate 2 hours 20 hours um there's a inferral way to do it in batches and stitch it together but like do we just keep scaling do we just continue long generation consistency all that would scale and like tying it into
36:39where we're at now from we the runway 2 to 4.5 like technically what what advancements have we made to today and then where do you see things still going?
36:48So part of the answer is definitely scale um and that was we we learned that lesson in a in a big way for with Gen 3. So Gen 3 was the model we released the year after um like in 2024 that was a few months after Sora was released um so yeah there's an interesting story of that that that that came to be as well.
37:08Gen 3 for us was you know the the first time that we need really needed to build. Basically we had to learn all the lessons that the language model world learned in it in in three years in the span of a few months. Uh one of the biggest changes of Sora was using diffusion transformers instead of connets. So a lot of the early you know latusion models were all um uh comnets for the diffusion model part. And the diffusion transformer paper came at some
37:40And it basically showed scaling laws for image uh diffusion transformers. And we we realized at that point that we needed to invest in infrastructure for model parallelism for really scaling scaling training to larger than you know a few billion parameter models. And we spent maybe the, you know, most of the fall of 2023 building out our infrastructure for distributed training. And we had a lot
38:07of false starts and a lot of failure in trying to scale um image and video diffusion transformers. And at that point, uh, you know, February 2024, Sora comes out and the results are very much superior to what Gen 2 could produce.
38:23There were a lot of you know a lot of chatter on Twitter about runway runway's done uh like there is uh there is no way runway will catch up and if you remember also open AAI in the early 2024 it felt very uh like it's a formidable opponent now but at that point it you know they were on the top of their game you know nobody could even get close to them there was maybe Gemini was just the first version of Gemini had just
38:52released So when OpenAI came with Sora and it was such a big jump of like quality it gave me there was like an existential crisis for for a few hours but that I think the the amazing thing about runway and like I think the you know we've been around 8 years now which is almost we're dinosaur in AI and we had to like there was a lot of those moments we had to learn adapt very quickly and build out skill set in the team that we didn't have and so you know if you ask anyone what is their favorite
39:21time at runway that was there during that time. It was that that push in like 3 months to get to a model better than than Sora. Uh and it you know we scaled 10x the model scale the the the model size and the you know compute that we were training on. Uh we figure out model parallelism we had zero expertise in that and then we came out with Gen 3 during during that summer. So um so that was a big turning point I think for the company where the the research work grew very quickly and we really started
39:51pursuing this vision of the general world model uh in in our nest I think after after gen 3 was out. Yeah, I mean that's the the amazing thing about building when you're building there's no stack to you have to invent everything yourself. You have to be completely full stack you know now now I think like there are inference specialists like foul or whatever that can help with like model serving and I I think you guys work with them as well. Uh but yeah like it's you know but at the time it was just it's very interesting to think
40:19about what you do when when Sora comes out and people are questioning whether your company should still exist.
40:25Yeah. And yeah, there was no you know there was no VLM of diffusion models that we had to build the whole model serving infrastructure and make things efficient. And a few months after we released G3, we released the turbo version which I think was the first step distilled model in production.
40:41That was a whole trend that we covered as well. Yeah.
40:43So that allowed us uh actually to serve uh to serve the those models at the the larger scale cuz I think the first version of Gen 3 was was quite expensive to serve. You know, I think the the whole like trend in like consistency models, lightning and uh turbo and all these things somehow didn't really stick around. I don't know if you have any reflections on this because at the time I I was like, well, obviously everything should start with a distilled model first and then you can upscale, right?
41:10Then then basically your your bigger models just turn into fancy upscalers, but like you should always draft with a a smaller model and faster model, right?
41:20Because you can get it so quickly like near real time. Yeah, I I would not be so sure to to say that didn't stick around. I think that's uh it's it's likely to uh I mean that there's a lot of step distill models that are actively used in production. Uh there's still obviously a gap in quality compared to the you know the uh nondistl model. Uh but in my mind we're still you know there is a two to three year offset from
41:47language models. So the things that um so it it's just a matter of time before there is better distillation techniques. Um you know we we use right now we have a real-time model core characterist that I think is the largest deployment of real time video models. Um that's a a step distill model and it's actively being used. It's a very specific use case compared to a general video model.
42:10So this is by the way this is avatar consistency character.
42:14Yeah. So this is a a talking avatar uh model. Uh we were able to you know we we we optimized the the hell out of it. Uh and and it uh it generates at 24 fps and it's a it's a step distilled auto reggressive video model. So if we look at our world model direction uh a big component of it is starting from the birectional diffusion that basically generates an entire video at once and making auto reggressive show. So you generate one frame or a few frames at a
42:44time. Um so there's a lot that goes into that pipeline of getting to a real time model. It's first you need to make it into a causal auto reggressive model and then you need to turn it into you need to do some additional step distillation to get it to actually be real time. Um and I think that part is actually just just starting. Um I would be very surprised if we're you know 2 years from now we don't primarily use real time models. To me, real-time video
43:12generation is just inevitable that you know it has much better user experience is much cheaper to serve and you know the quality gap between the base model and real-time model is only going to close as we uh as we figure out better uh distillation techniques and we made a lot of progress there internally on maintaining the quality of the base model when we when we distill them. How much of this is transferable? So is it the same base model like if you're doing diffusion across the whole sequence and
43:40you're converting it to step auto reggressive distillation is this like distillation where you still need to train both you can use the same base and converter what's that process like to go from regular model to something that's real time on a technical level.
43:55So the nice thing about diffusion models is you have uh two axis of distillation.
43:59So there is the you can distill to a smaller model which resembles what you do in LMS or you can distill in in terms of taking less steps uh less diffusion steps. So you could take a model that generates in 50 steps and and and generate in four steps and get to um you have some performance uh degradation but very often you get comparable comparable outputs. So you can even take the the you know the large frontier model and
44:28distill it with stepation and get to a real-time performance and that's what we've seen. So um depending on the use case in some cases we might also serve with a smaller model but in a lot of use cases we actually just use the the frontier model and we're able to make it work in real time.
44:44I think this might be a good time to cut over to his laptop to show off some of the real-time stuff that you're doing.
44:50This is one of the research updates that we uh we did recently. Um so we've been working in uh in getting our general world models to uh different applications. Uh one of them that we think is very uh is very compelling is using general world models as essentially uh an interface a universal interface to software. This is a a version of our uh world model that's called an interface world model. uh and
45:19the idea is that it uh it essentially uh replaces uh uh the a front end of a software application. It renders the pixels directly of an interface and it's trained to predict what happens next as a result of a a click or another interaction you have with interface. So this is all pixels uh it's there is no HTML CSS react that's powering this interface. This is directly at the output of our real time uh video
45:47generation model and it takes clicks directly as input and drags. Click and drag, right? So it supports uh yeah clicks, it supports drags. Uh it also supports scrolling. Um and the the the amazing thing about this is that you can effectively describe in the prompt how you want different elements like what do you want the behavior of different elements to be. So it's almost your you
46:14can turn uh an interface from a you know markup language description of like an HTML interface and instead you can just describe the interface you know if I press this button I expect this to happen if I press this button this should happen and it's useful we believe both for prototyping for like just testing like what different interactions would feel like um you can also add audio to it so it's a video audio generation model so you get you essentially and describe both what the
46:43visual outcome should be of your click and also what the if if there's a sound effect that comes out of it. So we we believe that's gonna be a much more flexible way of building software just render it just you know why why generate the code that generates the pixels just generate the pixels directly um uh it's the end to end philosophy applying apply to uh to front ends so we think there's a few interesting use cases so you can
47:12build creative tools on top of it um we think that you know for any kind of use case that involves a lot of exploration or like educational use case where you want to learn about a new concept and you want some kind of visualization and and kind of open-ended exploration. We think those you know this is a very powerful uh approach. Yeah, you can imagine new new forms of uh design
47:41industrial design software that could emerge as a result of uh of those models. And this is all uh generated in kind of in in in real time as well. So um you you can yeah you can build a lot of interesting kind of camera transitions and kind of forms of interaction that are very difficult to to build otherwise. And one way in which we evaluate this is what if you try to generate the same interface with uh with
48:10cloud by just you know prompting cloud here's an image reference of my interface that I made in Figma or that I created some somewhere else uh create this particular interaction which in this case it's you know drag that object uh upwards um and beyond it being slower it's also very difficult to capture some interactions by just fully uh with with just LLMs. So we think that this this is
48:38likely to be the way that a lot of the future like software in the future will be created. Uh and one of the additional benefits is personalization might be a lot easier done with with those models like you can essentially try out different prompts based on who's who's visiting the interface. you can uh more easily you know prompt engineer the the interface to have larger size um text for more accessibility reasons or you can make this or or like if you have a
49:07particular study preferences so so we're very excited about this approach it's obviously early days and I think we'll need to u make it more cost effective as well to serve those models because you know running a real-time video model versus just purely rendering HTML there's obviously the computational needs are much higher uh but we do see a lot of potential in this approach to building kind of front end interfaces.
49:31So we covered this similar thing with flip book before with our uh Ethan her episode with Grock uh video and I think it's very engaging visually. I think it's maybe good for education but it's it does sound expensive. I think there's an upper bound to how expensive it will be though right like you know the inference cost will go down over time.
49:50you'll figure out ways to optimize it effectively. When it pauses, you don't you're not receiving human input. You don't have to generate anything, right?
49:57So yeah, I mean you could also like in this case you have ambient motion. So there is parts of the of the screen that might you know if if you're let's say you want to um visit Paris and then you you get this interface that you allow to explore. You have people walking or like things happening. Um it's it's an yeah it obviously makes it more expensive because you need to run the model all the time. Maybe you have some looping mechanism so you don't need to do that.
50:25But all those things I think is stuff we need to figure out. Yeah.
50:29I think our first consideration is let's make this clearly find some use cases where it's clearly a much more compelling interaction compared to traditional interfaces and then it's a matter of time before it becomes more cost effective to serve.
50:42Yeah. when it comes to the people walking, you know, I think the approach that makes the most sense to me is basically Moon Lake Nick.
50:49Oh god, I I keep messing up their name with Chris Manning and Funen. I don't know if you've come across them where they basically map to some kind of game engine. I think it's Unity or something or Godo and they you can obviously obviously script some NPC behavior behind that and train on that. Whereas here you can really imagine whatever you want like that is a UI, right? like and and it feels like more tractable I guess to uh create a world model of software that is interactable because we have
51:17many of examples of that and you can you know do your fancy our environment stuff on that then it is scaling up to embody and real world physical use cases but this is a nice first step or you know there's the opposite of you have like 1B models 350 million parameter language models that just get so small that they're just predicting like you know fishes moving.
51:38I mean, small models are not 12b. So, ultra mini on device, but uh no, I think it like it puts it into perspective. At least the car one for me, like the applications, right? The amount of work to do that. Sure, you only make one model year car per year, but applying this, it's also a cost-saving to have to manually make all this, right? So, it opens up a lot of possibilities, too.
52:02I'm curious if you extend this out 2 3 years. So where do you see things going even further? You know, effectively the the the end game of something like interface world models is you have a fully neural operating system. So I think uh uh Andre Karpathy has written about that quite quite quite a while back. But it's uh you know I think to me it's it's a bit um it's a bit odd that you know we have for
52:31example with an interaction with with an LM of today you have this LM that can basically talk talk to you about anything. It can you can take the conversation any direction you can you can it's very general so it can solve all those different tasks but you interact with it through a very rigid interface. And so to me it's just a matter of time before the interface itself becomes learnable and becomes you know part of the the the whole loop of like you're not just delivering you're delivering an application end to end and that means you're delivering the
53:01language model but you're also delivering the the render and the pixels and then and then that's also a learnable component and the concept of applications might not necessarily I think we need to figure out new abstractions for software um the the concept of application comes from this idea that you need, you know, separate kind of code bases to describe to for to power each individual um tool and each individual application,
53:29but you might uh you might think of something a lot more unified if you're if you have a a video model that's actually generating the interface as you go. Um so it can take context from an LLM and allow you to combine kind of different uh different functionalities that traditional would live in different uh applications. So so it's a it's a way to solve you know software end to end uh effectively. We also see this as a powerful way to train computer use
53:56agents as well. Um so this is you know one way to see this as and in general with world models there is those two directions. One is world models for humans and world models for forgent to train agents.
54:10And so for every new work of uh world models that we do, we have this these both uses become possible. So this is a powerful synthetic data generator for training computer use models. It could become uh a live uh RL environment that you could use to do online RL with a with a computer use agent. uh and you can get wide diversity of different interactions, kinds of interfaces u just generate on the fly that um to improve
54:37the how robust the your your your agent uh becomes. So that's the same also with the world models that we're working on for a robotics use case as well.
54:46Is there a research breakthrough that you're waiting for that would unlock the next set of use cases that you really want to pursue?
54:54Long context is a very important one. So being able to maintain consistency for long periods of time and that depends on the use case. So for our characters model for example or for the interface world model it's easier to maintain long sessions of interaction. If you go into more open-ended worlds that you navigate and you take arbitrary actions in we like there is more the context at which you can and duration which you can generate becomes limited much more
55:23quickly. So we see more degradation and error accumulation happening. Um so the biggest challenge with auto reggressive models is error accumulation is basically you're fitting generative frames back into the model to generate the next the next frames and if there is any small errors they accumulate over time. That's not a new problem. It's a problem that LMS also have and we've seen the you know the ability to generate now really really long outputs.
55:49So it's a sol problem but it's definitely still still a challenge.
55:52Yeah. And what is the state of the art?
55:54Uh so for for Grock it would be like 10 to 20 seconds of context going in there for video with our characters models were able to generate up to 30 minutes of uh of video autogressively.
56:06Yeah. That's just for the avatars.
56:08Yeah. So, if we look at um GWM Worlds, which is more our open-ended world exploration model, um it's it's on the order of a few minutes, which is Yeah. Um it's probably enough for people because you have to cut to the next scene anyway, right?
56:24Yeah. It's not it's not the ideal game experience if you have to restart every few minutes. So, I think but uh I think it's yeah, for certain kinds of experiences, you can you can work around it. Uh ideally you able to just generate forever and it doesn't it doesn't degrade and I think that's a matter of time if before we get there.
56:43Yeah. Genie has like one max one minute you know. Yeah.
56:47This was your you did a study on robotics. I think I also have just your runway robotics page though.
56:55So last year we released Gen 4.5. So that was our our latest based model and we've been as I mentioned we've been doing all this work in world models and uh which essentially a lot of our approach to world models is how do you take a birectional diffusion model and make it auto reggressive and make it accept actions. So instead of being a video you watch, it becomes a a simulation that you step in and you can, you know, control it every step of the way. You can explore counterfactuals like what happens if I take this action
57:25versus if I take this action. And GWM1 was the it's the world model that we built on top of gen 4.5. So we did all this auto reggressive and like uh uh dissolation uh autogressive and then step dissolation on top of gen 4.5. And one of the biggest use case that we saw for GWM1 was in robotics. One thing we like to say is we as we scaled video models, we accidentally uh created a
57:53state-of-the-art model for robotics uh by just scaling video models. So we realized at some point um mid last year that robotics labs started coming up to us and kind of asking to use video models for synthetic data kind of asking us to post train our video models to work really well for robotics uh so that they can they can use that to basically generate variations. That was the first use case that we saw and then increasingly became clear that the models will be useful beyond just
58:22creating synthetic data to train robotic policies. They would also be very useful as simulators. So that means that you can use uh a video model online to test how your robotic action model performs.
58:35So you can take an action roll and then get the outcome of the action inside the world model and then and then continue that loop like this closed loop simulation and you can use that to evaluate how well your robotics model works. Um and the biggest thing that I think you need to solve if you want to build a simulator is establishing real world correlation that if you take an action inside the world model if you take the same action in the in the real world you get a similar outcome. So that
59:04was that was a goal of some work that we did earlier this year. So if you go to the first link. So that was uh essentially we wanted to establish that that you know uh real real to sim correlation for our world model so that if you do a series of actions inside the world model and if you do the same actions in the real world you get similar outcomes and we took our GWM1 model and we we used some benchmark data that there's this robarina uh benchmark
59:32that's very commonly used to evaluate how well do different action models perform and we use the same scenarios and ings and embodiment inside our world model. And we measured the correlation of how well did a action model perform inside a world model versus in the real world. And we saw that we could get very good correlation between uh between our world model and reality. And that means that if you want to evaluate how well your robotic policies perform, you can
59:59scale that much faster inside simulation instead of instead of having to do that with actual physical hardware. And so that was a first indication that the our models could be uh quite useful in robotics. And we saw as as we're working with robotics labs that that became like the first use case where they could use video models in a way that fit into their their their kind of training pipeline.
1:00:24Can I ask what the difference was from 4.5 to solving that? So the sim to real gap has always been the issue, right?
1:00:31you train a robotics model on video data, it doesn't generalize to real world and the simulation had an issue.
1:00:37So, seems like you solved it. But how?
1:00:40Yeah. So, a big problem with simulators is, you know, if if you're trying to simulate rigid objects, like it works quite well if you know if you if you can describe the physics of objects very accurately. Um uh then you're able to use um Isaac Sim or or Mujok or one of the traditional simulators. But for more complex interactions with with cloth for example or um like slippery surfaces um you know with the all the complexity
1:01:09that you you want to be able to solve with a manipulation with an action model that solves manipulation tasks. It's very difficult and so timeconuming to build you know for each of those environments and each of those tasks build the simulated version of that the digital domain of that of that environment. Whereas with a world model you you just need to provide the first frame and then you just can roll out the poles inside the first frame. So whereas you know we compared it to methods that
1:01:37required like 3D scanning an environment and then 3D scanning each individual object before you can now uh you know you can bring that simulation. Um whereas with a world model you just take a picture of the of the environment and then and then you're able to test how how your policy performs. Our general thesis on robotics is you know there is companies that are leveraging a lot of teleoperation data to to train robotics
1:02:05action models. There is now companies that are using um yumi data which is uh essentially human uh egocentric video where humans use kind of uh robotic creepers to perform different manipulation tasks. And then there's companies that are focusing on egocentric data which is you know you strap a GoPro on on someone's head and then you kind of capture them performing a task. We think that you know and all those are great sources of data for
1:02:34training robotics models but the most plentiful source of video data is third person video data. It's and if how how do we as humans learn how to perform different tasks? A lot of it is by observing others perform those tasks. We don't learn from first person. We obviously do some trial and error and like uh to learn different things, but ultimately a lot of what we learn how to do in the world, we learn by watching other people do it. And that's how when you're pre-training a video model,
1:03:03you're essentially doing that. It's a lot of third person video uh footage of people performing different tasks in the world, people doing sports, people doing household tasks. And the our main thesis is that video pre-training. Once you do that, you can then adapt a model to be useful in robotics use cases with way fewer hours of actual robotic data. So you require way less teleoperation data which is very difficult to scale. uh and
1:03:31and even if you look at egocentric data which is a bit more easy to scale compared to teleoperation data which requires actual hardware is still three orders of magnitude less of that that exists in the world compared to third person video data out there and so our thesis is and generally like the most plentiful source of data will ultimately ultimately wins third person video data pre-training is the right starting point for uh for models that you know you want
1:04:00them to generalize and be able to deal with new environments, new tasks, things that you haven't seen during training. That's the motivation for why we think our our models are especially useful in robotics uh settings. And we've seen that to be the case um as well.
1:04:16You said pre-training. So maybe it's like third person pre-training, first person SFT. Is that like a curriculum that you can sort of introduce?
1:04:26Exactly. So if we look at GWM worlds uh so GWM robotics so the GWM robotics it starts from Gen 4.5 it's the same video diffusion backbone right exactly yeah so you start from the base video model the one you're using to generate uh cats and dogs and other interesting stuff and then you uh fine-tune on a very small number of hours of robotic data so it's something on the order of kind of hundreds of
1:04:54hours compared to if If you were to pre-train a robotics model, the current pre-trainings go up to, you know, uh, hundred thousands or like millions of hours of of data. And you're able to get quite quite good performance uh, quickly uh, because the model leverages all the things that it has learned about the world and physics and human dynamics and the tasks that people care about from pre-training. And ultimately, you you want those models to generalize. You don't want to just be able to perform
1:05:23the task that that has been during training. And the diversity of actions and environments that that you have with a pre-training video data set is much larger than uh what you can kind of realistically capture manually. How's the scale looking like for the post- training? like you still want to do is it like roughly 90% of the compute in regular video diffusion model and then scale up a lot or do it like we want different robotic models for different
1:05:53tasks or just the one base really good world model can also apply to robotics.
1:05:58So currently uh we are postraining our models for specific uh embodiment that we uh for particular partners. So if they have a particular kind of single arm robot or bmanual robot or shimano robot, we would post train our GWM robotics model on their particular data set. Over time we see the different variants of GWM unifying like I would expect you know if a year from now or two years from now you have a single
1:06:26world model that can uh that can simulate manipulation task. It can simulate navigation which is a lot of the gaming world models are navigational world models you're moving around the space and it will also simulate kind of human behavior. So that's the character models. So instead of having three different models you you have a single model that's able to you know ideally you're able to simulate what it's like to be in the world. You're moving around an environment. You're maybe performing different tasks. uh you're talking to
1:06:55other people and that happens with the same uh a single kind of real-time video model that's generating that. Do you think you can solve self-driving? So if you are learning to drive a car in a simulator, you have a world model, your robot is basically, you know, car can manipulate so many axes. How far off are you from something like that?
1:07:15So a really good ADAS system. World models are definitely being applied to uh self-driving uh research right now mainly for evaluation use cases but our focus has been more on robotic manipulation. We've done some work on AV uh world models as well. Uh but yeah, we do think that world models are and video models are the best starting point for both simulators and also policy and the action models. So that's that's the
1:07:44other side to this is that once you have a great world model then you can just add an action head and it can predict actions as well. One way to think about it is if you take the starting frame of a scene with a robotic arm and you ask, you know, you prompt the model, generate the arm picking up an object, it would and if it generates an accurate enough video, then it should also be able to generate the exact poses uh in 3D that the arm should take to um to perform the
1:08:14same action. So this is the direction that's now the the popular term for it is world action models which is you're starting from a video model and then you're adding an action action head to predict the actions and it becomes a policy essentially. One thing I'm also impressed by is how much data you actually need to train these kinds of models. You probably can't say exactly how much but like you know like the the original uh diffusion models and from what I know even of the open source
1:08:42Chinese models is not that much data isn't it surprising what do you define as much data? Um yeah, it just comes goes in. Is the token count still relevant?
1:08:53So it's a bit more complicated and uh what is I mean just gigabytes, right?
1:08:58Yeah. Hours of video.
1:09:00Yeah. Yeah. I feel like something that's interesting is it seems like the let's call it tokens to param counts in language models has really you know maybe they're three years ahead or whatever seems to be a lot higher than uh video models still even though technically video has more information you know per per bit. I don't know if it seems intuitive or maybe there's just a lot of uh like the the variability between a pixel to the next pixel is not
1:09:28that high. So like maybe there's just a lot of information that is repeated.
1:09:32My answer would be it's still very early like the training video models will scale way further than it currently is and and you'll have capabilities that go much further than the current models can do.
1:09:45So one thought experiment that uh I like to use it's it's almost like the Turing test of video models or like the Turing test of world models. Uh um uh I I call it the lucid dream test is you have a you mean like the actual person lucid lucid dream.
1:10:01It comes from lucid rains, right?
1:10:04Lucid dreams is telling you're dreaming while you're in a Yeah, exactly. So lucid dreaming is when you realize you're inside a dream and then you basically be able to control what happens in No, there was also an inference guy called lucid dreams. Yeah. Quantization, very prolific person. Yeah. So, let's say you have a VR headset and uh and uh you're in a room with uh and you're wearing a VR headset and that VR headset, you know, most uh of today's VR headsets have a pass through mode so you
1:10:32can see directly what's in front of you in the world or you can obviously render kind of something inside the VR headset. And there's going to be a point where those interactive real-time video models become good enough where you wear the headset and you're in the same room and you're walking around and you're kind of and you're interacting with objects.
1:10:53You're able to kind of move freely in that room and do um and interact with any object. And then at the end someone asks you did you were you using pastry mode or were you actually or was this rendered or generated uh footage? And if you cannot tell for sure if that was what you were seeing as you were interacting with and moving around the world was generated or it was u kind of pass through mode and was just what was happening in front of you. That's an indication that the models have become
1:11:23good enough. And we're not we're not close to that yet. And a lot of it is just this idea of really simulating dynamics and counterfactuals. Well, like you know, if you ask a video model to generate a person scoring a goal versus a person failing to score a goal, it would do a better job at scoring the goal because there is a bias from the training distribution. There's a lot more videos of the person succeeding at scoring the goal. But if you have an interactive model, you wanted to be able
1:11:51to generate counterfactuals like if I take this action versus this action, you wanted to generate equally realistic outcomes. Uh so that's I think the big gap between video models and and and world models is that idea of the counterfactual generation. And if you want a great model for robotics, you want to simulate failure very well cuz whether you're using it for evaluation or you're using it as a in an online RL loop in the future, you want to be able to have the model kind of try and fail
1:12:20to do things and and improve. Um and so in order to do that, you need to be able to simulate things failing. This is the only domain where you have too many successful examples and not enough bad examples.
1:12:33It should be easy to generate failure.
1:12:36Oddly enough, I think like early image video models weren't good at being human realistic, right? Like you see a lot of the high-res 4K like professional photography, but not just everyday life like normal picture, right? everything looks like it's professionally generated like professional pictures but not just like normal like you know messy cables on a desk.
1:12:57Okay. So there's there's this stuff uh you know one thing we also covered that you guys have video agents that you launched I guess how does the traditional let's call it frontier like you know auto reggressive LMS like feed in you know to to all this they're driving uh your robotics models or they're driving other your video agent production um any anything where you you see the overlap of auto reggressive and diffusion let's call it yeah so harnesses are really important
1:13:27across all those different use cases. So we have this video agent which is essentially an LLM that is very effective a tool use of different uh image models, video models and kind of helps you through creating a a project end to end. So you know very often in like a traditional kind of advertising flow you have a brief you start from and then you generate um a storyboard and then you generate the video. A video agent and runway agent kind of helps you through that whole process and it helps
1:13:55you also analyze performance data. For example, how well did this ad perform versus this ad and then generate me more of the based on those learnings figure out what what what to generate. We think that the harness is a very important piece of the the pipeline. uh as I mentioned all the video production all the all the production video models use some prompt completion that happens and we expect you know that to become more and more complex and more you know you
1:14:23generate longer and more detailed descriptions before you use the the the diffusion transformer I do think eventually you know there's increasingly this unification into omni models where you have the you're training the models end to end to both do auto reggressive text prediction and also So uh diffusion as well. So you're predicting the next token uh of like you do maybe doing some reasoning and planning of the scene and then you're passing it into the diffusion head that's actually
1:14:52generating uh generating the pixels. Yeah, I I think currently maybe only Gemini and Quen do it. I'm not sure who which of the Chinese models are omni but yeah it's it's not it's not a very well um popularized modality. I guess it's an interesting use case when you think about it, right? Because not only do you have to end at like language model reason diffusion head generate you don't have to output there you can go back
1:15:19into feed that output to the same model reason again on improvements and it can do a lot of loops just in its own. Uh I guess the question is like do we need that or can we just do agent scaffold like do it outside the model? Is there a big benefit to doing it in I think there's generally the trend of something is first done by a harness and then it becomes part of the model right so you had uh chain of thought prompting
1:15:48where you have to do this super detailed system prompts to to get back and now the the model basically generates the reasoning trace by itself before it gives you an answer and in the in video models similarly a lot of the video models of kind of the the early days were singleshot video models and you had to use some kind of orchestr ators, you turn, you know, generate multiple shots in parallel, um, and then turn it into an actual video. Um, comfy UI just all over the, you know, spaghetti
1:16:17workflow. Um, and now you have multi-shot video generation where you have the you directly generate multiple shots and there is a benefit to that because then the video model learns some you know to generate a single shot well you need obviously to figure out a lot of stuff about the world. uh to generate multi-shot video. Well, you also need to basically get some like video editing instincts like you need to figure out what is the right pacing of shots and also LMs are not that good at it. Like
1:16:47they're not that great video editors. If you ask a LM to take some videos and then kind of auto create a edited video out of that, uh it would it would feel uncanny. So, I don't think LMs are actually that good yet at being video editors. And I think there's benefit to learning that end to end. Um so I would expect you know the train in general is the things that you know you need a harness for eventually get kind of injected into the into the model itself
1:17:15and you learn that end to end. Do you find that you need to hire engineers who can or or researchers who are also artists to infuse that taste or do you have kind of artists and residents to distill them? We have we have a a a re a large creative team that's very actively involved in the in training those models um like on the you know in every part of the way and like how do you caption video well so that you capture the stuff that you need for like the
1:17:43cinematography the aesthetics the camera direction in as detailed ways as possible so that you're able at inference time to actually uh elicit that through the model you know we have our creative team also does a lot of evaluation of like you know what constitutes a usable video out of those models. And so they're very involved kind of through every part of the the process. And I think that's one of the special things of Runway is just that that mix between like kind of creatives and researchers kind of sitting by side
1:18:13by side and kind of working together to to build the next generation of our models. I think that's been really really important piece to to um you know how we've operated as a company.
1:18:21Yeah. In some senses you can only do this in New York. I mean you have other offices but like you know I'm trying to find some poetic uh significance in the fact that you are a big New York company.
1:18:33I there's a few parts to being New York. Obviously there is that intersection of all those different industries and like media kind of advertising uh like this is very advertising the like the art scene is New York. Not not to not to say anything bad about the San Francisco but uh it's there's more more going on.
1:18:52There is that component. There's also I think we benefit from being outsiders and thinking of things a bit differently like not being in the same like hive mind of uh ASI uh of uh of Bay Area and like taking and also taking our time to you know to get we are where we are today like building the growing the team intentionally and bringing people who yeah both on the creative side and also on the engineering research side there's obviously huge talent pool of amazing
1:19:21people in New York so that that hasn't really been been a problem.
1:19:25I mean, congrats on everything. Uh what are you hiring for? You know, what should people look forward to uh for the future of Runway?
1:19:33We're hiring across the board. Uh I think this is probably the most open roles we ever had in the history of Runway. Uh we're growing our research team quite significantly. So, if you're um if you're excited about video models, if you're excited about world models, if you're excited especially about robotics, the robotics team, we're hiring also robotics across uh kind of software, hardware, and research. Um so, definitely definitely reach out and and you know, a lot of people don't have direct robotics background, but what should they have, you know, if they
1:20:03want to be useful in robotics? So ideally some experience with learned policies uh would be else uh good for for robotics. Uh but we we tend to hire generalists as a as a philosophy and like uh people who learn really really quickly. Uh but some experience and and kind of domain expertise in robotics is something that we're we're definitely looking for uh for the for the next months. Um and then we're scaling the the go to market team significantly.
1:20:32There is um a wide like very very active enterprise adoption happening around video models at the moment and uh we're we're really trying to uh respond to all the demand.
1:20:44Yeah, great. You want to talk about the uh open source robotic stuff?
1:20:49Sure. It was just random notes we had.
1:20:51Uh Nvidia launched Cosmo. I guess it's interesting. So uh your founding member AI labs to build open-source world models in physical AI. Um anything much to talk on here is open research. The biggest thing is that as I mentioned world models are still uh early like there's still so much that we you can scale and those models further so much more advancements and and things that we can figure out or how how to improve those models further and I think this is
1:21:20uh it's important that some of this research happens in the open and figuring out what is some incentives for different companies to come together to to actually bring some of that research into into the open and open source and so Cosmos Coalition was a initiative that we co-ounded with Nvidia to bring some of that research as open source and that could mean openweight model releases. It could mean benchmarks that measure physics and things that people care about when building world models.
1:21:50It could mean infrastructure. So really how do we grow the ecosystem of world models and make that something that also it's easier for a developer, a researcher that's just starting out that is excited about world models to kind of contribute to the field. I think is there there's some amount of like is this also our response against the Chinese world models that are being released you know or uh is that not part of the consideration I do think it's it's important for in video models you
1:22:18know if you look at the the leaderboards of video models I would say right now the majority of models at the you know the top 10 to top 20 are Chinese models there is uh only a handful companies that are made to the leaderboard from like the US or the West.
1:22:34We're doing better with images, but with video we're very behind, right?
1:22:37And so I think I think it's definitely important that we invest more broadly as a community to make sure that we we can those models can, you know, we have competitive models out there.
1:22:47Well, like what's to stop us from distilling from them?
1:22:50I don't know if that's the best long-term bounded by the performance that you can it's almost a bit of a pessimistic uh you know view that you can get better, you know, you can it's free data. like it's you know you might as well like if they're doing it for for the text language side they might as well do it for the video side the other way.
1:23:10Yeah. I mean I do think we're we're quite capable of training great models without without this solution at the moment. Yeah.
1:23:16Yeah. Yeah. So anything you have to say on benchmarks and evals like I feel like what I'm hearing is a lot of people really like arenas for video and image models. uh customers and whatnot as well. They only want the best on the leaderboard and they refer to arenas a lot more than language models seem to do. But any any notes on benchmarks, what's lacking? How does the average person compare? Well, these both look really hyper realistic. More than that, outside of we did talk about like
1:23:45robotics, simulation, um the physics and all that, but anything to say?
1:23:49I actually think it's the opposite in some ways. I think people generally creatives and artists and marketers like people that are using our platforms I I think rely less on uh arena scores and it's it's just so easy to you know generate with a bunch of different models and then compare the results visually. Like one nice thing about image and video models is you can immediately tell with your eyes like what what feels good from an aesthetic standpoint like any artifacts any issues with the physics of those models you can
1:24:18immediately tell. Um and so that's it's actually easier I would say to evaluate uh as a human. Uh there is also those models than it is in in language models where you have those those very complex kind of math and coding and uh uh tests where it it becomes a lot more harder I think for for humans to evaluate and discriminate between the the performance of our tier models at a time. So I think in practice people just test out the
1:24:47same prompt with a bunch of different models and see what the results look like and right now in runway you can use our models and you can use third party models as well. So it's it's very easy to do that.
1:24:56Amazing. We're going to end with the AI runway AI summit. The the last sort of so societal issue I guess I don't know if if this is a thing is the uh you are at the tension between sort of artists creatives and AI. A lot of people in that that community hate AI. Obviously the people that are in the runway community don't mind using tools. Uh it's just another brush. But um how have you seen the sentiment change?
1:25:22I mean our perspective yes it's just another branch brush. It's just another camera. It's you know it's the latest of the a long generation of of tools technology and art and technology have kind of evolved together. I think there's been a pretty significant shift over the past few months and it came some of it you can see with a lot of public figures speaking out in favor of of AI and being you know like in K you saw a few a few directors speaking in
1:25:50favor of AI we had Ron Howard in our film festival there is Mark Suchza also adopting AI models so you have more of those stories coming out every day of like a a well-known figure um kind of speaking in favor of AI and it's just a matter of in in my mind it's those models are becoming more and more demystified. I would say I have I have also a a bit of a hot take that one of the things that made the the initial
1:26:19response to those models maybe a bit more heated than it needed to be was this idea of text to video of you know you have a single text description and you get back a a full video. Um yeah there there was a misconception obviously you can generate it to our feature line film but the models of today now take a lot of references they take they are very controllable and I think when people see a tool that's allows affords many degrees of freedom
1:26:48and control they respond to it differently and it matters less that it's a generative model than um than the fact that you can actually steer it to the direction that you want. Um and so I think when people look at you know complex workflow on top on top of those models when they look at you know all the ways in which you can steer them and you can provide now with some of the latest models up to 50 references like the conversation becomes a bit different
1:27:17because it feels much more like a story a tool versus like something that a magical kind of entity that figures out like the your entire film for you. any notes on like workflows changing for people in the field like I think engineering at least has had a lot of people where they're like expectations have changed you know 10x 100x more productive and you can get a lot more done uh same thing is you know you're making dev tools for creatives um any notes there
1:27:46like there's some people that don't want to adopt some that do like anything yeah so I think in terms of like what people care about uh I see that We've gone through a few stages. So we started from the stage where the main thing that people were looking for was quality like as you know we scaled those models to improve the quality improved dramatically. That's something that people still care about but it's it's now in addition to controllability like being able to steer those models with references with um you know different
1:28:16kinds of inputs with storyboards. And now my sense is increasingly people are going to care about latency more and more. As those models become better, the ability to iterate very quickly becomes more important and like if you can, you know, with a single prompt generate 10 different uh outputs like almost instantly, you can explore way faster than before and you get some of the magic that characterized the you know the creative tools of the past like Photoshop was instant. Um, and we we lost some of that with generative
1:28:45models. You you're waiting for 2 minutes to get back in video. And I think we're going to bring some of that back now with the with the the real time models.
1:28:54Exciting and exciting. Uh, the last thing we'll plug is this one. Uh, runway summit. You're finally doing this in SF.
1:29:02Yeah. So, uh, we're very excited about this. So this is uh in late September, September 30th, we're doing a summit on uh primarily focused on physical AI and real-time video generation. We have panelists from Nvidia, physical intelligence, the botco, deep mind.
1:29:18Yeah, it's going to be I think a very interesting series of conversations. We try to make the the panels really technical and uh elicit actual kind of substantive discussion and hopefully some some interesting kind of disagreements and interesting debates on on things and uh yeah the there's tickets available. Hope hope people can join.
1:29:40Since you mentioned it uh what kind of disagreements and debates should people think about or do you expect? So it seems like you know there there is a one debate right now in the robotics world is uh VAS versus world action models.
1:29:56Um so there is labs that are really really betting on one of those two directions. Um there is like what is the best source of data to train robotics models.
1:30:06There's just the third party first party that we talked about.
1:30:08Yeah. There there is you know the people who really believe in kind of f further scaling teop data versus leveraging more large scale video data. So that those are kind of some of the and then there is you know the the world models debates of predict pixels directly versus something like ja versus a more 3D based uh 3D based approach. Um so I think we're at a nice time in world models because there is still that kind of active debate happening or like what is the best long-term direction. I feel
1:30:38very strongly that this video predict pixels directly and and scaling video generation models is the right approach but it's I think there is a lot of interesting uh debate happening uh by by researchers on like what is the best best best uh kind of path to take. It's interesting that it's all on like sort of let's call it the policy layer and the data model layer is the the physical side is completely solved like all the sensors all the actuators all these things they're we we have everything that we need
1:31:07I don't think that's uh solved either definitely um different problem you know like it's like I want to dream about all these things and then I get uh you know I I buy a robot or I buy I try to assemble my own and I can't even get the you know the the the motors to like work Right. Right. Again, it's you're dealing with very sensitive uh equipment that has uh voltage and power and like heat and all these things which you know abstracted the way we're sitting here we're talking about software and talking
1:31:37about models but like really you have to deal with those kinds of things too.
1:31:41Yeah. And uh I think I'm I'm generally also not um not opposed to incorporating other modalities into our models like we've seen.
1:31:49The simplest case is they can generate video and audio at the same time. So they can generate RGB and they can also generate uh they can generate sound and audio. But my you know I I've written about this as like what what does the maximalist version of a world model look like is you're incorporating more and more modalities from the universe and you're training a model on different scales of observations as well. So yeah, you got a good uh good essay that people should read on real
1:32:17world. Meta released a model that was like six modalities in one, right? I forget what the name of the the thing was, but it was like Yeah. Okay. Depth is one of them, but that is like a transformation of RGB and image bind.
1:32:29Image bind. Yes. What other modalities?
1:32:31They had heat, audio depth, heat text, uh whatever IMU is. Um I do think like you might as well do ultraviolet. you might as well do like whatever other modality you feel like because it's all data to the model.
1:32:45Yeah. And a big bet is also that there is transfer between all those modalities. So one of my favorite uh examples which is quite old at this at this point is there was this finetune of stable diffusion that was called refusion which was the music.
1:32:59Yeah. Just fine-tuning uh stable diffusion on specttograms and became a quite capable music generator. Right. Uh there is probably some spatial patterns or like spatial temporal patterns if we're talking about video that kind of emerge the different scales and different modalities. And so there is some degree of you know metal learning that the model has done that that allows you to learn faster if you if you start from a just a model on images and train it to predict audio than if you train
1:33:27from scratch on just just audio. Um and there is some other interesting examples. So there is this this project called the well it's uh it's a data set of physics numerical simulations in physics and biology and a bunch of other domains. So it's so it's essentially different physical systems across very different scales of space and time from like astrophysics to low-level kind of um like atomistic interactions. And
1:33:56we've seen uh we've done some some some some work on this and and we've seen that we can take our video model where you know real world video looks nothing like this and you can actually fine-tune it on on those numerical simulations and just treat them as RGB frames and you actually get reasonable performance much quicker than if you just train from scratch.
1:34:20Yeah, I think I think we've seen this across languages.
1:34:22Yeah, Deep SQL CR as well. Yeah, Deepse OCR like you don't have to tokenize text like you can just throw them in as images. There's a lot that happens in that base pre-training like there was an argument a long time ago of people saying oh humans have so many uh sensory representations right smell touch models have a whole two more modalities that will never like that we don't even have data for and it's like okay you take AQI sensor like you can you can try this stuff but actually you you know there's so much happening in just the base train
1:34:51run that you don't get as much from these little things. Yeah, exactly. And I think that's what it solves is data scarcity. So you don't have as much, you have so much video data available, but you don't have, you know, like all factory data.
1:35:05The cool thing is it goes the other way too, right? So if you want to do physics, like if you want to measure this or you want to have a diffusion model do audio, uh, it transfers really well. So like in your case, the little bit of post-training for robotics gets a video model to use its fundamentals in another domain. So we can apply that to other stuff too.
1:35:24Yeah. And if we look at you know like how do you make those models more useful for in scientific domains and if you know if you look at alpha fold it had all these very because of the data you know the limited amount of data that it needed to be trained on it basically it's very fine-tuned architecture just to solve uh kind of uh protein structure prediction. But if you take all those disparate sources of scientific data and you bring them together under a single
1:35:54model like I think that's an approach that can help us kind of solve solve new kinds of problems across across science by leveraging all the all the learnings from one modality or one set of data to to another. So very early days for for that direction but I do think that that's where ultimately what the endgame of kind of simulating the world is.
1:36:16You're not just using RGB, you're using RGB as a starting point, but you can incorporate more and more modalities of the universe and and leverage the transfer that happens from learning from one to the other.
1:36:26I guess the the followup there is what's the drawback of omni like why why is everything not an omni model? So why not now? And why would you start from language backbone or image video backbone and then go omni from there?
1:36:41Yeah, we need to take it one step. We have to solve robotics first and then we can go into solve everything now.
1:36:49Yeah. I mean I do think there is a lot of open-ended research that needs to happen for uh those omni models. There is you know there is a lot of things that require careful consideration when you're bringing multiple modalities into a single model to predict but I I think you know I expect those to be to be solvable.
1:37:07Wonderful. Uh you've been very generous with your time. Congrats on all your success and uh yeah, I'm excited for the uh AI summit uh or physical AI summit.
1:37:16Yeah, thanks for having me.
1:37:18And yeah, and people should check out the film festival if it's in town, right? Uh you'll be going to be touring all over the place.
1:37:23Yeah, ne next year we're probably going to do the So we do film festivals every May or June of uh and we did the last one in New York, LA, Tokyo, and and at the AI engineer uh fair.
1:37:38Yeah. Yeah. Yeah. Um, so yeah, hopefully even more place next year.
1:37:41No, I I think like as someday, you know, you will be hosting the Oscars of AI video and you know, I think people should like take this very seriously as like a potential career they can have.
1:37:51The Oscars of AI video will be called the Oscars.
1:37:57All right. Thank you.