0:00This going to be like the biggest biggest market in tech ever. A lot of companies are making routers because it's fashionable. The model labs have several incentives to go after you eventually. Today we have Alex Atala, co-founder and CEO of Open Routter, the unified interface, the gateway to the world of LLMs. They reportedly have had offers from Stripe for $10 billion.
0:21They've raised at a valuation of over a billion and a half. They are the market leader and this interview could not come at a more precient time. In July, we launched 70 models. About one model every 10 hours. America is very very behind. Still GLM 5.2 was a really big big step for openweight models. There are reports that you are selling to stripe for $10 billion. Is that going to happen? Ready to go.
0:59Alex, I am so excited for this dude. I have wanted to make this one happen for a while. I've heard so many things from Mad at Menllo. I've stalked the [ __ ] out of you speaking to Anonyie, even your roommate before this show. Um, so thank you for joining me, dude.
1:13Thank you. It's great to be here. Now I want to start with a little bit pre-open ruda and start on open c. It was a pretty incredible journey. What did you take with you to open ruda having seen all that you saw with open c? Yeah. So, OpenC uh we started as the first NFT marketplace and uh similar to open router it was very small for a long time. Um like we kept the team very
1:42small until the series A roughly or you know a little bit afterwards. Um and this was before AI. So, uh, right after NFTTS started blowing up in in 2020, October of 2020, we were like, "Oh my goodness, like we are underst staffed. Um, the servers are melting. All kinds of like our search index was exploding.
2:10Um, we had a couple big outages. It was tough to like keep the site up and it was like, oh my god, we're going to become like the Twitter fail whale but like applied to crypto." My biggest goal was to have us not be the Twitter fail whale um for crypto. And uh and it took a little bit to like create the team, get platform and infrastructure under control u like make sure we we could sca we could predictably scale. In other
2:40words, like do load testing to like help the site sustain 10x load um even when we weren't seeing that load. Because with crypto, you just don't know. There were like these moments where we would get these incredible traffic spikes and it would be very dependent on the content and the community and uh so I built like a lot of um infrastructure and scaling I think responsibilities then that I took to open router and spent a lot of time like thinking about
3:09okay how do we you know make something that is going to basically be always up and that that people can really count on from an infrastructure point of view. um even when there are huge surges in in really like tumultuous markets. Um, which has been very helpful for AI of course because like you know all companies like especially in thropic
3:36have seen like unpredictable growth and uh and we have as well and you know we've had like a couple bumps but overall it's been like significantly better and I like OpenC just kind of like drilled that into me in a way where I could like take it productively to open router. Can I ask you when you go back to the founding thesis of the company what has happened in the ecosystem in the model landscape that you did not expect to happen? Um okay well one thing
4:06that we did not expect was that a an ecosystem of companies would emerge to host and serve the openw weight models.
4:18Um like early on it wasn't clear that that that market wasn't going to be a monopoly where like just you know the three hyperscalers serve all the open weight models and uh and startups don't you know they're they're really far behind in reality like you know how often do you hear people running you know GLM on a hyperscaler never like they're using the the inference providers like fireworks and together
4:48And um there's a you know big list that we that we see doing the best job of hosting all of the the open weight models. And um in the early days we um we had I think we called it provider one and and provider fallback.
5:07We didn't like show which providers were actually doing the hosting cuz I mean what we weren't really a marketplace. So we were kind of a like we were an exploration tool for like finding and discovering new LLMs and we wanted to get we wanted to build like a marketplace of model labs but like the inference provider layer we weren't sure would actually be a marketplace and it turned out that those companies were doing a way better job than the hyperscalers were way faster to host the
5:36models and figure out these edge cases to hosting them and um and uptime was just going to be a a constant problem. It wasn't going to like magically get solved by the supply side of the market.
5:47A lot of people suggest that that inference provider layer is a commoditizable element or layer that will be removed or see margin reduction competed out over time. What would you say to that theory? Right now we're in a massively supply constrained market where um the and it's likely going to be supply constrained for a while where all the inference providers are are short
6:13pretty much constantly short uh and and you're like okay so GPUs are are really really beneficial and like why doesn't Google or Amazon or Azure run around and like buy up all the GPUs and take all these inference providers out of business. Well, the people making the GPUs don't want that. Like, one of Nvidia's top priorities is not having customer concentration. They want lots of customers to all have like separate
6:41like allocations of GPUs. Um, they want the the heterogeneity of the market.
6:49They want like competition on the compute layer. Um, and I and this is good for the ecosystem like like users also want this. This is it's good for Nvidia and it's good for end users as well. It like allows these inference providers to kind of like come up with new innovations on like how to serve the models better. Even a single model like Kimmy K3 um uh like Moonshot just posted a benchmark showing all the inference providers and how well they're serving
7:17Kimmy K3. Um, and uh, the they're they're pretty different numbers for like for benchmarks that are really static that are wellnown. We post this continuously all the time. We always are like benchmarking all of the models on all of the inference providers, all the openweight providers and finding really different results constantly. The results change over time. Um, these
7:44models are like very they're very emotional. They're they're very like they're very like uh non-deterministic.
7:54So, uh um I had Lynn on the show from Fireworks and she said that, you know, I said about Gavin Baker and a token is a token is what he said and she kind of corrected me that a token is not a token actually because one provider can make a token go so much further than another token. It's like how do you get to the store? Well, you can drive around the whole block or you can drive straight to the store. Tokens can be made more efficient and go further and that's the job of the provider.
8:23Yeah, I I um I agree with that. I think that, you know, in some ways we are providing a service to help people discover providers and um and and like ultimately when when one provider is making a token go further, we we spend an enormous amount of time on our router central router tech so that that provider immediately gets more traffic
8:51as soon as as soon as we detect that like there's a quality improvement or a speed up or a price reduction happening um immediately starts getting more traffic. Uh and this stuff happens like 24/7 every single like every every 5 minutes there are big changes for the big models. Um and so it like actually does make the experience better.
9:15You can only invest in one inference provider. Which one do you invest in?
9:20I probably have to stay, you know, stay neutral on this. I I do I really like the the the you know inference providers that are doing um that are doing like custom hardware and uh and very very like low-level optimizations. Um I like providers that are also trying to figure out how to make um customization
9:48easier. So like today you fine-tune models and um and you create this like new like fully independent model from the from the base model. Um, many in inference providers are kind of like, you know, creating these Lauras or some some call them like cartridges that are much more portable potentially between models and and we might see a future where like when you do a fine tune and you want to
10:16like change the base model layer, it only costs like maybe a few hundred, maybe a few dozen dollars to change it.
10:23It's okay. I understood that fire worse is your favorite. It's okay. I get it. Uh, mine too. Uh my question is when Lynn was on the show, she was like, "Oh, you don't want to rent your own rent intelligence. You want to own it." And we're going to see companies have specialized models which is trained on their own data and proprietary to them.
10:40In a world of every company having specialized models that's really tuned to them and their preferences, is that good for an open router business or not?
10:55Oh, definitely. I mean our because you'd stick on one model which is yours proprietary trained on yours and not be open to the diaspora of models that is available.
11:07Uh no I disagree. I I I think our mission from the very beginning has been to increase neurodeiversity and AI for the whole ecosystem. And we really believe that like a multimodel future is inevitable. And when you start, let's say like let's say there's one model that like you know hypothetically let's say you're right. Let's say there's one model that fulfills all of your desires um either within your company or like as
11:34a consumer. Um and every you know more and more people start using that model and then um someone decides you know what I'm going to like create a neurode divergent model. I'm going to create a model that's like a little bit different that like talks a little differently that has ideas that the the first model like could never have come up with cuz it's like completely different data um that's being used to train it. Then it kind of creates inevitable demand to use
12:02both models. But creativity is not a um it's not verifiable.
12:09There's there's you can't really put an easy number on on creative ideas. And when you use two models together um you're more likely to get creative ideas than if you just use one. It's it's just a fact if if that other model was trained in a different way on that different data set like or has like made a big update. So um consolidation on one model just seems like it just doesn't make any sense to me.
12:35Totally get you. So you'll have companies which have like a core workflow or their core which is their own specialized model and then they'll use a plethora of other models and they'll use open router for those other model selection.
12:47Yes. And and I think that when companies make like to get back to your question when they make their own you know their own model trained on their own data um like you are the the ecosystem around you is all doing the same thing. You have to like play out the game theory for these things a little bit. Like if everybody is doing this as well and create and all the model apps are creating new models constantly using new data that they've acquired that they've bought from other companies that's all
13:16like potentially data that's valuable to you. What is in your best interests?
13:21It's to go and try out those other models and like see if you can be more productive with them. If you can like you know merge them together to get better state-of-the-art performance. If you can reduce your costs using these other models, they're all if you whether your goal is to like reduce your or um improve your margins or grow your company like you are incentivized to go use what the ecosystem creates. So the the the model that you made, you're
13:50going to have to continuously improve it to keep up and it's never going to win the whole market. This is going to be a massive market. This is going to be like the biggest biggest market in tech ever.
14:00and uh biggest market probably in human history. No one's going to win all of it. Um you're not going to build a model that wins all of it. So you might as well build a model that like is known to specialize in something very useful and that's very important to your company and your business and be known for that specialty. And I think a lot of enterprises are going to move that direction, make their own models, make their own um branded intelligence. So your brand is a big part of your moat
14:29and that model will like be a way your brand carries around.
14:33You mentioned the immense time that you spend on the routing technology that you have. Um a lot of people are thinking that we're seeing the commoditization of the routing technology. You're seeing ramp release products like this. I mentioned earlier of merge a company we invested in release that product. Um several are releasing kind of routing technology similar or claiming to be similar. Are we seeing the commoditization of this layer?
14:59I think a lot Yeah, a lot of companies are making routers because it's fashionable. Um, I think they're, you know, they're seeing growth happen here and uh or they're making gateways at least. Um, first, I think there's two there's two issues with that. First, it immediately puts you in the mindset of copying instead of like, you know, winning something. um you're sort of you're playing to play. You're playing to exist
15:29rather than playing to win. Um and uh and maybe you're just trying to like play to serve your your existing customer base. Uh and you you want to see some AI growth happen. Um I think you know immediately kind of like puts that gateway like many many months behind um the companies that are fully focused on it. like I am 100% focused on building the best router and gateway and and LLM marketplace um and it shows in
15:58our product and um and you know the the benchmarks that we create internally and how we see ourselves compared to the competition. Um this is not a side quest for us like it it may be for some other companies. Um the other problem is that it uh it reduces the leverage of all of your users. So um like I really deeply believe in giving users and developers more leverage. Like fundamentally giving
16:28them access to more models is about giving them more leverage over all the innovations that happen in AI. You want to be able to like access them all. You want to reduce your dependency on any individual one. if um you know you build on top of a a router or a gateway that you know doesn't give you access to the full market or full flexibility or full customizability.
16:55Um it doesn't give you like the the full leverage of the whole ecosystem, then you're you're kind of like being cut out. You're cutting out all your employees at your company of things that they they need. And so like open router is fundamentally about giving people more choice because that gives them more leverage.
17:11You do that at a price at 5.5% take.
17:16That was sort of our our pay go plan. Um we then added an enterprise plan with like a totally different pricing model and um it's been very successful so far.
17:28You it's kind of based on like committed spend and then uh you know no fees on that committed spend. Cuz that was going to be my question. Ultimately, companies will like love it small and then as you scale like [ __ ] this is really freaking expensive. I'll just build my own rooting tech now because it's become such a significant part of my cost base actually. I mean, I kind of think it like some of those companies don't, you know, just haven't realized we have like
17:56an enterprise plan and some of them uh and some some of it is like our fault for not having like a better I think more detailed pricing model. We're soon going to introduce like a kind of a business uh self-s served plan that um that also just you know makes makes it make a lot more sense. And if you if you have your own inference like if you bring your own inference to open router if you bring your own keys um that fee
18:24goes away. So it's a fairly like you know for inference that we are providing you like when you go into open routers capacity and you're not on our enterprise plan that's when that fee comes in um otherwise like you know it we need to be able to like predict demand a little bit. So that's why we do these committed spend. What will be the main revenue line of open router in three years time?
18:48I mean I think it's going to depend on on the economy in so many ways. If the overall AI market keeps growing the way it's been growing over the next four years, you know, you know, with like 10 to 15x every year um or potentially more. It's a lot of growth. Um, you know, I think under that
19:17under that world, um, I would expect people to continue to underestimate how much inference they're going to need.
19:27And thus our our revenue is going to be dominated by, you know, the same things that dominated today that dominate it today, which is like in, you know, like us helping people with with an unplanned inference capacity. um both enterprises and startups and uh like that that's what Open Router is best at like when you when you need to try models that you weren't expecting you need to try when you're like um
19:56using more inference than you thought you were going to use on particular models like we make we make sure that that is not going to be an issue for your company um by providing the best failover and best uptime and and this is really really a good thing to do when the market is like continuous continuously underestimating its inference needs and growing at this rate. Um if this uh growth rate continues over the next four years um I
20:25mean it's going to be a a wild amount of growth like the economy and the economy has some limits to it. I can see like major SMB you know SAS like growing for us. um and needing to grow for us whenever uh you know you know if you know if growth like does not keep going 10x 12 15x per year
20:52we we've seen token prices fall 90% give people or take in like 18 months is the reduction of token prices helpful or hurtful to your business because obviously you have a take on spend if they come down and spend is more efficient seemingly it's bad for your business you have shrinking pie to to take from.
21:13Well, a lot of people talk about the Jevans paradox that, you know, when when uh prices go down by 10x, the usage increases by more than 10x. Um, but like no one has really done a great job modeling it there. Uh, we do have a lot of spot stories that confirm it. For example, uh, GBT 5.6 6 Luna on Open Router. Um,
21:41OpenAI cut prices by 5x and then in coordination with us by another 2x. So in total price has the price of Luna has dropped 10x on open router over the last 2 weeks. And guess how much usage has grown? 13x.
22:01So, it's a close to perfect Jevans paradox story where you drop prices 10x and usage grows by more than 10x um just a bit more. And uh and the also the the usage is pretty stable. Like it like grew, you know, it like flattened out but at at 13x and then, you know, it's been kind of like growing at the same rate that it was growing before it hit the the 13x multiple. Um, so that's
22:31pretty interesting and it's a pretty like uh low variable like there there few other confounding variables in the story and and it was also done in the middle of Deep Seek launching and having a really really good price and GLM having a really good price. Like now Luna is being used more than GLM on open router. GLM used to be like one of the top like three four models by token volume and now Luna is past it. This is
23:00the first time OpenAI has had a model on our platform in the top you know three to five models by token volume in an extremely long time. So it was a really big and interesting move.
23:15How reflective of the market are your token volumes? because it's about I I'm may get this wrong about maybe one and a half 2% of say token volumes and so how reflective are they because a lot of people when I say oh the top five models when I look at open router are all Chinese what does that mean they'll go oh well Harry no offense to open router but like it's not reflective of the market and most people who use frontier it doesn't go through that like they use frontier APIs and so it's
23:45not counted to what extent are your rankings refct active of true token usage.
23:50We try to estimate how they're off um by, you know, just surveying people. Sometimes we're looking at like the surveys other people have done. Um I think we we have a we definitely have a bias to people who believe our thesis which is that the future is multimodel and companies who want multiple models and there are still
24:18companies out there I basically rarely very rarely run into them now but there's still companies out there that are just like oh yeah we're an open AI shop like we've won you know we only do open AI models and so we're not going to see any of those companies Um and I think those companies are primarily focused on like the you know the hyperscalers, openi anthropic and gemini.
24:46Um so we do probably like underount the the frontier models. Um but I think like over time our thesis is becoming more and more common to see in other companies and the moment that like they're like oh yeah like we need to use other models. Um then our data becomes more representative and as we scale up the data becomes more representative in general. Um so my hope is that like that we that the it just becomes like better
25:15and better data over time. Can I ask you? Alex Cop said on CNBC in his rather wonderfully energetic way that companies are terrified of working with frontier model providers.
25:26Do you think they are?
25:28So I haven't seen what he has what he talked about there when I talked to our customers. Um, but there was definitely like a little there was some skittishness that the uh particularly when when Claude design came out uh around Figma and that part um I did see and I do think that there are like real concerns for a company that um is kind of building like a you know thin
26:07you know for like Figma is very different but if a like a a startup is only building um a like go to market wrapper around intelligence like hey we are you know we're a company that kind of like brings AI to this market and does so by like doing the right integrations and like customizing the system prompt um you're going to be fine if the model labs don't
26:36care about that market which there will be many markets like that. But the model apps have several incentives to go after you eventually.
26:45One is um getting multiple teams within companies they do care about to be dependent on them. So th this is my theory behind why like claw design was strategic. While it's not like a massive amount of revenue for enthropic, like a not probably not a significant amount of revenue, um it does get the design team to really care about enthropic models.
27:10And so the companies that like they want like they now have another team that really wants to stick to enthropic. So those team that team strategy uh makes this you know like can make you compete with the model labs. And so I think like companies like that that find themselves like oh we're like building a product for a team that has now become strategic for the model labs for like companies they actually care about um that's where I see probably the most
27:39near near-term threat.
27:40Do you think claw design will have a meaningful impact on the Figma business?
27:44I speak to many founders today who are bluntly switching from Figma to claw design and it's cannibalizing their Figma usage. Do you see that and do you think that will happen? So, I I I saw a lot of designers try out claw design, including our own. Um, but so far, I haven't heard of the repeat story. I don't know, honestly.
28:10Like, I have not talked to very many designers about this. Um, I certainly haven't heard a lot of chatter about claw design. Um, and like if you just look at the numbers for Figma, they're quite good. like very very incredible earnings. So, um this is why you don't want to be public, dude. You see like you great numbers, Figma down. I'm like, poor Dylan. Like, give what?
28:39Yeah, that was crazy.
28:41Do you know what I mean? Like, really?
28:42Come on. Um we were talking about like the the different models that we have on offer and where the company's willing to work with Frontier models. The rate of model development feels immense. Do you think we will see the same rate of model development continue over the next year, 2 years, 3 years?
29:03Frontier model development or general model a general model both frontier and open.
29:09Just cuz I mean every single day there's there's two, three, four new models.
29:13In July we launched 70 models.
29:18It's about one model every 10 hours.
29:21There's some agent labs starting too that are all that are all going to kind of like probably make models eventually like Jeff Dean is starting an agent lab right now um from Google. Uh the the companies that are like that are known for making agents have an incentive to create their own model, a very clear incentive to create their own models and distribute it through the agent. And and we haven't even seen the start of that.
29:47Like sorry we've seen the start of it but we haven't seen it really pick up like cognition has a model cursor has a model does lovable have a model yet I don't think so not publicly yeah so the agent labs are going to I think develop models this pressure from both the GPU uh makers like Nvidia to like create more competition in the
30:14space and create more diversity iversity in the in the space plus us plus investors who just want to try new things that like all could like improve intelligence in some neurode divergent way. Um I think those are those are strong incentives. I think that there's still there they're enough to like incentivize more founders to make Neolabs.
30:42Um and the and if like if American openw weight models pick up in steam then it gives these neolabs a base to train on that's not Chinese um which will then probably create more American neolabs.
31:00Do you think we should be concerned by the rate and quality of Chinese open models?
31:06We should. We're behind. America is very very behind still. Um I think things are picking up and I think I think um you know we we have poolside we have thinking machines we have RC. Do you feel a sense of responsibility for that? And what I mean by that is like you know you are a routing business and you could route a company to a Chinese model that who knows people are worried
31:35about back doors ch ICCP involvement you could be the the deliverer of that to those models. Do do you feel a sense of responsibility for that? So we we do feel a responsibility to have safe access for all of these models like um customer trust is like our you know paramount goal. Uh if if one of these models is unsafe to use you know
32:04generally considered unsafe we pull it from the platform. If there's like a way to use it in an unsafe way, I mean, there's a way to use like all the models in an unsafe way, then we believe in using technology to make it safe and to like work with the model labs themselves to figure out how they're doing it on their side so that we can be state-of-the-art or better. We spend an enormous amount of time um making sure
32:32that like that our practices like match what the best things that we're seeing coming out of the the labs or are better. Um and because we're like a very good because we're we're a way of like exploring all the models and finding them for the first time. Um we're a good focal point for like deploying safety measures across your whole company. For example, we have prompt injection protection. you can just turn it on and
32:59immediately flag prompts that look like prompt injection um that's trying to happen. Um we have PII redaction. We have uh we have like a couple different things that you can automatically just turn on with a click and and get an added safety layer on top of all of your inference. Um, and we build that so that enterprises feel like they can safely like deploy new models and that their their um, employees can try them out. I think of the models a little bit like
33:28the internet. You know, you can't you you can't just like ban the internet at your company because there there's some like bad things on the internet. Um, you can create guardrails and you should. You need to use AI to build the best possible guardrails that you can. Do you think you actually know what's going on within Moonshot or Alib Baba with Quan? Like these are these are incredibly secretive organizations in the depths of China.
34:01Can't pretend I know like what's going on inside of them. You know, as a US company, like we're going to follow like like the best practices of what happens in the US to make sure that we're not doing something irresponsible.
34:13What do you think US companies are more nervous of? frontier models or Chinese models.
34:18I think they're they're more nervous about frontier models usually part because there's just like a a much there's much more confusion around the data policy about what's like actually happening to the props that I'm sending and um where they're being stored and how they're being looked at. Um and you can't run them on your own machine or in a provider of your choice. Uh and so that just immediately creates all of
34:47this uncertainty in a lot of enterprises and it's uncertainty that they can also pattern match. It's very similar to like you know running on their own infra versus running in their VPC um and and knowing like who can see the data. Um, how extraordinary is that they? Like they're more nervous of like US companies headquartered in Silicon Valley where you can see and touch and feel the headquarters and the leaders.
35:16It's just like what a strange world to be in. Yeah, it it is very strange especially with these with the frontier models like doing the like having the biggest cyber posture right now and like the the being the best.
35:30What do you make of every company kind of posturing? Haha, we hacked someone. First you had Open AI, then you had Anthropic and then you had Zuck coming out. I don't want to miss the party. We did too.
35:40Yeah. Well, I think they have they they have to talk about it. Like the right thing to do is to reveal when there's been a cyber incident involving your model. Um covering it up doesn't work when it's not going to work in the long term. And it certainly looks like they're all bragging about it. Um, but really if you were in their position and you know something happened with one of
36:08the models um and you had to make the choice about whether to publish it or not like I think the right thing to do is to publish it regardless of what like how people are going to spin it. So, I don't know. I like I I doubt that that they're, you know, they're actually, you know, thinking of the felony bench or whatever it's called.
36:33How significant was the latest Kimmy model which got so much attention? Was it as significant as everyone thought?
36:42It's quite good. It's it's uh it's not cyber capable in the way the same way the frontier models are and long range long horizon tasks. I think it's still a bit behind um the frontier models, but GLM 5.2 was a really big big step for openweight models. Kimmy was kind of like moonshot
37:09getting up to that step. That's a little bit how I see it. Um, and Kimy's also a very good writer. Like the voice and tone are both pretty good. Um, whereas like some of the frontier models I think have like voice degradation that happens when they get better at coding.
37:30Especially I know it's like oh my god like the I can't I can't read this output anymore. the output sounds like three of the four arguments you made are right and one is a turning point, you know, and duh here's the rub like I it's just it sometimes just impossible to read what what they're saying. Um and and this stuff is fixable, but uh but Kimmy I think has always had
37:58pretty interesting writing. In 12 months, will the chasm between US open source and Chinese open source be bigger or smaller than it is today? So my fear is that it will be bigger because when you have deepse it becomes a national champion in China and I mean Xi Jinping is going this is our AI horse. I will concentrate all of my money and efforts behind this and I will supplement this ecosystem to the end. This is the winner. And then when
38:26you see another moonshot come out, suddenly all regulation gets moved aside, all policy gets pushed aside, all funding becomes available.
38:38Everything is allowed. You are free to run. And these guys are unabbridged in their ability to do whatever they want to get to the end goal.
38:51Whereas open AI and Anthropic and all and all the other providers in the US especially open source [ __ ] you going try raising billions of dollars for a US open source model bit tough actually um not impossible at all but tougher business model questionable uh AI research is super expensive and you're competing against open anthropic I think the comparative landscapes they sit in mean that the Chinese open source providers are just inherently advantaged sadly they have like very very good
39:19researchers And I think Americans underestimate that a lot. I do think they're going to be concerned about the cyber posture of their models and they do seem very concerned about like censoring the models and and censoring the um the information that the models can provide to people. So, you know, while today people complain about American models censoring more due to cyber,
39:48I'm not sure that's always going to hold. And as like the Chinese models like grow in importance for China, I mean, what what are they going to do?
39:59Are they going to like drop the Great Firewall? Are they going to like give up on on putting the firewall around the models? like like I I don't know that much about China, but it does seem like kind of strange that they don't seem to care more that the models are I've never seen anyone do a profile of like what you can do with deepseek that you can't do with the internet in China that's available to you within the border like what information you can access. I've never seen anyone kind of like do a real
40:27deep dive. Like how past how far past the firewall does deep seat go. If the firewall matters to China, if it's going to matter in 10 years, like that's going to something's going to change.
40:39Well, what's interesting is like obviously the abilities of the Chinese models outside of China is immense. The abilities of the Chinese models inside China is actually relatively limited.
40:47The guardrails, the guardrails are incredibly stringent and prohibitive. So it's ironic that they are incredibly superior to us [ __ ] domestically. Terrible.
40:59I literally just had my dear friend Jason like you run SAS to come back and be like couldn't figure out what time Starbucks opened on Deep Seek like wasn't on offer would say like not allowed.
41:12Wild very basic rudimentary requests. Um we're speaking about all of these different models and the thing I think is like What about loyalty? And you have this incredible seat in the ecosystem where you can see everything.
41:29Do we see any developer loyalty today with models?
41:34Honestly, we do see some. We, you know, we we try to make switching costs close to zero so that when new models come out, um, people can try them out really easily. But we also measure retention and churn from all the models. We share this with this data with model labs too when they ask for it. Um so they can know like oh you know for my model that just came out like which models drove
42:01traffic to it and like for those users like when they leave which models are they leaving to? Um and we'll like make this more and more uh available to the to the world um soon. And and we do notice in the churn data there are developers who kind of like continuously stick to models even when there are better models out there better models for their use cases. Um I think it's a combination of like a couple probably
42:30root factors. One is like my app works and I don't want to break it. you know, if the support bots start saying something weird that I didn't expect, like why why add more headache? I've already done all this optimization and like I've already put all these guardrails around it. Um, another is uh new new models are not necessarily going to make your pricing better. In
42:57fact, um, in general, what happens is that the current models like price goes down over time and especially when new advancements in in, uh, in the labs happen, you know, you you'll see like intelligence jump, but like the price curve like also jumps and then we'll start going down over time. So, it's not necessarily the most like price effective thing to do to like shift over to the the the newest model even for open weights. The third reason is
43:26I it's just fundamental like trust in the outputs. Like if I'm using a model to do my work um and I like the way it talks um I probably have like some eval like a personal eval. A lot of people have these personal evals that are just these random tests that they give the models and if the random test doesn't look really good on the new model it'll just be like good I liked you know Kimmy K2.6 Anyway, people thought before that memory would
43:56be the retentive mechanism. Well, OpenAI has all of my previous quer like prompts. It knows that I live in London.
44:04I do podcasting and that will make it a better model for me moving forward. Is memory no longer a retentive mechanism?
44:13Memory is really interesting. I um I've always thought it like is a retentive mechanism and the question is where it lives. Is it going to live with the model? Is it going to live with the inference provider? Is it going to live with the app? Is it going to live with the infrastructure provider, the the router? You know, my my guess is that all of those layers are going to try to own memory in different ways. Um there and there are going to be advantages to
44:42sticking your memory in each layer. You know, if you stick it with the app, then the memory has like the most app related context and is model agnostic. If you stick it with the model, the memory might perform the best on personalized benchmarks and perhaps have the best like ultimate intelligence. And I think the model labs are going to work on memory and and then
45:11the ultimate thing might be like is there a good combination like can I use memory in the model and memory at the infrastructure layer or the app layer at the same time like is that going to confuse the model? We don't know yet. I I do think that like it's impossible for one layer to capture all valuable memory because the apps own so much important context that the model labs don't have.
45:34Um and they the model labs in order to get this to work they'll have to incentivize the apps to like give them that context.
45:42Apps and the models that claw code cursor bundle model and harness is the routter absorbed into the agent framework before it ever has the chance to be independent when you have the agent and the harness together. The harnesses are pretty interesting because like in our early days, one of our early bets was that most apps were underestimating the
46:11desire for users to choose the model.
46:15Like most apps in the very early days in like 2023 and 2024, it wasn't even clear which model was being used under the hood. They were like, "Oh, people are not going to care about that. They they just want AI." Um, and our and one of our like strong convictions then was that no, like people are going to want to like use particular models. They're going to care about who they're talking to. It's like, you know, I want to know which employees I'm talking to when I'm trying to solve
46:44a problem. And models will be kind of like that. Um, and that has played out, you know, like in notion you can like choose the model that you you talk to.
46:53Um even though you would think an app like that might want to like obscure it completely. Um sim similar a similar thing happened with harnesses where particularly with developers um they started to build an affinity to different harnesses and uh and that's cuz it's like it's a user experience. So I think that is my favorite argument for why harnesses are going to stick around.
47:17um not that like they're being bundled with the models because in fact like as models get better they get more resourceful and the the junk that gets thrown in the system prompt just becomes a handicap. Um Enthropic I think published like a good a good uh article about this where they showed that like oh we got like we got rid of stuff from the system prompt and suddenly fewer contradictions showed up later on with user prompts and the model performed
47:44better. Um, and we and we're seeing like a lot of the harnesses right now are like deleting code in order to perform better with the latest frontier models.
47:55I don't think means that harnesses are bad. I in fact I think we'll see more harnesses come up in the future because it's a way of building a user experience on top of models. It's a way for developers who are not model labs to own a user relationship and that is just going to be, you know, incredibly valuable for the economy to have that layer. I'm going to get killed for this.
48:19What's the difference between a harness and an app? Feels like this word wank of like everyone's talking about harnesses and the harness. I'm like, is that not an app? Like hello. Yeah. The nice thing about the harnesses compared to the apps is that they're more composable.
48:35Like I can have a harness call another harness. I can have a harness spin up another harness in a sandbox in the cloud. Um is that what APIs did for apps?
48:46Yes, but uh it's much more reliable and and uh deterministic and sort of easy for users to gro harness because um the harnesses are Unix based. they all have and and the models are so well trained on on Unix on bash commands whereas like you know if I'm like telling a harness to go orchestrate an app in the cloud it's going to be like oh boy does this
49:14app like how do you log into this app is like do I need your password do I need a uh do I need to fire up a virtual browser it's going to be pretty slow I'll figure it out okay I I like fired up a browser and like now I need your password and I'm going to like try to find the input where to put it in. And I there's probably an API in this app somewhere. I need to like look up the docs to figure it out. And okay, now I've got the API, but there's so many
49:42like unknown unknowns when you're composing around an app. Very very very very few unknown unknowns when you're composing around a harness. So I I think it just gives developers more flexibility um and flexibility that they can inspect like API calls. you're just seeing a whole bunch of code flying around the screen. A harness, oh, I can like jump into the harness and like look at what's going on and talk in English about it. So, it's much more user friendly.
50:08We've seen Meta and Muse really be a focus for Zark. We've seen Alex Wang front and center much more. Were you impressed by what Meta delivered with Muse?
50:17They've been doing a good job. Yeah. Um, I mean like it takes a while to set up a whole new model lab from scratch and I'm sure a lot of like uh organizational debt to deal with. Um, you know, like do you think they will be a serious challenger?
50:35I do. I think they're I think they they're they have the resources. Um there's uh I think there's some competitive things they can do uh around the model that like helps people in ways that the the model labs are not as interested in doing like just having like a social network
51:02um and like a focus on people.
51:06uh you know it it's like something for the brand that maybe Grock and like spa like X a SpaceX AI have it too. They do need to find their niche like I'm not quite I think people don't quite know what to do with Muse Spark yet like when to use it or when to go for it or what it's like like core advantages like they just released a coding harness. Uh they are they're trying to be like a generally capable model right now. Um, I
51:35expect that in the future they're going to be like, "Look, we are way better at this thing." And that's that's going to be a really important moment for them.
51:43Fantastic. I said I was impressed by it actually. You know what I use now? Maybe plugging one of our our mutual friends, but Anastasio and Arena. And it's so weird. So, I'll put my prompt in Arena and then obviously it comes back with a load of different model options.
51:58And you know, I come back with I used one the other day, Pergamom.
52:03Yeah. And it it was like Kimmy and Perom and they offers you four different options and it takes me to models that I would never have used before. And actually Muse has come up a couple of times being pretty impressive.
52:14But I love that in terms of this like discovery mechanism to models that I would never have used. I would never go to Kimmy. Honestly, dude, I just go [ __ ] chat GPT.
52:24It's really interesting.
52:25Yeah. I um it basically though goes to the point of the model layer just becoming a utility layer.
52:31What do you mean by that? But actually, I have no loyalty to them.
52:34I have no affiliation with brand. I go to Arena and I want to see what you got for me.
52:42Show me the results.
52:44I don't care if it's Kimmy or Muse or Claude or Sonnet or whatever. Do you know what I mean?
52:51And actually, I just want to see the options you got and I'll pick the best from there.
52:56I'd rather run four in parallel. Do you buy this whole we're going to have uh one frontier model run four open models and the frontier model might be 160 IQ points and the open models might be 120 IQ points but that will be a a model infrastructure structure that we'll work with totally think that that is a great architecture that everybody needs to explore um and we prov like we've been helping lots of
53:24developers do this where like you you have sub agents. We have a sub agent server tool um that we like tune to be really really good at using models generally. Um and then you have a an orchestrator model that calls out to the sub aents when it wants particular tasks to get done. And these uh these sub aents are just very very low cost and they're focused on deterministic tasks.
53:50This is what open weight models are generally really good at compared to frontier models. When you have a deterministic task where you know the shape of the output, you know the type of problem that you're working on and it's a type of problem that has been solved like classifying some text for example, then you should definitely use like a lowcost like you know model from open router and uh and then have the orchestrator model read the results and
54:18then go and continue working on the like unknown non-deterministic task that it was set out to do. I want to create a open American ecosystem. Yeah. On more amazing open American models and I make you head of this program. What would you do to encourage incentivize the open US ecosystem to compete more viciferously with the Chinese? I think I would spend
54:46time talking to the current American labs a little bit more to figure out what distilling the Chinese models looks like for them and how effective it is.
55:00Um, you can you can probably get pretty far distilling the Chinese models. Um the nice thing about the openw weight models uh and uh and the Chinese models that they allow distillation and they like most of them and that means that you can like take the outputs of these models to do reinforcement learning on top of the model that you're building and and you know this is just like a very important
55:28and common practice in AI that all labs do. Um, so you know, like it's it's one I I I would want to learn a little bit more about like how effective it is. Um, but it's one way to catch up uh with the openw weight models um because they allow it.
55:50The other thing is they uh the the American models and also when you distill you see the output so you can like inspect them to make sure that they're aligned. So if there's anything about the like you know the open weight models that you're worried about um not being aligned with like the voice or constitution of the model you're creating um you have a much better shot at catching it when you're doing these RL rollouts. Uh the the
56:20the other thing I would try to figure out is um is the compute question. like compute is just a huge advantage that I think we still have relative to China and these Neolabs um need a shot and there like there needs to be an easier way to like get compute to the right talent
56:47um you know in all countries um but like especially you if we're trying to create like a competitive American neolab system you know like Nvidia's been doing a good job of this, but there, you know, there's like there's Google, there's TPUs, there's Tranium from Amazon. I would like work with all of the hardware companies and also like the super like, you know, the the the Neo chips um to help with compute.
57:16I don't think we will have that compute advantage for long. I think you see Deep Seek and Bite Dance both aggressively pursuing their own chips now. The export controls mean that they have to and this is like the number one problem for Xi Jinping in his race to win the AI war.
57:32I agree. I build a bridge in four weeks. I think they'll manage a chip in 6 months.
57:37Yeah, it's it like staying ahead on the on the chip war is critical for America. Um is distillation wrong? I mean distillation is it's a technique to build models like you know but people view it with cynicism and shade. Well they just distilled models.
57:57It's a technique to build models. The closed weight model labs distill models too like like you know sonnet is a partially distilled version of opus and like you you this is how you like make smaller models out of bigger models.
58:11Yeah, it's an important way to just teach your model new things when when you find uh like something useful in the ecosystem. I we do think that labs have a right to say it's not allowed in their terms of service. A company can like cut off access to someone who is trying to build a competitive model. Um if you're just trying to build like a smaller model that's like really focused on doing one specific thing, that's not competitive.
58:41Um, most of the frontier labs don't prohibit that to my knowledge, but the there are going to be markets for companies that allow it and companies that don't. And um, and we we make sure that we help like both companies uphold their terms of service.
58:56I have to ask you one question before we do a quick fire round. I'm going to get killed if I don't ask it.
59:02There are reports that you are selling to Stripe for $10 billion. Is that going to happen? I can't can't comment but you know like whatever happens we're we're going to execute on the vision what we're doing is critical for the ecosystem and and we believe for like safe access to AI where you know one monopoly doesn't take over and uh where
59:31we have like a vibrant ecosystem of of models that are that are that everyone can explore and and when new providers and new server tools and new inference adjacent like tech comes online, there's a really easy way to discover it and connect it with all of your um with all of your existing AI.
59:49I was thinking these situations like my response would be like, well, I own 22% of the company, $10 billion, $2.2 billion. Oo, now I'm a venture capitalist. But like is it hard not to think like that?
1:00:02I don't really think about it. Um do you know I don't spend a lot personally. what what I think about when I what I do with like personal capital, I I really want to help people work on problems that are not that just don't lend themselves very well to venture capital. They're sort of falling in this gray area of problems that like people need to solve but are really
1:00:30tough to fund cuz they don't come with a business model attached. And I think they're like very cool things to do now in in the nonprofit space because you can use AI to review way more data than you ever could before. Um I'm not quite ready to talk about it publicly yet, but I do want to like do something that um that like helps researchers work on those problems and like get grants to do it. One really cool example I think of
1:00:59this is David Fialcow who's one of the founders of General Catalyst who basically finds incredible stories that won't get funded for movies and funds them to shine a light on them because he thinks they're very important. So like the dissident which you know obviously um told the story of Kosogi and Kosogi being you know and then you know Icarus which is the story of the Russian doping and like these were films that would not get funded had it not been for his funding because they were politically
1:01:28sensitive charged and he's like I'm going to enable the stories of these forbidden tales.
1:01:35Yeah, it's kind of Yeah, kind of like that. I love I love that stuff.
1:01:38He's great. He's [ __ ] awesome.
1:01:40Anyway, are you ready for a quick fire round?
1:01:43Okay, so what is the most underrated model on open router today?
1:01:50Good one. I mean, first I like poolides models are great. I um that's probably like my fire round answer. good like New American Lab um building interesting coding models that are very they're they're small um but highly effective and uh um and they're like like building a lot of useful tools for accessing them um good team.
1:02:1670% of Neolabs will die in the next 3 years. Agree or disagree?
1:02:21Disagree. 70 seems very high of Neols. There aren't that many Neolabs. if like getting acquired by one of the model labs counts as die. Uh I do think there'll probably be some like potential consolidation.
1:02:37Um if you if you include the consolidation I I would put I would say 50.
1:02:45Do you think Dario should be less negative and more positive as a voice in AI?
1:02:50I think it's important to have somebody who is very paranoid about the future and and how things are going to shake up. And I appreciate that like I personally appreciate anthropics paranoia. Obviously, there are there are areas where like I want like other model labs to um not feel like they're just being like
1:03:18pushed off the table. But I'm a big believer in in neurodeiversity and like enthropic, you know, is a part of the neurodeiversity map that really matters.
1:03:30And if um if no one is being extremely paranoid, then you know like no one is like offering that voice. Um and so I appreciate that they're doing it. What's the craziest thing that you see in your seat on top of everyone's usage that you don't think people talk about enough?
1:03:50I mean, a lot of companies are obviously worried about cost management and uh and freaking out about the amount of inference they're spending and they don't know how to think about it. It's like a whole new way of of like doing business and thinking about your your opex. Like the old way of thinking about how your how much you give your employees, you like give them a salary and you kind of forget about it. Like someone knows what what everyone's
1:04:18making, but like it's a static number that like gets readjusted on on a quarterly basis maybe after performance reviews. really your your employees all cost totally dynamic different amounts now and uh I think a lot of like companies are putting it on them to do routing and I think in the future there's a good chance that it will like get pushed downwards to the employee level. Your employees should like figure out which
1:04:48tools and models to use that are best for their tasks and then we should figure out how much you're costing be like due to the choices that you make as an employee and you know your your cost as an employee is going to be a dynamic number and it's going to be you know dependent on how much that employee is like effectively using you know expensive and cheap models to do their job and uh and then I would I I advise
1:05:16companies to kind of like still do their normal management work, like have their managers kind of assess how how effective and productive employees are, but also line it up with how much their employees cost and then kind of come up with, you know, a quadrant of of celebration. Like these employees are like doing a good job and they're pretty price effective or cost effective and then a quadrant of concern. these employees are kind of maybe doing a so-s so job and whoa they are not
1:05:44cost-effective at all their AI use their AI psychosis is off the charts like and then you address the quadrant of concern so I don't think people talk about like basically how you think of like employee cost in the age of AI and that it's it's really it should be a dynamic number and not a a static thing that like only a few people know about and it's gone wonderful but can you imagine going to someone oh I'm sorry you were worth 100 and last month now you're worth 50. Uh I think it would make planning
1:06:13well they are in control of how much they cost. That's the great thing like all employees are in control of how much they cost and and can like influence that. Um it's now you get to think like okay how good am I as an employee and how uh efficient am I being as well.
1:06:33Final one when you look at the landscape today there are so many things to be excited about. What are you singly most excited about?
1:06:41One is rare disease research, which I think is one of those things that has been intelligence bottlenecked or or really just the inference bottlenecked like it involves like trying out lots of ideas and seeing if they work. Um the other is crowdsourcing productive urban life improvements.
1:07:06So, for example, like imagine if somebody had like someone was curious about finding every lead pipe in America or every lead pipe in the UK and like had a an approach to it, but they really need like to make it mature um and and stress test it. Like now you can use AI to do that and we just might solve some weird problems that everyone's just kind of given up on
1:07:35because like you you need like a crazy idea to come from somewhere like brilliant ideas are sort of evenly distributed all over the world like they they they can come from anywhere. Um and now you just give them leverage to actually work. So, I'm excited about sort of very like broad kind of urban um like or urban or or rural like quality of life improvements that we'll be able to make.
1:08:03Dude, I've wanted to do this one for a while. I'm so glad we could do it in person as well. I was worried that we were going to have to do it remote. It is so much nicer to do it in person. You've been fantastic. So, thank you so much for doing it with me.
1:08:13Likewise. This was great.