0:00Hey, there's a new type of unit of value that's being streamed across the internet called a token. And over the next 10 years, the entire internet value chain was going to have to deal with the fact that like the more valuable tokens got, the more bad actors were going to go to try to get their hands on those tokens. And anytime you scale something and the payload gets more and more valuable, more bad things people try to
0:29get access to that value. Online payments, you know, is started roughly in the 80s and '9s, right? And grew to over a trillion dollars over the next 10 years, and we needed to build entirely new payment solutions to deal with online fraud. um where we are today is roughly there on tokens, but over the next even five years, we're expecting the token economy to get to like roughly $5 trillion. And over the next 10 years, I'd be shocked if we went to 10 trillion
0:58of token flow. We we blocked 10x as much dollar volume last month as the month before, and the types of token fraud are diversifying quite a bit. Before we get into today's episode, I just have a small message for listeners. Thank you. We would not be able to bring you the AI engineering, science, and entertainment content that you so clearly want if you didn't choose to also click in and tune into our content.
1:25We've been approached by sponsors on an almost daily basis. But fortunately, enough of you actually subscribe to us to keep all this sustainable without ads, and we want to keep it that way.
1:35But I just have one favor to ask all of you. The single most powerful, completely free thing you can do is to click that subscribe button. It's the only thing I'll ever ask of you, and it means absolutely everything to me and my team that works so hard to bring the Inspace to you each and every week. If you do it, I promise you, we'll never stop working to make the show even better. Now, let's get into it.
2:00Okay, we are here in Anj's house, which is where all great startups in San Francisco start.
2:07U and congrats on cursor mist. Uh I don't god knows what else. You got so much stuff going on.
2:15There's there's a lot going on. Well, open router is probably the most has been the most I would say one I'm excited about recently.
2:22Yeah. And we have Alex uh first time on the pod, but uh you've been a a few times. I appreciate every time you shown up uh for the community. Congrats. I just like what a journey. When I was looking back at your past posts, one of the earliest principles that I saw you write as a sort of product person is pub sub as a product principle. And I wanted to you to maybe explain how you think about what should exist in the in the world.
2:47Yeah. The the pub sub piece which was early 2023. I didn't think about it until we we we talked like 10 minutes ago is about how there there is like a way of thinking about products as an intersection between subscribing to data and publishing data and marketplaces are an easy easy example of this. You have suppliers that are publishing some kind of product to a skew. And the skew is
3:15kind of like a pub subtopic that a consumer is subscribing to and just going to like consume whenever they want. And humans consume in a very like discreet ad hoc way. It's not very scalable. You know, all their attention is on the topic when they're buying the thing and their attention is nowhere else when that happens.
3:35um agents and and consumers of inference don't act like that. They're consuming continuously and they're changing the SKUs that they consume from all the time. So open router is sort of like a blend between a you know a normal API experience and a marketplace where we create model slug. We have the auto router. We have all kinds of like product SKs that you can subscribe to and and then you can like continuously add like uh derive value and make
4:05decisions based on those uh those consumers.
4:09Yeah, this is something that was more consensus now but not consensus when you guys started which was that there is such a demand for swapping models and changing things out and uh that people would not use use the native SDKs. I guess uh for each of you, what was your sort of realization moment that this would this would be it? I you've you've given a talk at EIE about alpaka as like one of your inspiring moments?
4:33Alpaca I can like rehash the alpaca moment for a sec. Like the very beginning at the end of 2022, OpenAI was the only game in town. There was like OpenAI Coher um and then a a smattering of of like early attempts at openweight models. Um, when Llama came out in January of 2023, it was like, "Wow, really exciting. This is really big. It outperforms GPT3 on one or two benchmarks." Uh, but you can't chat with
5:03it. It wasn't like an it wasn't actually an engaging model. But it seemed like someone just needed to fix a couple things and do some RHF on it to get it all the way there. And Alpaca was the first model that I saw that did that. It only took $600 to do. A team at Stanford generated a bunch of synthetic data, fine-tuned llama, and made alpaca 7 billion parameter model or was it maybe it was 13 billion parameters and it was
5:32so good like I was just like on an airplane using it. I I you know in many cases I like you could not discern a chat GPT versus an alpaca result and I figured if if it was this easy to make a model um one we have a whole new way of monetizing data for the first time um you can just like take really valuable data and turn it into a service in $600 uh and that cost will probably go down
5:59over time. When you say so sorry uh when you say monetizing your data as uh what eventually will become an FCP endpoint or or as a training data for a model.
6:09Yeah. Training data for a model like an abstract way of saying like hey I have this data compress it into a model like it makes sense for me in my product but like I could repackage it in the form of a model and sell it. And so it's just a whole new business model for the economy. It also of course provides like you know a way of following what frontier labs are doing but in a way that like a single developer or a small
6:39team of developers can roll on their own. And so that whenever you have an example of that like a breakout app that's doing really well and then some kind of framework for imitating it with your in your own flavor. You have an immediate ecosystem of like like an an immediate ecosystem like should arise because there's just a huge gap between the like decisions that the single company is making and all of the variations in those decisions that like
7:08a wider ecosystem can can create themselves. And so then you know you need a marketplace to like discover all of those uh services and all of those products. There wasn't any place on the internet that like was like a home base for LLMs in terms of seeing how much they were being used and seeing who was using them and why.
7:27The closest would be Hugging Face.
7:28Hugging Face was the closest.
7:29They just started hub like a few years ago before before that.
7:32Yeah. And Hugging Face also didn't have the close source models.
7:36Um and they didn't uh you couldn't use the models at the time. Um and there wasn't data about who was using that. There there there are like a bunch of differences between open router and hugging face and those differences felt really critical to me especially when I was just trying to learn about LLMs and like why people are choosing like these different little ones that are emerging over time.
8:01Got it. And then an no stranger to wanting more model diversity uh at the time uh you know you're a couple years into your anthropic journey which we covered in a previous podcast as well.
8:11What was your introduction to Alex?
8:14Well, the introduction was I think 13 years before that. Oh.
8:18Um, but the open router handshake actually happened right over there if you remember. Yeah.
8:21Um, which was Alex and I uh met I I believe it's sophomores now if I remember correct review um meeting for the first time.
8:32Yeah. So Stanford Review was the libertarian newspaper on campus at Stamford that Peter Teal had started back in the day. And uh forever for whatever reason I you know Alex and I both showed up to one of the meetings and I remember um the editor-inchief was a mutual friend of ours. Lisa was really a really great editor-in chief. You know part of part part of an editor-in chief's job is to assign responsibilities to people and make sure the work gets done. Um, and I I I may be misremembering the details, but I I
9:02remember wanting to it was kind of surprising to me that at the time there was no dedicated technology section in the newspaper.
9:09Um, you know, because it's political, right?
9:14Yes. Yeah. States and things.
9:16Yeah. But it it to take us back in time, you may remember this, but um there was there was this technology uh kind of legislation that was being debated called uh the net neutrality act. And net neutrality is like inherently this political concept, right? It's it's about the regulation of internet broadband access. And so there was a community of us who are kind of technologists but also debating the politics of the technology. And I thought the review would be a great place for to like write about that. And
9:45I was working on I think a net neutrality article and I remember proposing well maybe you should start a technology kind of section and Alex was one of the only people who said yes that would be cool and said I I forget whether we ended up writing stuff together but that's when we first met um was 2011 or 12 I forget which year it was. It was one of those.
10:06Um it was at Old Union if I remember correctly. That's where we used to meet. But, you know, along the way, Alex and I have had a chance to to to hang out often. And probably the the the time when we had the most professional overlap was when I was running the platform of Discord. Um, and it had become this explosive kind of platform for crypto.
10:31And NFDs in the middle of the pandemic, which also, by the way, you were in charge of safety and security as well, right? I was the head of platform which meant all of the crypto DAO and NFD launch security debugging fell on me and the fishing the the the social engineering attacks like a ton of DOS that we were getting hit by um around the time I started teaching security at Stanford CS53 and Alex was in the at openc at the time and I was trying to
10:59figure out how we we could defend against all these attacks that we were like and at at peak I forget if you remember how much NFT volume was running through discord but it was like a meaningful amount of like se it's like several billion dollars in NFT volume of GMV so to speak running through the platform and it was all coming from openc it was these like buy sell trade I mean the server the D and D is discord yes and so that's when I think we had hung out professionally but a year after that
11:27openai gave Discord early access to GPT sorry GPD3 no it was GP3.5 actually yeah GPD 3 which is the RL version of GPD3 and that's around the time we we made a Discord bot with um OpenAI for internal deployment and that's when I I realized we would need like since I since I was part of the deployment team what was the use case there were two that were and there there's actually a post now called Discord is your place for AI with friends that somebody sent me recently
11:56that I wrote um and published in 2023 but there were two use cases one was Clyde which was that in like a first party friend inside of Discord that could help you set up your Discord server and talk to you about onboarding and get your friends to hang out more.
12:14Um, and then there was content moderation and one of the realizations we had with content moderation was the it would refuse to moderate like it would just refuse our prompts because the RL the post training was we were very early in the post- training era and it would just our prompts would trigger it. uh it's like guardrails. And we told OpenAI, hey guys, we need access to the weights because if we're going to be doing content moderation at scale, we had 250 million monthly active users. We need more reliability that the model
12:44will do what we need it to. And they said, well, sorry guys, that's not how this works. We're a closed source company. And so that was my first realization that we needed open models and the enterprises would need more control o over capabilities and then ultimately would need some kind of control plane or management system to orchestrate these open models. But there weren't no good there were no good open alternatives until maybe 6 months later when Llama came out and 6 months after that I led the series A into Mistral
13:12which was started by Giam and the Llama team and I that around that time is when I remember hearing about Alex launching open router and going these worlds are going to collide and I don't know when it'll make sense to team up but Alex was so early and could see I think he was totally right about this ecosystem starting with Llama uh that then needed like a an easy layer to to manage for especially for I was approaching it from the enterprise perspective because I'd been that like the as the VP of platform
13:41at discord it was my job to ensure that when we deployed models to like 250 million users they did what we wanted them to and that was very hard um cuz if you outsourced it to the labs and they controlled the the guardrails and their guardrails or their safety policies forbid the model from responding to your prompts that was quite catast catastrophic.
14:03Yeah. But well, you know, a moderation is a thing that they want to support and obviously beyond that they would work open would work with you uh you know presumably to give you a moderation endpoint which they offer for free.
14:14It was an interesting use case um that they so they did give us a moderation endpoint. However, as you guys know, every Discord server is like a mini deployment of itself. And so the use case was instead of having human moderators that have to interpret the norms of the community, you just give the they often like every you know subreddit discord servers public ones have their own rules that the user the randomly space discord in yeah
14:40and then humans used to read those norms and then enforce it every day manually like observing each message in these communities and these communities are like millions of users. So we had a 5,000 plus person team globally in the on the Discord content moderation team.
14:56These are outsourced contractors who had a really tough job. And so the idea was instead if you could give the norms of that server to the LLM, then the LLM would do custom moderation for that server. It's almost like a like in context moderation for that server. And many of those servers norms just violated OpenAI's rules. And so that it it was like we had our own custom eval.
15:19So we used to like each server had its own custom eval but this at the time OpenAI's eval so primitive in our thinking about how to deploy these LMS that often the um the post training prompts were super heavy-handed. It said oh anything about Harry Potter anything that has trademarked content you know don't refuse. And it was a if it was a fan Harry Potter fan community this is a real use case that had a content moderation the LM would just refuse.
15:46Yeah. And that was just not precise enough.
15:49Another one that that we heard was like if someone was trying to write like a detective story and there's one chapter with a lot of violence, like maybe someone like kills someone, the LMS would just refuse to like help with that part of the story.
16:05And then they like the user would be like, "Okay, this this is not like structurally inherent to LLM. There must be like some choice out there so that I can like switch to another model." um when I'm getting like a refusal or a bad result from the the main one that I have and and that like tension also drove me for a marketplace.
16:26Yeah, I think that is well accepted now. What was it like back then when you were raising or you know starting this? Um did people get it? Um you know what was the some of the struggles? Basically I like I like getting stories out of him about how other VCs don't get it.
16:42So like anything anything you want to you want to uh talk about now you know now that let's let's call it the you know that the early journey of open router is done right you can obviously talk about some of the early days stuff.
16:52Uh well I was going to say that like the the biggest objection we got is is big model win which is all the value scaling loss. Yeah, scaling laws um and natural network effects are just going to kind of acrue to one company which will be like it'll be a Google style monopoly just like how Google won the search market um by a large large margin um and you'll just be fighting for it
17:20scraps at the end basically. That was probably the biggest objection we got and it is interesting that Google won the the the search engine race with such a huge margin. Um, you know, I think like had there been more interesting benchmarks or had like search engines been, you know, a bit, you know, have people like seen them a little bit more like LLMs where they're services that you can build companies on top of. That might not have been the case. Um, but
17:48LLMs don't merely have a user interface.
17:50They're also like ways of building entirely new businesses. And you know a Google level monopoly would be like the Dutch East India company times you know quadrillion in magnitude cuz the whole economy ends up like depending on the one monopoly as well. So it didn't seem like you know a like would be a really crazy outcome if that happened. And it's also less likely because the the economics of like creating good
18:18competitors are are much like much more decentralizable. Everything Alex said is true and I came at it from a completely different perspective which is yes this is why we're here. Um the scaling laws were never like in my mind were always a feature not a bug for why open router would be very valuable because I was one of the first investors in entropic and it was obvious to me that other researchers in our friends group and I went to grad school for machine learning and I just had a lot of
18:46friends in the ML community who it was a very obvious to us that the bitter lesson holds and so I was like oh like fantastic now we have at least two pro proof points that compute scaling works. It was open AI and anthropic. Um, and by the time I think we decided to team up on open router, I had already invested in Mistral and Blackforce Labs and Luma.
19:07So there was multiple model companies and teams that I was uh working with.
19:11But you did other modalities whereas this is different modalities.
19:16Exactly. And it was so obvious to me that an ecosystem of different kinds of models were being created. um and that this whole do narrative of like only one company will dominate like Google was um well like may be true but one I don't believe that but two there was so much extraordinary innovation happening across several different research teams but the shared problem I was noticing across all of them was often you know the research teams were fantastic at figuring out how to reason about new
19:45capabilities they think in terms of capabilities but never like are not developer mindset oriented like what happens s after the training is done and the checkpoint comes out like you'd be shocked how how like similar the early pre-training teams at OpenAI uh sorry Anthropic BFL Mistral uh were in in in their like default approach to taking their research out of the you know lab and kind of scaling their impact which was often oh the checkpoint is done put
20:13it out as an API done and then there'd be crickets um in in the case of claude the first cloud checkpoint And it was actually done a year before they released it internally. And then Chachi PT came out and we decided, okay, yes, it's a good idea to to release a cloud version externally. And they had no plan, like no no no plan for how to get developers to actually try it out. And so if you go to the Claude one blog post, you'll notice they like three kind
20:41of developer examples for users of the API. and one is a discord bot and the second is Vivian my wife's startup called Junior Learning because and then there was like notion um because these are all friends of like the entropic team because that's how like last minute the planning was around hey once the model's done training how do you get it out to the world there was no distribution platform that understood what developers needed all the the key management provisioning like simple like you know endpoint management versioning control like all these things that the
21:11scientists and researchers go I mean that's plumbing I don't detail, right?
21:14And instead, Alex came at it from that perspective. And so, you know, it was so obvious to me that like every single lab I was funding would would spend like literally sometimes billions of dollars into training and then a checkpoint would be done and there'd be crickets like doing early access cuz they're like, "Oh, that's right." Like it's hard to use a checkpoint to make anything.
21:35You actually need a whole bunch of plumbing around it to make it usable by a developer. And so by the time I think we it was so obvious to me that a distribution platform like open router was critical to have in the ecosystem.
21:46If we wanted there to be competition to Google like unless you know with Google Deepind is done training a new checkpoint and then they push a button and it gets blasted out across all their surfaces from Google Docs to you know everywhere even if I don't everywhere you want to know about like on Android like like overnight they can deploy a new checkpoint to like a billion devices right and that invisible infra advantage distribution advantage most people don't realize but until open router showed up it you had to think about all of that yourself as a model
22:15lab And it was very daunting, you know, at Anthropic, I think it took more more than 12 months to get to our first 10 million in revenue. And in contrast with with Black Force Labs, I remember the early days, you you guys had a conversation with the BFL team and uh it it was so simple for open to say, "Oh, no problem. Like the day you launch, we can send a million developers to you."
22:39You know, that that was crazy. That was like a step function change in in like power.
22:44Is that a real number? million.
22:45I I I think today it's like million. How many developers are on open router today? We over 10 but over 10 million but um but like it's it's hard to I don't know how to we do a lot of like you know account dduping work but you know no one if you could get a thousand developers just to put in context if you get a thousand developers to actually try the model on day one after you release it and just like do inference and give you feedback that's a thousand more
23:14developers than they knew how to get to on their own. Well, you know, BSL had a reputation. Yes, they had one with stable diffusion.
23:21And with Mistral, uh I don't know if you guys remember, but the first checkpoint they they released was like torrents. It was like torrent weights.
23:29Yeah. They just put up a magnet link.
23:30Yeah. There was no API cuz they didn't they weren't in for people, you know, like, okay, download these weights and you guys go.
23:37He has a story on his side. Yeah.
23:39Yeah. I I mean in addition to the like building a really good developer experience around it, the marketing that we do uh like for different models is totally different and perceived totally differently from the marketing that a model lab does for itself.
23:54Yes. We are like a, you know, neutral layer looking at this market like it's a big dark room with all the corners completely obscure to users and users were walking into the room and like feeling around and trying to figure out what objects to grab off the tables and like build into uh their companies. It's an insane way of working. Like models are not products where you can just enumerate all their features onto a web page. They're all black boxes, including
24:24the openweight ones. So, you need to like shine lights on all corners of this room um so that people can see what makes this model good. And you need the company shining that light to be a neutral third party, which is what we specialized in. So the like it's in addition to developer experience there's also like a very important like marketing and product packaging component and a way of like um routing
24:51and and discovering models becomes like critical to your go to market as a provider or a model lab or a server tool and more in the future and and this value to your earlier point about how many VCs like you know just don't one of my biggest frustrations is venture capitalists many of them like just don't have any operating experience in the field you know so unlike a traditional investor who's just maybe come up through the ranks as like a
25:19associate working on financial modeling or maybe hasn't been a real operator in the field for like more than 10 years which is a big part of the industry now I had just arrived at A16Z like a year after running the platform and so I knew what the challenges were of like building a real developer experience and actually like being able to create a working piece of software with with a model and there were a few I won't name names but there were investors who were
25:46looking at open router um and you know felt at the time like when I would compare notes with people that it was just I quote just a marketplace yeah just a thin layer just a proxy just a wrapper or whatever on other people's APIs and I was like you have no idea how strategic the value that open router has created by being able to orchestrate even three APIs in production. The amount of both engineering work and community design that goes into getting
26:16that actually live and running in production at the scale the open order team had started just doesn't happen by default. You know, and that was one of the things that stood out to me about Alex from the earliest days, like he just understood like these from a systems perspective, like how do you get these flywheels going? Like that stood out to me with OpenC when we were working together on the NFD integration at Discord. Like Alex had a level of community like systems thinking around how you get these flywheels going that most scientists and machine learning
26:44people just don't think think of. Like we often think in terms of pre-training, mid-raining, post- training, it's a linear stage. It's this linear pipeline. There's no loop.
26:52Yeah. It wasn't until much later that the modern context feedback loop cycle really got standardized in the industry. But at the time, if you remember, machine learning was like like mostly we did a lot of ML like when I was in grad school on a laptop. So you just like download a data set, ran some ablations, and you looked at the loss curves and you're like great, I made AI.
27:11And the idea that you have to like deploy those capabilities, collect feedback trajectories, then like put those into a continuous loop came much much much later. And it was very counterintuitive to the science like the traditional AI mindset. I do remember during the investment phase for um open router. I just didn't try and re-educate a bunch of other VCs on why it was not just a marketplace. I was like, you know what? I'm just going to invest and I'm going to like take the
27:40opportunity to partner with Alex and if not no other VCs get it that's totally fine cuz at the time it was not obvious I think to several other investors that like open router was not more than just a rapper around APIs and that infuriated me and I was like you know I don't have time to debate you I'm just we're going to we're going to invest and then I think like a month later Matt Murphy marked it up by 10x like our I think I forget what the exact post money was and so on but you know to his credit Menlo Ventures realized Okay, there's actually much more strategic value here as well. Maybe you
28:10didn't hear all these conversations behind the scenes. Um, but that that frustrated me a lot. You know, there's a lot of this like opining about rappers. Um, and if you're like, oh, an app is just a rapper on no model then like and you open router is like this rapper on top of other APIs. This is the most stupid reductive framework. So, it's clearly somebody who has no experience deploying.
28:32It's the thing you you dismiss other things with like you're everyone's a rapper on everything, right? Like and there's there's some point some some rappers have value.
28:38I mean investors are rappers and LPS right like venture capitalist. So I mean yeah it's all rappers down all down to bare metal I guess and when I started the whole engineer I guess the coining uh in 2023 like that was the number one push back is that this is no value. You should actually just train models right uh and yeah I mean obviously this is like you guys are one of testaments to the fact that like you can actually build very valuable rappers but also very valuable model companies. It's so
29:04um hard to be like the the day a model launches, the fact that you have an open router um endpoint for that model frequently at the top of hackernews on day one. People don't realize the amount of work that goes into accomplishing that. An open router used like that would happen over and over again and I remember going people have no idea how hard that is. you know that's not yeah we've covered some of the inference engineering that goes behind uh some of the with base 10 and all those
29:34well today you have you know all those like cool code name things and people guess what oxy alpha is and all those things but like I I I guess one of the things that you're teasing is how do you get that initial flywheel going right because today you have your scale and your reputation all these things so obviously you get you drive immense distribution but when you're early on when it's most the bootstrap yeah how what is the bootstrap like I mean to bring it back to early Discord days. I think we like initially connected with this is an open C story
30:03technically, but we initially connected when you were at Discord and we talked about like Axi the Axi Infinity server.
30:12This server was like the biggest server at the time at Discord.
30:16And you were kind of like constantly bumping up the limit.
30:20The limits on the server for those who don't know like like 10% of Philippines was actually I was on that server. That's it was like a meaningful contrib like crypto game. But there's like a Pokemon breeding thing similar. Yeah. There was battling, there was breeding and and then there was like a marketplace for trading, play to earn as well.
30:43Yeah. Play to earn. And like the graphics were really cute and fun and you kind of like you you you know you get kind of emotional about your axi that you make. So to like start a community like that uh which we had to do many times at OpenC with basically every early project for us to create a marketplace for it. We need to make sure that the like the community actually wants it. And it it's kind of like building something that people want and going and telling them about it. Like
31:12you can do that on a one-on-one basis, but it's way higher leverage to do that in a community where everyone can talk to you at the same time. So we we spent a lot of time like building things that the community really wanted. We did the same thing for open router and you know like the AXI community was one of like a zillion communities we did that with and an like saw us doing it and cuz you could just see people sharing open C links constantly in that discord. Like
31:40users sharing links is a really clear clear indicator that like something important is going on. So we spent you know a lot of time like first figuring out what the gap is in the technology that people care about like what was the actual problem that needs to be solved.
31:58You know in early LLM days it was you know open AI refusing to finish the the prompt or like to like complete the task. It was also you know inability to customize models. Um, and so there are communities that like are just completely blocked on that issue and those are the communities that are most useful to sort of learn about and and dive into and explore.
32:25Something that really struck me at that time you as as I was just hearing your talk, I remember noting how you may not remember this, but we were we had these like working uh Zoom calls that we're doing a sprint around for like this OpenC integration with Discord. Um, and you know, we we'd get it would it was myself, my engineering team, I think you were there. And I remember, you know, um, Alex in the middle of one of those calls just like there was like silence,
32:54uh, you know, we were we all like, "Oh, yeah, this totally makes sense. Let's do this." And then there's some like everybody aligned and Alex was like, "No, this makes no sense to me." And everyone's like, I remember going, "What? What?" Like that it works. like you click on a link and this then it bounces you out to like open C and he was like it's not a good user experience. Yeah, we should not do this.
33:16And I remember going, you know, he was the only one person out of all of us to actually raise his hand and go, yes, it made sense from a technical implementation perspective, like we were bouncing the user out into the into OpenC. And so it kind of checked the box of the product manager requirements on both sides. But Alex went one step further and was like, you know what would be better, guys, if we just embedded the experience right here inside of Discord. So the link opened up as an embedded iframe and you can just check out right there. And not one
33:46person on the call, like seven of us who had met like, you know, week after week.
33:50And it's the guy who doesn't work for Discord.
33:51And it's the guy who doesn't work for Discord.
33:53Like technically you benefit if they bounce.
33:55Exactly. And that was like adversarial to keep the user inside of Discord would be adversarial to OpenC. And yet Alex put that user experience first. And I was like, that's special.
34:06Cuz it's very hard to have somebody who's technical like Alex and understands the developer flow, but also understands the best user experience and wants to prioritize that. And that's two sides of the fly. that you can get spinning like is often hard to stop and you just reminded me like that one was one of those moments where I go I went I realized I got to be better at user experience cuz I should have been the one who came up with that and I didn't and I learned from you and um I think that went into one of our case studies for the PM training program at this I
34:33don't know if it's there because you need an Alex conclusion yeah yeah you need an Alex and and this is why I'm not you know nobody should be surprised why Stripe decided like they had to buy open router because it's a really rare combination of people who understand the machine learning community, the developer experience and the end user experience and putting all that together has resulted in this extraordinary scale that very few other marketplaces have been able to achieve over the last you know 5 years.
35:01Yeah. Well, we should talk about the other reasons for acquisitions uh which you've written about. Uh I want to sort of proceed somewhat chronologically as well. So that there there is a point that you know one of the questions that uh Dave from HFZ sent in was when did you know it start really started to work and you brought up Mix Draw I don't know if you want to bring up that.
35:20Oh yeah which obviously you overlap with so yeah thee was I don't know when I mean the there's no like one moment where I was like oh this is you know officially starting to work. It was like moment increasing like really super early on.
35:37Oh yeah. Well, yeah. So, before open router, I wanted to like explore a bring your own model experiment and um anyone familiar with crypto is like you know phantom and all these things.
35:48Yeah. Yeah. So, it felt like doing a meta mask analogy for AI would be kind of a fun way of exploring that. And at the time there were no AI apps. There were probably as many AI apps that were like hitting AI via like hitting an LLM via an API call as there were like games just doing it in JavaScript. You basically like there there was a there was a moment in time where it could have
36:18been the case that web apps call LLM through the browser like through some kind of desktop managed app that is controlled by the user. Um, and of course there are like I think many reasons that that did not happen, but back when the when the days were that primordial. I built a a a Chrome extension called window AI and with plasma which I had come across early on and I was like who's going to actually use this? You did
36:47plasma had a couple like I think Phantom was using it. Um, there were some other like like real companies basically react for Chrome extension. It compiles to all these kind of like Nex.js for and yeah built window AI on top of it. The creator of plasmo like started contributing code to window AI and uh in GitHub and that turned out to be Lewis Vichy who is the co-founder open router.
37:15That's you have told me this is how you met Lewis. Yes. Okay. So, um, that that allowed users to kind of like configure which model they wanted to use for a web page in their browser and then like the the app would just call out to that model when it needed to do things. You know, not the right form factor for LLMs, but you know, it's like fun experiment. you learn a lot and like I you know open sourced it and uh and you know the main learning is like okay this has to be an API and it has to look a
37:44little bit like there has to be more of a developer experience here and more of a discovery experience as well like I don't know where to use these models and a little Chrome extension is not going to help me discover it's not enough real estate I need more space I need visuals I need graphs I need you know examples I need images I need to like I need to be able to like explore both as a human and as an agent. So that's kind of how how open Rider came to be.
38:11You know, a meta point that I think is underappreciated, but Alex is reminding me is that we were quite lucky that we were so we were like adjacent to the crypto community in those days because in hindsight, crypto ended up being kind of like a dress rehearsal for generative models, right? If you if you think about the the AXI experience, uh you know, Alex is totally right. There were not that many AI apps at the time.
38:37And while I was dealing, you know, my job was to be the head of platform at Discord, which meant to be a general purpose place for communities and friends to create uh for developers to create apps and bots and you know, other services that could be deployed across Discord. And while 80% of the attention of the time was being spent on crypto because that's where all the NFT volume was, there was like 20% of my time of my time I was spending with a friend u who would get hotbot with me and ask me for we would play Magic the Gathering on weekends. Um and he was working on a
39:06little Discord bot that could take a text input and turn it into an image and it was called Midjourney.
39:12You know, was that David?
39:13That was David Holtz. He was a good friend and David and I have both been sort of failed ARV VR founders, you know, in the last before that. And um I remember this you midjourney was one of the fastest growing communities we had after axi infinity started to peter off and many of the the like the abstractions and the infrastructure decisions we made to scale Axi happened just in time because you know Axi did this and then fell off a cliff and then as Midjourney was taking out we like
39:41explicitly decided to help David make the server the midjourney server as the primary place for interaction with the with the model because it was very hard for people to understand how to use the model if they couldn't see other people using it and copy them. And so the single player Midjourney web app on its own like majour.com had like terrible retention cuz people would show up they'd see this empty field. It's kind of like Dolly 2 um and they would type in like cat or dog and it was like paralyzing for them to have this blank
40:10canvas that they had to fill because they never used an AI model before. but instead in a discord server you could see other people using it and riff off of their prompts and the engagement was off the charts and so scaling you know midjourney from zero to like 10 million monthly activives was a much smoother approach post Axi infinity and so don't forget the best of four pictures any which is the feedback loop the RHF feedback loop um which by the way separately like Tom Brown David and I used to play Magic the Gathering on weekends and so like it was one group of friends would hang out like these
40:40concepts were all being discussed all the time But you know there was I I I think there were few of us who bridged both the crypto worlds and the AI worlds and compared to crypto where it was the question was always what's the use case you know for this technology there was never any need to ask that for yeah cuz there use case was so visceral it was like I can create now anything my I can imagine I can write novels I can code and the infrastructure that those of us who believed in the distributed systems
41:09like value of crypto like the censorship resistance part found this use case that was explosive and I think between midjourney you know claude was a discord bot pre-launch you know that we were using internally as an LLM um 11 labs had a TTS model that we had on discord as well like discord became this petri dish for like early apps to innovate and I don't think it's a coincidence that they found a home there before open router gave the world like a public home
41:37store or like a um you know storefront discord was this like almost kind of petri dish storefront that was kind had kind of like piggybacked on the infra we'd built for the for crypto communities and then I think Alex was one of the first people to realize wait a minute like these apps need their own home um on the internet and then open router to me was a continuation of that of that community's needs and of course there was the crazy distribution that you enabled for a lot of these developers
42:05so so then my question is how come you were my perception is open router is not that discordcentric right you have a discord Yeah.
42:12And you use it to engage your community, but it's not like midjourney where like no that is like the primary way people experience open router.
42:18Yeah. Midjourney like it really helps to see visually really quickly how people are using the model and how to prompt it and I think that is partly why the server was so critical. It's like it is the user experience. It actually adds a ton.
42:34Yes. Um, and you can go the whole mile with just like prompting via midjourney like the via the the midjourney discord server, getting your images and then sharing them and having fun for open router for LLMs like you need a lot of user experience around LM to make them like really usable and yeah the like seeing the examples of other people is also not as useful because it's a lot of stuff to read. It takes a long long time. um you need like codebased
43:03integration not not possible to do in a discord server you need um or tech it's possible I shouldn't say that it's just not a great you great developer experience um you need like comp you need governance for at the point where you got codebase integration now you need governance for managing the LLMs that have access to it the data policies which teams all that stuff needs a lot more than than a Discord server can provide so it's just like it's not the right well in addition you're not wrong but
43:32also there's the very important distinction that you know midjourney was an enduser application right and you know that's why discord which just has 250 million monthly end consumers you know made it made sense for discord to to kind to be ao host for that application experience what I I knew was going to happen soon after midjourney found explosive product market fit because we I think when midjourney launched from launch 100
44:01million revenue run rate was less than eight months and shortly thereafter stable diffusion launched and you know all of us used to hang out in the discord server it was the um the stability discord uh it was the lion yeah lionage community that stable diffusion and so when stable diffusion came out I realized oh now other people can build their own midjourney because until then mjourney did not have an API so they were a full stack company right they were training their own
44:30models and they were deploying them as an application. But if you want to build your own mid journey, there was no API of that quality. Um and I think Dolly 2 was still quite primitive like midjourney actually had great quality and then when stable diffusion came out suddenly there was this new person who could there was new this new capability in the world which is a developer could create their own midjourney and that I think created the the need for something like open router because then you need an API to if you were if you had the kind of creativity of David Holes and you had stable diffusion as the model and you wanted to put these things
44:59together how could you do that without having to figure out how to host the weights and what open router uh the shape of open router enabled is is that right when you have open models alternatives to closed sort of applications open router's value in the world becomes extraordinary because now any developer can just show up and model did you just say the shape of open router oh no you this the real I I'm misaligned now I've
45:27been overtrained I've been using cloud way too much haven't I claudish is what people say claudish oh god I got untrained myself Okay. And I just want to cap off the mistral side. Uh my my my TLDDR is there was a mixture price war is what they called it, right? Like uh roundabout Europe was 2023 or four. They launched uh the MR 8 by 8 by7B uh and like the price went down like 80%. To me that's very positive because it's like the first like real competition to to host
45:57Mistral. Is there more?
45:59Yeah, that was I'm like trying to remember it. all the things that happened it like we saw that model come out and immediately saw people say that it was the best model in the world like this was to my knowledge the first time an open weights model was called that in real seriousness it's hype right is it you know it was it was hype it was hype it was also like hype from AI influencers at
46:28the time and there were many examples where it was like outperforming GPT4.
46:35So people really wanted to try it out and see is this going to be true for me too and if so at what price? And uh the the like inference landscape was really messy. Yes, we cleaned it up. Um the it allowed like providers to compete on price so we could give users the best price in one spot. And so it was I think the first clear example of like a provider marketplace working in a way that adds
47:04value to developers.
47:06Sean, you may not remember this, but I I think we met for the first time a few days after Mixtra came out at Nurups at a lunchon.
47:13Yeah, that's where I also met BFL as well. Yeah.
47:16And Gom was there. I was at Nurups at that time.
47:18You were there too. And um G we had just announced the Mistral investment and I remember Giam was over there and I remember turning to Guiam and asking him like is it is all like how are you feeling after the launch of Mixrol in 7B and you know him in his typical French fashion was like I mean it's a it's an okay model it's not that good and I was like it was so you know in contrast but I I remember him also saying that part of the reason he felt a lot of people
47:46thought that it was better than GPT4 before was because of the speed. You know, they it was an MOE model that they had like absolutely kind of figured out how to make super efficient. It was on the period frontier and and this is an important thing with LLM, right?
47:59Sometimes when they're faster, you think they're smarter. Um even though like if you did end of, you know, these common like eval do tries and I don't actually remember. I think we should go back and figure out what the data says, but I wouldn't be surprised if it turns out on an end of seven attempts, GPT4 was smarter on EVLs, but the perception of on on like or correctness would be smarter or more accurate, but you know,
48:26people like from a human preference perspective felt that it was faster because it was smarter because it's so fast.
48:34Um, and actually most queries do not take that level.
48:37Don't take that, right? This is the start of humans as router which then eventually becomes open router as router of like the auto you know what I mean like because humans are the routing mechanism like I will ask the fast model first and then if like oh not good enough I'm going to upgrade manually but then he's going to auto it I didn't I hadn't thought of it that way but that that I mean that makes sense which then there's there's a lot more techniques like fusion fusion is a thing that we should talk about before I move on to those things I just want to close off the sort of early early years uh one thing that I observe which you are also
49:06an investor in arena Right.
49:08And we talked about midjourney having that that feedback loop of ABCD and choosing that very being being very important. And you understand the flywheel. So how come you didn't build arena and how come Arena didn't build open router?
49:21Well, Arena started before open router, right?
49:25They had they had the school project and then they became a Marina. Yeah. So, but and I know you had some Arena experiences like the the heads up comparison type things, but you never really went as hard as Arena did and doing heads up experiences and and LMS actually did have a router project based on Alam Marina ELOS uh which they never commercialized.
49:46It's hard to do a company that does both because one company is taking data and selling it and the other company really can't by default. So you know I think there there is like a branding reason that there are two companies here. Um like when you set up open router there's no training or no prompts like aside from what your provider policies set like open like open router can't see
50:14your prompts or completions. If you want to see that as an org, you have to opt into it and enable it. And so we're like pretty conservative and careful about data policy and security and privacy.
50:27And Elm Marina is like their business model is is like oriented around the labs and and because they give it for free, right?
50:33You don't give it for free, they give it for free.
50:35Yeah. But I mean, we do give some we like have free endpoints too, but like those free endpoints, I think we're not collecting any of the prompts. We're not like monetizing the data unless you you know opt into it for some reason. Th this this comparison I mean you're not the first person to ask me this and Alex knows this but I you know I was the the inter like the founder like first C CEO of Arena for the first 5 months when we were helping Anastasia and Whan kind of spin out of Berkeley and uh I did invest
51:02in that before um open router but it was it was very strange to me the comparisons that outside you know folks would make between the two projects because the missions were completely different. the founding entity for Arena, we called it the AI reliability institute because it was it was actually there as an eval service like the data so to speak that they they were originally um kind of offering the labs was how do you make the evaluation of
51:30models more reliable than kind of like the state-of-the-art at the time which is like really just finger in the wind.
51:36Um that that's kind of what Anastasio and and Whan's PhD work was as as scientists at Berkeley was on statistical methodologies for sort of uh correcting you know eval estimates um based on like intrinsic biases and how you collected the data um and style control style control and stuff like that and which is very much like a hey how if if you're a scientist and you're trying to kind of um the highest expectation customer for Arena
52:03was always like a a post-training and uh like a a researcher at a lab. Whereas the highest expectation customer from from my perspective that that Alex like really understood and and was the mission was to serve was was like a developer, right? Who then takes the result of the research and then produces an application that's deployed to the world. It's actually a completely different problem and person that these two teams were focused on. And so from the outside in actually I don't know if
52:31you remember this but I have a distinct memory of a few weeks before we did the term sheet uh together for open router I'd given you a call because we were trying to get a pool's data set together from open router and from arena to uh create like an open- source repository of prompts.
52:49I mean these projects were so kind of different in their goals that it was totally normal to me to be like oh yeah let's call Alex and see if you'd want to team up on on pooling data cuz they're so different. we need we we actually don't have that kind of data at all. We like we didn't have API prompts. We we didn't have like what developers want to do with the models which is very different from what researchers inside a model lab want to do before releasing the model.
53:13Does does that make sense? And so to this day I I think you see that this difference even though at a 30,000 foot level you could I guess you could kind of conclude that arena and open router are adjacent but uh the road maps the missions and so on at the time at least were like in very different sort of directions.
53:34The ideal customer I get I I totally get that as a founder I want to own everything right. So like this is clearly the adjacency then I'm like I'm going to explore that. uh own everything meaning like you don't know what to do yet so you want to like make sure you catch PM I think what he say you you want to own the entire infrastructure space and so you'd kind of expand to whatever demand yeah I I think that's that's hard you know in reality because serving multiple customers is is
54:03clearly you know this is the only one of focus right yeah I I still think even in the age of AI like focus is is underrated and critical Not just because you end up with a better product by focusing your humans on it, but also because the world knows what your focus is.
54:21The world can map like, oh, I have this issue. Which brand out there is going to help me with that issue? This is the brand that's known for that focus. So like if I want real attention on this issue, like this really matters to me, I should go with the brand that cares the most about it. to underscore Alex's point about how important focus is. In the early days of Anthropic, it it was not easy to like people think that the early days of Anthropic were like super easy because they were on the GPT3 guys who left. But it was actually very
54:50competitive. The company was starting $10 billion behind OpenAI, right? And so to get to the frontier like the big question was what what do we want to be known for? What's the mission? And the mission was AGI pair programming. And so to the um exclusion of all kinds of other things that were really shiny at the time like image models and video models that were getting lots of you know momentum the anthropic team was like we just got to focus on coding like that is the core capability that we're focused and today you can see the results right it's a
55:19trillion dollar company within 5 years and that focus I think like the high the focus on who your highest expectation customer is and how you exceed their expectation because exceeding anyone's expectations is hard and doing it for multiple like customers is so even more difficult um is part of the reason why opener succeeded and entropic as well.
55:37Was the focus on coding that early though or did it come later? Literally from day one it was AI pair programming is responsibly commercialize an AI pair programmer was the seed memo. That was when I invested. Right. We we actually kind like refine that memo a lot. Well, you got to ask Dario and Tom for permission on that. Um but it's an extraordinary piece of writing that they had put together and AI you know commercializing responsibly commercializing AI pair program was the mission you know from day one. And I
56:05would say there was maybe like a couple moments in the company's history where like they did experiments to kind of see if like little detours made sense like a general chatbot like Cloudi when chat GPT was really taking off but at the end of the day especially once um they got their significant pre-training computer online I think like the all the main evals at the company for example have always been coding evals long horizon agentic programming I mean from day one that was always when like flawed instant came out and
56:34Claude 2 came out. Yes, I remember the marketing mostly being focused on pros like this model write better and long context um this directly affected me cuz I told something on that what did you make uh small developer which was my Devon before oh yeah small yes uh and uh you know so I think like there's there's all that really like good like focus is another thing that is a question that people do want to ask uh you know you could have built any any
57:03other things like and obviously open roto was working working working uh were there other ideas that you wanted to pursue that you turned down you know just the paths roads not taken we made a couple prototypes for things that we didn't launch one was a fine-tuning model as a service lot of that open pipe and all those things but it was it it kind of was in a very consumerry form factor where you would give us a YouTube video or two or three we would then extract all the
57:33transcripts from it and try to fine-tune model to talk like the person in the YouTube video or the people in the in the videos that you sent. So like a really really easy way of creating a fine-tuned model based on like some kind of videos that you like.
57:46That would be so useful.
57:49We we we made it too. It was like I it was and nobody used it. was we didn't actually like test it with that many people because the model marketplace was our main our main focus and it and it was like growing and we're building more conviction in it over time.
58:06Um just just as a creator I've been pitched many like uh you I have 500 hours of recorded voice of myself make a thing of you charge access to it.
58:18Uh works for only fans doesn't work for for us as regular people. I think this is mostly uh it's just a glorified rag bot whether it's in the weights or it's outside the weights doesn't really matter. You're just doing rag on the videos and people ultimately always just want to find the source video uh that directly answers it.
58:34My use case was mostly to practice with from with myself cuz I often like to see what like the way I practice for a job interview or if I'm hiring a candidate or public speaking or whatever is I I wish there was like a good mini me that I could like critique cuz it's kind of hard to pull yourself out. I would never get I would never offer it to other people service like pick your top five mentors that then talk to them instead of talk that would be cool too. Yeah, that was that's a replica and that was the use case we were aiming at.
59:01It's like you want to create an experience like AI Steve Jobs and AI AI Steve Jobs was the the initial use case that's a even though it's not allowed that's a that's a common prototype. you know talking about adjacencies fine-tuning as a service as part of the router service is something that I would typically think about as well right like like why don't you do that because if people are running already their inference through you store everything log everything uh fine tune to a smaller model that is cheaper faster all these things that that's within your control right uh you didn't do that but like
59:30other people would have pitched that in the general state of of infra startup I think you were just maybe a little bit early because today that's an extraordinarily fast growing segment like you know from astral where they do a lot of enterprise deployments I mean it's often fine tuning is you know custom models for ASML or whatever often but not as a router they're just like I come to you because I like your ML models I want custom model right it is not I want uh to run all my openi prompts uh get store all my results and then just move off of openi right they're not doing that
59:59uh as a as as like a way to export off of dependency on a on a frontier lab I have not seen that yet which which was your kind of decision to do I mean we we decided really we like leaned into our focus and figured that like there are like we just saw the ecosystem develop over time. All these inference providers that that do want to help companies do that. Um be like like it makes sense for us to partner with them and to like give users
1:00:27lots of choice and to like you know figure out what makes them um what gives them competitive advantages. It's it's a whole new business basically and there's there's value in being a neutral marketplace that just kind of like works with those companies. Um could you share a little bit um to to Sean's point like how you prioritized what what are some ways you prioritize features cuz you've always done it so elegantly. I never you know it just happens and you make all the right decisions that always have product market fit from the outside looking in.
1:00:57consistently you seem to have prioritized you know a lot of hit features that worked and maybe I have a sample set bias or whatever but Sean can list what you think hit features worked well like oh the leaderboards like leaderboard okay yeah you know like from day one the feedback charting B okay but like he had like plugins uh you know he had like uh and I think there was a whole thing I want to get into about like completions versus yes check completions versus completions and then also uh let's call it like the the rise of the reasoning models and how you deal with
1:01:27multimodality all those all those things by huge one there's one like I think it was in early 2024 very early 2024 we thought it might be interesting to fuse the results of multiple models together and we launched a prototype called mom mixture of models that let you like pick a couple models we'd pick them for you and then it would fuse the results
1:01:56together at the end and it would show you all the intermediate results in this like big conbon board looking product. What does the fusion at the end? Another model.
1:02:05Another model. The the smart the smartest of of the three of the set.
1:02:10So, this is like a council idea.
1:02:11It was a model. It was like a very early LLM council.
1:02:14This is a multi- aent swarm as as like they would call it at one of the frontier labs uh in the early days, you know.
1:02:20Yeah. Like some of those ideas are like going the right direction, but the devil's in the details. There's a lot of like product refinement needed to make them really work. um they take your focus away from right you know whatever else you have going on and there's a lot of like community building and learning that you need to do and the technology might be too early so there like all kinds of reasons they might go wrong and in our case the technology was a little too early in other words the fused
1:02:47result was a little bit worse sometimes the same as the best model that was being used to fuse because the best model was so far ahead of options two and three at the time. You know, over time, the top three or four LLMs have gotten closer together. Still neurode divergent, but like all capable of inserting like pretty interesting ideas.
1:03:12Like RL has basically like expanded the surface area of creativity for machine learning researchers within each lab and so they can, you know, diversify the reasoning power of different models more effectively. At least that's my my theory for why uh Fusion is it like works better than it used to early 2024.
1:03:32And um so the technology was a little bit too primitive. The form factor was was not right and and so we would have had to go through a couple more iterations. And so we decided to just delete all the code. And uh then years later, middle of 2026 um or early 2026, we're like let's bring it back. like the research is looking kind of promising for fusion. The models now have like two, three, four top frontier models that are all really good
1:04:02and like like I'm I'm frequently trying to like consult multiple models to get the best results like and then I you know I ran a little personal experiment where I was like I'm going to like do a an architecture plan for a code change. I'm going to give it to all the models.
1:04:19I'm going to fuse the result and I'm going to ask all the models if the fused result is better than the individual result each model came up with. And they all said yes that the fused result was better. And this happened a couple times and I was like okay spot check pretty good. We should like benchmark this and that's how we built fusion.
1:04:38Yeah. And it came on your fable. So you were like this is fable level.
1:04:41Yeah. Yeah. Let's start leading up to to this year which we haven't got to gone to this year. you know can you mark out the main milestones in the journey? I I think um it seems like your your promise uh was you know routing you decided the business model very early you take you take a cut and like you know what are the major milestones that uh inflect the growth right like you're you're growing like 9% week on week now is that's the official number in terms of token volume I think that
1:05:10sounds about right yeah yeah so just like can you can you mark out like the the sort of brief history of of open router up to up to you know the acquisition let's let's call it we we're we're just we're just talking about uh you know people are uh have you have sort of your your birth moment with um the misreal stuff where where people are really competing you have your state of AI thing where where it's very cute you have 100 trillion tokens haha uh because now you're doing 10 a week uh you know um we're doing 10 a day 10 a day now
1:05:40yeah more so so yeah you do this in 10 days like what are the major points there you know I just want to like there's a smooth curve but like you you you feel the infections A lot of this is is kind of oriented around model launches. Um, we had, you know, a huge focus on pros all the way up through May of 2024.
1:06:04Um, because coding was just not there and no apps were able to build much on top of it. So, um, you know, a diversity in models but not a not a wide diversity and not a wide diversity in use cases.
1:06:17Dream Tavern was one of our top apps at the time. The creator of Dream Tavern Dream Tavern now runs product at Cognition Devon. Um the then we in in the middle of 2024 we saw Claude Sonnet 3.5 that came out incredible leap forward in coding. Um and we saw the dynamics of like apps building on top of us change. um we saw a huge surge in
1:06:44volume in like users uh using open router and and this is when I think people started to look at the like money that they were spending and and get a little bit like whoa what's going on I might need to like think about like more costefficient but equivalent models and shortly after that I I think it was after Sonic 35 mix uh mixtrol 8x7B came out and everyone was like what this is the model like the open weights community delivered and so it was really
1:07:13good timing from the strong.
1:07:15Basically all of an cos are just helping you out.
1:07:19It takes an ecosystem to grow an open router, you know.
1:07:22Yeah, that was the Yeah, it was it like it was the this early ecosystem. It was like a a swing action where like model labs would come up with some sort of front tier innovation like usage would surge then users you know look at their invoices 30 days later and like whoa what's going on here and then open weight models would deliver like a like cost effective options to 3 months later. We saw that happen several times.
1:07:51Uh, one thing one thing you also did with the coding agents was that you broke out which are the top coding agents and they love that they love that leaderboard. The the client versus the ru code versus the what have you.
1:08:02Yeah. Yeah. Like Klein was like the top of our leaderboard at the time. We we then at the end of and I'll skip forward a little bit. the end of 2025, there were quite a few coding apps on the leaderboard, but they were all IDs or or you know, terminal based agents. And at the end of 2025, we saw Open Claw appear. And openclaw was like particularly interesting because one it
1:08:31was like a a new form factor that like brought in a new type of user not just a developer but like a a productivity or sort of a like an internet creator came to AI for the first time. And it also had an interesting architecture where it was like calling your chosen model for these heartbeats to see if it was still alive in addition to actually using the model for real tasks. And the heartbeats are like they're kind of you
1:09:00don't want to pay a lot for a heartbeat.
1:09:03So, um the the auto router that we provided was really really useful to this like wide range of users all of a sudden and um and so we just saw it rocket exponentially and then we saw you know like open claw just blow up and a couple other um apps lean into that new paradigm and do something similar. Um, Hermes came out and really leaned into things like the auto router and built like a really good community and leaned
1:09:33into like like basically skill management and making it really easy and effective for people to like like set their memory in the agent and and build really good skills which another thing you never did memory skills sandboxes all these like adjacent things you could have done could have but like it's I think like it's hard to bet they're also very there are things that developer that really matter for like the developer use cases that were coming out at the time like developers wanted
1:10:01to architect those things those were kind of critical to building a good user experience. It's really it was like hard it's been hard for companies to find abstractions that work for all developers on the memory layer.
1:10:13It is it is you know there are some um like Mastra has done a pretty good job for example but uh but like developers have like lots of varied preferences for them and then we saw you know our the way our our leaderboard has has changed over time is kind of like a movie of how the AI space has changed over time. If you just sort of like go to the wayback machine and look at the the the rankings leaderboard and the app's leaderboard
1:10:41over time, it sort of shows you like what's happened in AI over the last couple of years.
1:10:46To me, the coming of age moment was uh Andre Karpathi was like, I no longer read local llama cuz like I just go to open routers leaderboard which I remember that I think he probably like said like sorry guys like I'm going to send a bunch of traffic to you.
1:11:01So I also want to bring it into the Stripe uh thing. Uh how does that kind of conversation start?
1:11:07We had this longstanding relationship with Stripe though from you know like many different projects that we had worked on with them. We invest you know a lot of effort in countering abuse um token fraud and token fraud.
1:11:23Can you give some numbers just to so people understand?
1:11:27I think I like I I posted about this. We we blocked 10x as much dollar volume last month as the month before. And the types of token fraud are diversifying quite a bit. You know, there are like fraudsters going after typical stolen credit cards, but they're also, you know, people trying to resell traffic against the terms of service. There's like hacked accounts. There's people who just lose act, you know, like their
1:11:56their whole company is compromised and they don't even realize it. Um, and we help them like regain control and detect it. There's there are accounts that are like reselling inference on the side.
1:12:07There's there are accounts that are dealing with, you know, like an accidental runaway agent and and they don't realize it. Not not a hack, but it's something that blows up and um and the company doesn't want it. And so our trust and safety team like works a lot on all of these like categories of problems and and helps block it and detect it. And so we've built these, you know, we have models around them. We
1:12:36have we worked closely with Stripe for a while on this. And I think it's going to become a huge problem in the ecosystem.
1:12:43Like we're already seeing a lot of companies start to see these fraudsters like spread and look for other ways other you know other than open router to other fraud vectors. And if you're making a gateway or or selling like generalized inference you are a target for fraud. If you're selling very discreet like intelligence products, intelligence products that are like doing something pretty specific but not
1:13:11like you know just reselling inference with some added capability, then you're way less likely to get these fraudsters. So it I think we'll see companies also move away from just reselling inference with some sort of like added capability and move towards sort of like discrete tasks and charging for those tasks and charging for those enhancements and letting people bring their own inference like in a third party way.
1:13:37Whoa. Okay. And yeah, obviously you would power that. Uh but you people pay uh for outcomes or per task? I think people will pay um you know I think like the the the data dog pricing page is a is a good look at like the future to come. It's like companies like infrastructure companies will like charge for different types of events that they're providing and there'll be lots of like continuous
1:14:06pricing models that look like that. And of course there will be like if you go down uh towards consumer apps you know simpler pricing more subscriptions um you know fewer events to worry about um and ones that like are not focused on just adding a markup on top of inference not just because fraud is hard but but also because the pressure from the labs and from inf like good inference
1:14:35providers to like do a commit and then bring your inference elsewhere is going to be very high.
1:14:42Any comments? uh two one you know I I think Alex has done a very eloquent job of describing something you know counterintuitively I knew would be a thing at scale like four years ago because of discord and the particular experience that taught me this was um you know as we started scaling midjourney you know the one of the primary ways that we used to give away or like get people to try midjourney early on to get to their first 10 generations cuz you know 10 generations of 10 images generated was roughly the
1:15:12magic moment activation point we found like once you done 10 you were like this is extraordinary um but for that we so we had a free trial with midjourney and one day I woke up cuz they had a platform and had to monitor I had all these dashboards um you know I had like three missed calls from David and it turns out like they had there had been this flood of new users overnight and and we were like this is great and he was like no actually we shut down the free trial and I was like why is that um and he said I want to look at the
1:15:40geoloccation IP addresses And basically somebody in China had started to resell midjourney free you know subscriptions with the free trial as a way to like basically you know it was fraud abuse right and even for a specialized model like mid journey yeah and that was an actually an application so this idea that I I think the big picture realization I had back then was hey there's a new type of unit of value that's being streamed across the internet called a token and over the
1:16:10next 10 years the entire internet value chain was going to have to deal with the fact that like the more valuable tokens got, the more bad actors were going to go to try to get their hands on those tokens. And anytime you scale something and the payload gets more and more valuable, more bad things people try to get access to that value. And and so it, you know, it was very obvious to me back then. And so look to this day I don't think there's a free turn like I don't
1:16:39think majour's ever actually turned on the free trial since then because it was really not an easy problem to solve in terms of trust and safety. That's why I try started teaching the class security at scale at Stanford like this like one of the that and the anthropic learnings to me it was clear that the need for security at scale was going to be enormous a few years from then cuz if you just do the math right think about like if where you know online payments you know is started roughly in the 80s
1:17:07and 90s right and grew to over a trillion dollars over the next 10 years and we needed to build entirely new payment solutions to deal with online fraud. um where we are today is roughly there on tokens but over the next even five years we're expecting the token economy to get to like roughly $5 trillion and over the next 10 years I'd be shocked if we went to 10 trillion of token flow and so if we were starting to see such aggressive abuse and fraud at
1:17:36subscale mid journey remember mid journey at this point was like less than 300 million revenue run rate a year I I just realized we were going to need like entirely new like systems to deal with the fraud that was going to happen for trying to get into the token flow. So um my I I you know when I I forget the board meeting it was when you brought up that you know Stripe wanted to to partner up and it made so much sense to me cuz Stripe Radar when I was a cliner 10 years ago we invested in Stripe and the whole pitch that you know Patrick
1:18:04and John communicated so eloquently was like hey unlike traditional payment tools like Brainree that do a 7-day verification like KYC and email to get get the fraud out of the way we actually just bite the fraud cost up front as a customer acquisition cost and give tell a developer or like just use five lines of code and we'll start accepting your payments in 5 minutes and what'll happen is over time we'll collect all this data on the developers. Cloudflare model is the Cloudflare model, right? And and they did. 5 years later, they launched Stripe Radar and Stripe really today is
1:18:34a security company. That's the real people think it's a payments company. No, the reason there's lots of other payments providers today that give you like cheaper payments transmission, but the reason stripe keeps, you know, being the dominant one here in Auden and Europe is because they have extraordinary fraud detection that they've built, you know, with over the years.
1:18:50The same story with Elon and Max Lechin and and a firm. Yeah.
1:18:55You know, I think the story shows up over and over again where every time you have value streamed across the world in large amounts, you need new protection and security infrastructure to fight to to keep the bad guys out and allow the good people to like have their transactions happen really fast. And so I think you know the this is why the from my perspective like the stripe and open router story is a security story for the internet ecosystem for the frontier AI ecosystem. without a partnership like that, it becomes very hard to defend the quality of experience and the speed and all the good stuff
1:19:25without letting the bad guys get in the way. Um, the second is that, you know, there's this underappreciated thing about like the fact that you need to like these all the bad things that that Alex described as being perpetuated by humans right now is going to be perpetuated by AI agents over the next 10 years, right? So, think about the like recursive scale we're about to see of bad actors. It's not just bad human beings. It's it's all the bad agents that are going to be attacking the token
1:19:54flow. And and there's it it's very hard if you're a researcher and or an AI lab to reason about that problem because the only data you have is how agents you're training are going rogue. But that's just a fraction of all the bad behavior on the internet that we're going to see.
1:20:10And so what you need is is defenders, new sheriffs in town, which cowboy hats, um that can can see all the bad behavior from AI agents across the ecosystem, from different model labs and different post-trained deployments and different developers and take all of that data and say we're going to build a shield for the entire token economy. Because without that, you know, the amount of fraud we're going to see of this 10 trillion dollar in GMV and global GDP growth is like a huge percentage of that, I think, is going to be fraud,
1:20:39abuse, and we might never get there if people just don't trust tokens, right?
1:20:43Um, and I don't think this infrastructure exists. So, you have your work cut out for you with at Stripe, but I don't think people have realized the scale at which agents agent agentic fraud like bad behavior perpetuated by AI agents is about to hit us like a tsunami.
1:20:56Yeah. Uh I mean there's a lot to dig into there. Uh I want to give you the last word. We do have to wrap. Um um what can people expect from open router and stripe?
1:21:04I mean I think this is a really good way for us to accelerate go to market and um to go up market more quickly. It's also, you know, as an eloquently described, this is there's a really clear better together story here when it comes to improving trust and safety and making it really easy to like accept tokens and let people bring their own inference to your app and to help developers just like build on top of inference. Uh,
1:21:33going forward, we have a really strong brand with Open Router and we're keeping the brand. So like open router like as a as a product and the road map and the name and the brand like you know is staying the same and so what like you should expect you know in this next 6 months is that most things will be like what we would have done had we been independent except everything will be moving faster and that that's kind of
1:22:00like our you know near-term goal longer term hopefully I can comment on it soon but I can't now. Okay. Well, uh, we'll hopefully do a follow-up at some point. Uh, but thank you for being so generous with your time and, uh, congrats on the partnership. I mean, this is one of the most beautiful bromances I've seen in a just starting starting from Stanford to here.
1:22:23Lots more to do. Lots of sheriff policing to do of the of the town the token economy. We need new we need new sheriffs for sure.
1:22:31Yeah. Awesome. Thank you.