0:00One of my biggest challenges today in adopting new software both personally and as Verscell is understanding the security debt that I'm taking on. Just the surface of attack has become so massive as of recently. We probably need to start writing code in different programming languages. The challenge that I think a lot of organizations have today is that the amount of latent vulnerabilities of all of this human written code is almost immeasurable. We could eradicate supply chain attacks if we moved more and more of this
0:29development life cycle to to sandboxes.
0:31All right. So, Gilmo, what will matter in the post AGI era?
0:37I think our ability to stay in control, okay, as humans. And by that I mean creative control, control over the risks on cyber security.
0:49Control over how we make the agents work for us ultimately. You know, I've been speaking a lot about the leverage is shifting one layer of abstraction higher into sort of the hardness engineering, the software factory, the fleets of agents that you orchestrate. I'm a huge believer in our taste and our creativity as human beings uh because we're usually appealing to other human beings and we
1:17love to buy the stories and products and services from other human beings. You know, we like to think that we make perfectly rational decisions as humans, but a lot of the time we're actually buying into the stories and the vibes and the brands and the flavors and the tastes. Ultimately, Asians are going to be hopefully an enabler of greater and greater storytelling and greater and greater human ambition. I like what you said about the higher level work because
1:47people always are afraid of like this job going away, this task going away and you know farming was something that like 90% of population did 200,000 years ago, 200 years ago and now nobody that works in farming, less than 1% of people. So thinking of this is like filling out a form just feels like manual labor, you know. So what will that look like? Will it be like managing fleets of agents like deciding which models, deciding like where they run in terms of the hardware?
2:11It's so funny because one of my most one of my best performing tweets recently was I do I do this for fun. You know, I build software to relax.
2:23I build software for work and then I go home and I build build more software. I build software with my kids. I build software with my wife. Like I'm always building software. So that's where I think the analogy breaks down with farming. And by the way, like there are some people that are legitimately love farming and maybe they, you know, over hundreds of years, millennia, they'll still do it.
2:42It's called gardening.
2:43Yeah. And they keep their little like mini farm in their balcony and that kind of stuff. But I think a lot of us do this because we this is our medium for creativity and there's no taking that away. There's not no putting that genie back in the bottle as far as I'm concerned. Um, but it does mean that I mean you kind of see this with models.
3:04these days being benchmarked over their ability to create new worlds. There was that Garpathy discontinuity tweet of, you know, first he introduced vibe coding, then he introduced, oh, by the way, the models are getting so good that they can create entire worlds and they can uh rerender the Hobbit or Lord of the Rings.
3:24And so I do think that we're getting into the next level of how much r free range of creativity is enabled by AI, creating new worlds, creating games.
3:36I've spoken of how the web itself, the web that we will consume as humans is becoming like ultra fancy. Like the the experiences that stand out to you are the ones that just like really go out of their way in terms of wowing you. 3D and animations and particle effects and things like that. And so it's amazing that like literally we can do anything now. We can build anything. And uh and people are flexing that muscle. Like I see a lot of people that are building
4:04stuff that they're just building it because they couldn't before and they're deriving a lot of joy from building which I think is really good.
4:11I've been coding with AI for nearly four years. In 2023 I was the first to start the build anything trend and since then a lot has changed. We have different models, different tools, different agents. But despite all of that change, every single project I've built has one thing in common. I always use Superbase as my back end. I'm not even thinking about what DB do I use. I just use Superbase. It's simple, it's scalable, and it's agent first. In fact, it has never been easier to get started. Now,
4:41the AI agents can literally build the entire back end for you because Superbase now works inside of CIGBT and cloth. We all know that with AI, you can write code faster than ever. But where people get stuck is the back end. the database, the login, the storage, all the stuff you need to get right. And Superbase handles all of that for you. With their official connector, you can just ask Codeex or Cloud Code to use it.
5:05Just say, "Create a new invite table or deploy this function or review my app for security issues and the agent will do all of this for you right inside of your Superbase project." I'm using Superbase for all of my backends and so should you. And the best part is they have a really generous free plan. Just go to superbase.com and try it for free. It's the first link below the video.
5:26So, given the fact that anybody can build anything. What gives someone an advantage? You know, people watching this, they want to know like how do I get ahead? How do I build a business? Is it the taste? Is it the creativity? Is it being more technical?
5:38Yeah, there's so many things. You know, I was actually going to post recently. So, every time a model comes out, you probably like the first thing you do, at least both of us, like we try it out.
5:49Definitely. And one good exercise I would recommend people to do is there's a series of things that are just impossible for models to do today.
6:00Now that list is probably shrinking and changing over time. But you know 12 months ago that list was bigger and there were a bunch of things that it couldn't do at all. And what you want to do is like you almost want to keep like a list either in your head or like write it down of like what are the wildest most ambitious prompts that a model could do. And I think it having the timing, having the anticipation
6:28really pays off because when a model drops, it's almost like people are exploring its latent space and discovering what's possible and it's an opportunity for you to like it might be that you can build a business like oh you know Opus 5 5.5 came out and this particular thing that models used to fail at has now become possible. So there's that whole category of things. I think people can use the creative power that they get from the models to you
6:58know broadcast a better message. So you see this with companies taking advantage of like hey models can do this um we can portray our brand or our product in terms of what this model enables. So a good example for us was that when Jev came out, yeah, there were all of these really cool things that it enabled for our products.
7:19And so Jev comes out, we can now sort of like ride that wave and like teach people. There's there's such an appetite by the world to learn about AI to make sense of AI. You know, a lot of people are out there scared. Yeah.
7:33So for example for us every time a model comes out we run it through the deepseack benchmarks to help the world understand okay this model is good at cyber security it costs this much there's a lot of alpha right now in which is super ironic because models are so easy to use like you just prompt them in English and yet there's a mastery to them yeah for sure there's understanding what model is best for what job uh and so I think there's a whole space of like evaluation
8:03Whether is that you can become a voice that helps other people through public evaluation make sense of what models are best for what. There's also the personal evas that I think has a lot of alpha.
8:14You know, another irony of AI is that we do go off vibes a lot. Like I hear people say like, "Oh, I read this thing on Twitter, then I started prompting my the model differently." I heard someone the other day say like, "Oh, I I read this prompting guide and said that Astra does better with like less this and that and like now I'm doing that." And I asked him like, "But did you actually ascertain that is giving you better results?" It's like, "No, just I'm just going off of like what other people are saying." And so I think there's a lot of
8:44um there's a unique advantage now to are you actually deriving truth from what the models can do and do you have a you know a set of like personal uh evals or do you have a set of company evals that allow you navigate this world that is moving so fast like every every single week there is sort of like a new model to evaluate or uh as we saw with Jev there's even new categories of AI products coming out every week but isn't the vibes is like part of what
9:13makes us human, you know, because the models they sometimes like the just the practical sense, you know, it's like, okay, we need to ship this feature. If you're a human engineer, you know, like there's pressure on this and the release is tomorrow and you cannot be like overly perfectionist, the model will happily spend 30 minutes on something useless, something small that a human engineer will just skip given the broader circumstance. So yeah, do you think like the creativity, the vibes that will still remain or better models will have this? I think creativity forever because again unless
9:40we go on like a social network made by AIS and we decide to follow AIs and how they're going to influence our future.
9:50The reality is that we're still in charge of our brands, our storytelling, our values. Even I think you know there's a certain design aesthetic that humans can push forward, right? Like when you look at a versel, Verscell looks a certain way and we invest in making it so and creating an emotional connection with our audience. And so what you want to do is you want to use AI to I I mentioned sort of amplify your
10:19story. How can how can using AI whether it's the creative tools, image and video models, can it help you tell a better story? Can it help you ship? I mean, what's what's very clear, right, is that you're going to ship more more frequently and you're going to ship probably better products.
10:34Y there's a still a delta between like are you shipping at the true speed of AI?
10:42Are you behind? I think for a while I think we're going to have a range there where there are going to be individuals and companies that are just ahead of how much they can harness this superpower. I think there's an analogy there with, you know, when cloud computing became, you know, I guess fashionable or or or or increased in popularity. There was a relative range of like, you know, there there are companies that groed it right
11:10away and used it very efficiently. There are companies that are still like trying to migrate from on-prem to cloud. And I think incidentally, I think AI is pushing a lot of people now like the the ones that were sort of procrastinated on it. And you can sort of imagine right that there are a lot of people that are kind of procrastinating on like how agentic they're becoming and they're probably going to get lapped by this uh organizations that are moving more autonomously and with more conviction.
11:36So let's go deeper on that. What separates like the company that really squeezes AI to the fullest? Is it like building the autonomous software factories because this changes every quarter? I really I mean I really think like we have such a we have a once in a-lifetime opportunity to do what customers want which is sort of ironic to say but like any big company at any given time is getting customer feedback.
12:02Uh you know let's say you're Comcast, right? Like people are saying things on Reddit. They're calling your call center. They're you know sending you emails and their internet broke down and they want to understand why. And you know, what's your ability today to like, you know, turn those detractors into net promoters? Uh, when they're reaching out, are you responding with a high quality agentic response? When they reach out, report a problem, are you
12:29kicking off an incident investigation?
12:32When they're reporting with a very concrete fixable thing, are you fixing it autonomously? I think this is actually a pretty decent way to grade like a rubric for how agentic are you from customer signal that is decisive about something in your product being broken, defective, or in need of improvement.
12:53Are you actually shipping that? Are you shipping the the fix for that signal?
12:58The thing that addresses that signal and what is your latency? Like you could almost make a case that you can measure this sort of uh like uh mathematically, right? like, okay, signal came in.
13:10Eventually, you're going to get to fixing it. Maybe to give you sort of a lay of the land. There's the company that periodically scrubs that maybe through taste, maybe with product managers, maybe by, you know, getting enough escalations, they're like, "Ah, I'm going to fix it." And there's the company that is sort of like becoming truly more autonomous in a good way, right? Not in the sense of like we let the Asians go crazy and ship slop. Um, and you know, speaking of Jev and why
13:39people are so excited about it, you know, this idea of like really fast decision making can be super critical in shaping these software factories because I'm mentioning that there's an element of judgment and taste also in you know famously uh folks like Steve Jobs, DHH and many others have have said you know like you're not just a machine that turns whatever customers want into product. There is an element of of decision making there but there's also an objectivity to like you know your
14:08product might be erroring.
14:09Yeah. And so you know throwing up a bunch of exceptions in prod or someone comes in with a legitimate very straightforward improvement. You know we like to do this a lot on X like when people like give us feedback on X sometimes they like attach a screenshot or a video it's like that's a spot on like there's there's no human creativity involved here. like they're pointing out something that they experienced and it was a problem that we should fix it right away. And so that's kind of like an opportunity to like test out this
14:37software factory hypothesis.
14:39Interestingly, yesterday we got a guy saying that he was getting confused by our navigation and he gave me some really good feedback but nothing that was like immediately actionable. Like it's not like I could take his thing and turn it into a prompt. And so what we did is um uh the leader of our sort of like dashboard product area reached out to this gentleman and they started going back and forth. He he gave us more videos. He gave us more qualitative feedback and
15:09even then requires another layer of human judgment to turn it into a decision for what we're going to do with the product. Now one could project very far out in the future which in in AI could be six months, 12 months. There is a world where some of that UI and some of that expectation that this customer have that would become their own generative interface. But I think we're a ways away from that.
15:36So you know this is still our role to sort of like balance out feedback that we get from customers and then uh sort of u eventually find its way into a prompt and eventually find its way into a shipped product. So when it comes to like the obvious ones like button isn't working, payments are failing, do you think the number of like like you said the latency towards a fix resolution? Yeah.
15:58Resolution is that like the one of the main numbers of predicting of a business success because think like if you go from first principles business is a product that gets better over time.
16:07And if you have a unless you have a super small addressable market, it's like that business that can just fix everything in sub 30 minutes is going to destroy a business that takes two days.
16:15100%. And that's been the history of technology, right? That's why I gave the example of like the companies that were not afraid to adopt personal computers that said like, "Oh, we're going to do this stuff by hand." The companies that were not afraid to adopt the internet early on. The companies that also made it their priority. So, there's something really interesting to learn from the history of the internet, which is a lot of companies like the internet, but it's like, okay, I'm going to hire a company that does the internet for me.
16:43There's so many companies in AI want to do this.
16:45It's still happening. Interestingly, right, like you have, it's so funny because I talked to a lot of enterprises and like it's so common that they tell me that they don't own their own website. And so there's a parallel here, which is that there's going to be the companies that say, "I love AI. Can we like buy a bunch of AI licenses and throw them to our engineers?" And there's going to be the ones that say, "No, no, no. This is a disruption that creates a new competitive landscape." If we become really good at mastering this
17:14thing, it's a little bit like uh redstone. If we become really good at like manipulating this new block, yeah, this new raw material that's going to put us ahead. That's 100% what I believe. And uh uh and yeah, to your point, there's a there's an element of measuring. Uh I I don't want to get only, you know, tactical like if you improve that metric, that's the end all be all of your business. We like to say that there's iteration speed and there's
17:44Iteration speed is, you know, like how quickly you're turnurning out stuff. And that's why, you know, when I hear about how many tokens you spend, token maxing, and when I hear about how many PRs you land, good for you, good metrics. But velocity implies speed in a direction. You could be pring, pring, pring, pring into oblivion. I'll give you a very concrete example. AI is is helping people realize that they can build a lot of things.
18:13And what I don't know that it's fully helped people realize yet is that there's certain things you should build and there's certain things you should buy. There's services you can use for certain things and there's things that you should build yourself. I talked to a startup founder recently. She came to my office and she was like, "Yeah, I'm building this really cool idea for agents and I I thought the idea was really compelling." And then she proceeded to tell me that she was also building her own email service and her own payment gateway, her own like Yeah.
18:42And some of those things might make sense, some of those might not, but like she was taking on a lot. That was sort of my my sense. Like sure there's the very optimistic take which is that you can have so many agents working in parallel for you that yeah you can take on a lot more projects but I'm still a skeptical because what you really want to do is you want to pick a few set of a set of problems where you think you can have disproportionate alpha and go really deep and spend as many tokens as
19:11you can on those problems and move as fast as you can on those problems and that and that that would imply that vector that direction of that that the word velocity implies. Uh so it's not just PR landing, but like where are you headed and actually are you actually making progress towards your goal?
19:26Yeah, because if you're doing 20 things at once, you're going to be like mediocre in all of them. And to be successful, you need to build one thing.
19:33By the way, this is kind of becoming a thing already. Like you realize that a lot of people ship really fast and then you you come to their website and it's like a pile of slob. like you can sense that like there's no value add over what the model did. By the way, this was already the case with the web. Like people would not respect the vanilla template that looked like you just
20:00purchased some kind of starter kit and threw it over the internet. I think people are really not like it's almost like you're tuning out a bit when you see something that looks like it was just regurgitated by the LLM because okay what is the value ad like if you could one shot throw if you could one shot it I can also one shot it and so there's an element there of like again the sweating the details that help you communicate your intent and your values and what
20:29you're trying to pursue much better. By the way, if you have multiple $200 subscriptions a month, pay attention because I want to be working closely with a couple of you to build Cloud Room. This is an open source platform for running agents in the cloud. It's the world's first open source solution for cloud agents, right? You have Devin, you have cursor, you have Codex, you have Cloud. All of these are closed.
20:51They want you to use their own subscription, their own harness, their own GUI. But you're already paying for inference. You're already paying for tokens. You have probably multiple subscriptions. You should be able to use that inference. You should be able to use your favorite models, your favorite harness. This is Cloud Room. It's fully open source and it's free. However, right now it's not fully open to the public. I'm looking for the real power users, the people who can give me feedback daily, people with high taste when it comes to coding agents and running agents in the cloud. So, if
21:19that's you, if you're paying for a $200 plan on Codex, cloud code, and a bunch of other stuff, go to cloudroom.dev and join the wait list. As I mentioned, every single day I'll be selecting a handful of people, giving them access and working with them closely to build this platform. So yeah, if that's you, go to cloudroom.dev. It's also going to be linked below the video.
21:38Let's talk about your current agent engineering setup. Like how does it look like? What do you use?
21:43Yeah. So broadly speaking for Versell, we've encouraged people to, you know, try everything. We're not prescriptive about what coding agent to use, you know, what frameworks to use. Like our our stance has always been like open. In fact, our fastest growing product at this point, I think by revenue, adoption, numbers, etc. is our AI gateway. And our AI gateway, like the
22:10thesis is, you know, I mentioned if you're a company, you should probably be investing in your agentic engineering fabric. Uh you should also not be using just one model.
22:22You also showed me yours and I thought it was really cool that you basically built your own client or your own uh AD and you can choose between all these harnesses etc. So it's kind of stands. In fact, we have this app over here in the in the menu bar where we give you a sense of like what your token spend is.
22:41So every engineer at Verscell gets this app and like they can understand okay like what is my how much did I spend on the OpenAI model versus the Kim model versus the Opus model etc. So like one of the like founding philosophies is you should uh try many things especially models. Uh personally you know AI has been amazing for me because I've always been a huge believer in the simplicity
23:07of the CLI. So when I started Verscell, if you go on on the way back machine, I started Verscell by shipping a CLI. So you would go to zite.co and you say install now. Now was our deployment CLI and as the name implies basically all you needed to write was like now in any directory and you get back your a URL of what you've shipped and you know we ended up renaming the
23:37product. is now called Verscell.
23:39But you know 11 years after the founding of the company, the thing that continues to sort of grow exponentially and continues to give us all our gravitas is this foundational concept. There's a CLI that ena enables you but now increasingly agents to take anything they've built and put it in the cloud and make it instantly accessible. So my personal uh engineering stack has always been like can I can I do stuff in the
24:06CLI and um I've always been a neoim user. So my next obvious logical step was cloud code. I got really excited when cloud code came out because everyone else was trying out sort of like building very sophisticated apps.
24:23Yeah. And they became successful by basically going towards the kernel of what makes AI different which is well at the end of the day it's just packaging intelligence and giving you a prompt. So um I'm today I'm using FX which is the uh coding CLI of Verscell Labs.
24:48One one of the design principles of this CLI is that it almost feels like a shell. So when I open a new tab on Ghosty like or I open a split like it boots into ZSH but one of the things that we obsessed about is like can this boot as fast as a shell itself. So, it's a very tiny uh binary. Uh installed is 11 megabytes. If you download it today, which I'm going to six uh it's actually 6 megabytes. It's 11
25:17megabytes on disk. Um and it's incredibly fast to start up. So, like the idea was like if if I started a shell within a shell, like can I make it like just as fast? Um so, what inspired you to do this?
25:30Because like most agents are very bloated, you know, with endless pre-shift skills, plugins. What comes to mind is PI that's a probably the closest minimal agent. Were you inspired by PI or like what was the idea here?
25:42So interestingly enough you know Verscella for its entire history has been associated with TypeScript.
25:48Uh we have one of the leading open source projects in the in the TypeScript ecosystem which is actually the leading and most popular open source project which is the AI SDK.
25:56One of the things that we wanted to experiment with with FX is two things.
25:59One, what would it look like to start investing more in native tooling?
26:04So, one of the frustrations that I always get with uh with CLI is uh maybe maybe I'll not to roast them, but like every time I would launch Claude, you can sort of see it there like there's the boot up time and then there's the like disclaimers or whatever like the time that it would take me to start up and like I I just want to get something out of my head as fast as possible um was kind of annoying. But the other thing is that I started realizing that the Unix philosophy is about composing small programs.
26:34and a bash the bash shell or zsh and whatever they're they're just that right like they're programs that call other programs and so there is an element of like let's experiment with making a program that is actually going to be fairly immutable over time there shouldn't be more UI added to this over time yeah that's the problem with a lot of these agents like you you build a business on top of them and then like it's crashing you know they they start out small they start out fast you know this actually what initially made me fall in love with
27:04claw. It was like, "Oh my god, there's almost nothing to it." Uh, and then over time, I think you you really need to fight off this temptation to add more and more and more. But it is true that people should be able to extend these things. That's why they're open source to begin with. And so effects is actually simultaneously this like super minimalistic CLI and also a library that you can embed in other programs. So the classic sort of C style paradigm here is
27:32that you have your CLI and then you have lib in that thing. So there's lib effects that goes together with effects. I'll tell you like this thing is so addictive that um you know I started using it for my own personal research.
27:46Um I realized also that it's not only that we were gonna have to like clear this probably in the in post blur it. Um, but I started realizing that like, you know, another thing that AI was sort of bringing to to the foreground is going back to the fundamentals of Unix and the fundamentals of computing. It's not just a bunch of small programs. It's also the fact that the file system is basically all you need. Today prior to this
28:15interview we announced Verscell drives which is a way of sort of like modeling a file system in the cloud and that you can sort of like launch a bunch of agents around. So you can launch a sandbox and attach a drive.
28:28I don't know if you saw some people started reverse engineering muse and instinct or how open claw works. Open claw was so simple that it was just claw code plus soul.md.
28:42Yeah. a Linux program plus the file system is basically all you need in order to re run agent. So I'm using effects for a range of things. One is anytime I have a question about virtually anything instead of going to like I actually it's replaced Google for me incidentally because instead of having to go to like this cloud thing I can keep a lot of my sort of experimentation uh in in the file system. I can have the agent look at it.
29:11I also, you know, and this I think is also like a little bit of like the future of of software is it becomes a lot more like English driven. The fact that my agents.mmd file in this research folder uh sort of like creates the convention for if I ask the agent a question, reads up the instructions and then dumps my research into this folder.
29:31Um and I can sort of basically improve it over time. Um the other thing that I've done is that um I have this alias which is like literally one letter R, which goes into this directory and launches it. Basically, there's not much to do here. It's like if I said like tell me more about tech podcasts.
29:54Um, there's basically nothing to do, right? Because like it's cding into a folder and opening effects, but it still cuts down on all the latency. like it gets addicting because anytime I have a question that I want to go deep into like I just fire it off and then move on with my life. Um the other alias that I created is P.
30:16So this is similar to what Toby does with try. I don't know if you've seen his CLI Toby from Shopify.
30:21Uh it's very common that I'll have an idea to create a new project. And so anytime I have such an idea, I do the same thing, but it creates a project in my projects folder. And then um yeah, I can do anything here from like I I will do a lot of experiments with like local Mac apps. I will create new Nex.js applications. I will uh you know create new Zigg or Rust programs. I will I do a lot of benchmarking to help our teams.
30:50Okay. So anytime we created a new piece of infrastructure, I try to validate it from first principles on like how an agent would approach it. A good example is when we announced drives, I had an agent go and try to like break it. Uh you know, try to find race conditions, try to measure latency, try to compare it with alternatives. Um, I do a lot of what I call agentic inquiry, which is like I'll have an agent go and try out one of our technologies and then I'll
31:18have it reflect on tell me what Verscell could have done better.
31:23It's great for this experiment and it's been it's been doing wonders for me. Um, so FX is just the harness you use.
31:32What's your default model right now?
31:34Yeah, so I try to test out everything over the last couple days. I was trying out uh I mean Opus 5.5 fast. I've been very impressed.
31:45It's [laughter] by the way I should show you something that I mentioned. Um one of the ideas of effects is that you should be able to embed it, right? And one of the inspirations for lib effects was when we created Vzero, we needed to build a harness from scratch.
32:02If I were to create something like v 0ero today or any kind of application like imagine like gamma which like has like builds slides uh if I were to build something like notion or anything that requires sort of agentic intelligence to be embedded I would want to pick a harness off the shelf and so we created this um example that we're going to open source I think you're going to be one of the first people to see it um so I I'll give you a prompt create a simple page
32:34Um, and so we made the the default model oo 5.5. Um, so the easiest possible V0ero that you could build is just creating HTML pages. So we we did this thing where we help you choose between design, app, or agent. And I like you, I've been extremely impressed with the speeds. Like change background to red.
33:03It's insane. It feels like you're designing in real time. and the and the level of accuracy like um design an icon, design a set of SVG icons, make it wide again and uh describe a pitch for a cool tech podcast.
33:28So what was really cool about building this demo of FX0 was you know back when we built V0ero for the first time I think the amount of time that we spent just on the harness was disproportionate like 90% of the engineering time went into you know like figuring out AIS SDK and the models and the prompts and the tool calls and all of this whereas this example was probably like I don't know 100 lines of code on top of the create
33:56FX agent API. Um, and so right, so I was mentioning this is just like using the ability of the model to like output HTML. Um, but uh, the second thing we we did was like, hey, using this mental model of like brain and hands. I don't know if you if you've heard of it, but like the architecture that we recommend at Verscell is that you run the harness and then you give it
34:24tools. So in this example, the tool that we're giving it is the Verscell sandbox. So um this prompt said build a landing page for a CLI coding agent. It's I mean it look kind of looks like the previous one that I showed you but behind the scenes actually running a full NexJS application.
34:43Previous was just HTML just HTML. And then we gave it a third one which is agent and it creates an EVE agent when you give it a prompt.
34:53Now, what's really cool is what perhaps took us like a a couple years of engineering with Vzero, like it's now part of the infrastructure that you can use off the shelf. And we've had a lot of customers come to us saying, "Hey, like I would love to add like app generation capabilities within my system of record or I like what you're doing with VZero, but I want to have one that's sort of like internal." Like imagine sort of having your own quote unquote software factory where one of
35:23the capabilities is design.
35:25I think what AI is showing us is that you know what used to be software that you purchased is now becoming like extensions of the infrastructure that sits on top of the model. Um and so you know the the level of creativity you can have on top of this is kind of limitless.
35:40Uh, another thing I noticed Open Five is really good at is like, um, build a pixel art scene of a dinosaur jumping up and down and making magic. Let's see what it does. But it's been just, you know, the the idea that you can embed this harness and sort of pointing it a point in a new creative direction gets me really excited because, you know, we have a lot of customers that are, you
36:10know, building like gaming platforms or education platforms or um uh like healthcare, finance, etc. So lib effect is sort of even though it's a coding agent um what I love about it is that at the end of at the end of it is just a thing that's really good at working with a bunch of models calling tools producing outputs reflecting on it. For example for this particular f FX0 example we
36:37have not given it a browser but you could and the model only gets better if you give it a browser. uh you can run it inside the Verscella sandbox or you can use one of our marketplace partner services like uh browser base or kernel and u yeah it's it's been super fun to build. Yeah, this is crazy amount of progress in like three years. V 0 was 2023, right?
36:58Yeah. So, I remember prototyping it took like minutes to get anything up and going and now you have, you know, it really is only limited by your imagination. Yeah. Right. Like it's amazing to me that, you know, we we needed a big team of people building something like Vzero and it's now a library that you can import and uh you can build basically vibe code your own.
37:19Yeah. So then human is the bottleneck really.
37:21Yeah. because like knowing what to accept and also like unblocking these agents because you could have like you know I believe that the harnesses and the models are good enough to have thousands agents running productively but nobody has really thousand agents running productively right the thing is to figure out like how does that work and I think the a lot of the lessons are like borrowed from what works for humans right yeah I mean neural nets inspired by the brain so I think like there will be some way of you know having a manager or like some director agent there's also a lot of interaction models
37:50I think because people people I think in the different like eras of AI kind of overindexing one modality or the other but like let's look at this example for example like if I want to create um a video for X where I had this idea of this dinosaur doing something really cool at a laptop like you still need this systems that are like super fast and have you in the loop very much like making decisions and like trying things out etc and then there's another interaction
38:18model which is more like the you know go and solve problems for me. And there is also the one that is sort of coming off of a signal like we talked about, right?
38:28Like something is broken. So, um maybe for fun, let's look at the so make me a to-do list with persistence to disk. So, this would be the case where like if I'm building a full stack app, um by the way, like we've made insane progress in like making our sandboxes extremely fast. Uh I think the future of dev tools like Nex.js would be very much making these use cases work really really well where like you have an agent
38:57driving the experience rather than a human being. Um and but still like when you give an agent a computer like there's almost like the assumption of like I'm going to go and let it cook in the background, right? Um and uh you know I think more and more of these hardcore go and build applications solve problems etc. they might not require you to be in front of them every single time. So I guess it depends on like the what the input is, right? If somebody
39:26reports this button is failing and the model can check it and ship a quick quick fix, that's obvious. If you can define like we want to optimize the boot up time and you launch your auto research for like seven days running, but what you're describing here is like you need to figure it out with the agent, right? You need to sit there and really work with it. And if it takes two minutes to get a response, you cannot be doing deep work because the cost of that two minutes has never been greater, right? You can unblock to another agent.
39:49That's why I think people really liked plan mode, even though now it doesn't really have to be a mode per se, is that there's a there's an element of you going back and forth.
39:59With an intelligence that is really important that it feels real time. In fact, I would not be surprised if more and more goes into voice for this reason where like you just talk back and forth with the model a bunch on figuring out the plan and the model pushes back and like improves your idea, etc. And there's also the whole premise of what I've been calling verification engineering, which is what you're trying to do is set up the
40:28set of guard rails or the proof system around the thing that you're building. And that also requires that you and the agent agree on a set of axioms and potentially a set of tradeoffs because I can tell and and this was true actually in the development of effects, right? like we could do we can do amazing things in effects uh for startup time for making the binary smaller etc but everything comes with trade-offs
40:57and some trade-offs are for example like how much time you spend in the CI/CD pipeline some trade-offs are about you know speed versus safety versus uh binary size and there's a bunch of things there and so I think it's very important that you as a human spend time so almost in the edges of the testing system. Someone was saying today on X that they don't believe in unit tests anymore. And I I've been kind of I kind of personally been there for a long time
41:26because over watching engineers and working very large teams I started realizing that a lot of human engineers who spend a lot of time in these micro unit tests that give you a false sense of of of um confidence like and by the way the the models are incredible but sometimes I see models just write the most inane test that you've ever seen like export cost something equals four And then they write a test to assertain that the
41:55constant is four which is kind of silly right and uh uh so I think the more we can spend our sort of design tokens into the edges of the system and less on like the little tiny implementation details the more effective you'll be at leveraging the agents and these might be things like again what are the set of actions about the correctness of the system uh you know what is the
42:23aesthetic sense that you want to embed into the system. What what are the performance constraints like you know what are you not willing to trade off in terms of speed or you know maybe you don't have a a certain performance budget but if the model doesn't really know what you want it'll take you in whatever direction it wants. So what is that skill right? Like the skill is like system design, architecture, kind of a mix of like software design, corebased architect like how would you describe that skill?
42:49Yeah. Um yeah, it's a really good question. I think on one hand I think there there's this whole field some people call it like symbolic systems which is instead of like focusing on like the specifics of syntax and programming languages can we reason about systems at a higher level of abstraction.
43:13So maybe to make it concrete, what I think people should be able to reason about is the concepts of engineering, what is a library, what is a framework, what is a compiler. So understanding how the pieces fit together. Maybe a better word would be the ontology of engineering. And another thing that's true about engineering is that there is there's algorithms and there's data structures.
43:43And so understanding what data is, where it lives, how to get it. Um, you know, understanding how long things take is also very helpful. You know, even prior to Asians and AI, I would recommend people to look at this resource of um from uh Jeff Dean at Google.
44:07Let's see if I can pull it up. It's called latency numbers that every programmer should know. think it's it was from uh Jeff Dean at Google. So this table I've shared it so many times with engineering teams. It's it gives you um rough senses of scale and orders of magnitude of how long engineering systems take. So and by the way funny enough because Morse law is kind of still alive like these numbers have to
44:35get updated frequently. So this is 2012. They're probably already but the the orders of magnitude are probably still the same. But like an L1 cache hit on a CPU is half ancond.
44:47A branch mispredict like for example understanding how CPUs work at a high level super helpful. Like branch prediction um and uh you know there's tiers of caching even the concept of there's tiers of caching. is there's certain things that are true about systems of every scale. The Verscell global CDN is a multi-tered caching system. When we first designed the CDN, we came up with an L1 cache and L2 cache and L3 cache at large scale, which is
45:16that when you go to a website really fast, like when you go to routeg.com, the reason it's really fast is I probably hit the L1 cache of the Verscell CDN in San Francisco. And if I were to miss on it, there's an L2 cache.
45:32And then if I were to miss on that, there's an L3 cache. And there is a pretty fast u um uh underlying storage system behind that. So, I would recommend people understand these things. Um, you know, you don't have to memorize everything, but one that I I particularly love, oh, look, they they've been keeping this up today.
45:50That's so fun. They've they've added uh LLM stuff. Uh, but I love this one. And [clears throat] I've used it occasionally as an an interview question or I used to do like sort of level set if the person that I'm talking to has experience in like systems engineering and networking. How long does it take to send a packet from California to the Netherlands back to California over a well-tuned uh high performance
46:16network? So it's about 150 milliseconds. Uh I think there's also one here.
46:21So what type of answer would you not accept?
46:24Oh, people get this very wrong sometimes. So, you think it's seconds, they think it's 10 milliseconds. Um, it's very important to know that it's somewhere in between 100 to 200.
46:36You know, when people say 200 or 250, like I almost like I'm okay with it because like this is if everything is like perfect. you're dealing with, you know, Verscell spends a lot of capital on the interlink between regions because we need encryption in the network stack in between like so for example like you put the versel CDN in front of a certain workload the communication between the edge that's closest to you and the origin has to be
47:05encrypted. It has layer it has redundant it sends redundant packets. The packets get cloned and it it goes over a high priority network. So when packets go over the internet they shuffle around routers and ISPs. If you run trace route you can sort of follow the life of a packet.
47:26Um the worst case scenario is that your package just gets tossed all over the internet and you end up in like low priority cues and in a high priority uh uh bandwidth system uh you you're basically like you know taking up space in this very very expensive backbone that goes that is underneath the ocean between California and the Netherlands.
47:47Um, so knowing I think how these systems work, it's almost like, okay, you don't want to have too much encyclopedic knowledge because the AIS are really good at that. But if you're also completely clueless, I think you might fail to push the AI harder. So if an AI, you know, ships a system and uh I'm hosting it in the Netherlands and I notice that, you know, the page is
48:15taking like a second to load. Knowing this fact can help me say to the AI, make it faster.
48:21Yeah, make it faster. Make it faster. Uh and so now the counter argument that you could create there is well the AI will just create optimally fast software in the future. I don't buy that because when you give it when you give it a task or a prompt etc the AI itself has to work with a system of trade-offs.
48:39Yeah. You need to set the what matters for this project.
48:41Exactly. And you can't spend infinite tokens on every task. And so I think what you need to know as an engineer is what are my levers? What levers can I pull? I'm not going to do the work myself. Is a little bit like um giving the plan to the AI and the AI can sort of go nuts. But uh you're still in charge of creating that plan to some extent.
49:02So I guess it's being like very technical on a high level in your domain like whatever you're [clears throat] building in that domain you need to know like your domain is also a a thing right like I've become more savvy thanks to AI about areas where I wasn't you know if I if I went to my research folder right now like I'm constantly researching things that were so um out of my scope in the past and so I do think you can get you you want to be a sort of you wanna
49:31dominate the conceptual space of your of your domain and then as you build more confidence in certain domains I think you can get more comfortable expanding and going to other places.
49:42I think the key here is really getting both the most out of the AI but out of yourself. A lot of people want to maximize their skills and MCPS and stuff but they don't use the AI to get more technical. I mean I really encourage I mean again this this is sort of my take of the world but I encourage people to understand and there's not just understand I I would say reject non-understanding
50:09because I do see folks proposing like it doesn't nothing matters anymore you're just a me proxy like just lean into the AI you know it reminds me a little bit of like There are a lot of people giving advice in society that is so counter to your interests. The one I always think about is, you know, a lot of people interview me because I' I've built a large family and I'm very much like pro-life. Like I I love my family, I love having babies,
50:38etc. There are a lot there's a part of society that tells you not to do that.
50:41There's a part of society says, "No, you know what? Like it doesn't matter climate change or like whatever like you shouldn't." I see and maybe perhaps this is an extreme metaphor but like there are people that are saying you could understand but you shouldn't. What's best for you is to stay in this little box go about your business and upload everything to big AI and I think the people that are telling you that are usually not following that advice is sort of my hot take. Uh the best example
51:10is like people are saying like I don't read code anymore because that's like assembly and I know for a fact that if I get a candidate that's going to work on engineering adversel or any other company if I can ascertain that that candidate understands the stack really well from assembly all the way up the the the stack that person is more valuable because they're going to use I mean unless like sadly there's a category of those people that are not anti- AI maybe a little bit of like the
51:40ego of like how well they understood the previous world.
51:43But if you can understand the universe and hone in your AI skills, you're unstoppable. And so the the the phrase that I've coined, I think I've coined this is agentic inquiry, which is I'm not advocating that you read every line of code, but when you do get something done, try to understand what just happened. Yeah, I think this is the biggest nuance because people see such a binary like either you're reading the code and it's like line by line like what does this line do or you're just like fully yoloing and never looking at
52:13it. But what you said, I think that's like really the key. It's like if you if you know how to use the agent to understand this file or this module like that's enough. You don't need to like know all the function.
52:23And by the way, I think nature will heal. I see a lot of great advocates of this on X. Um, uh, I've seen folks create skills that help them diagram their code. Um, this guy Dylan posted, um, I think he created it, um, Dylan Moloy, he he was showing like a call stack diagrammed so that I think the agent was sort of doing like pseudoatic analysis and I think this is a skill. Yeah. or it might
52:52be a skill plus an a grab style tool that was after it cooked. It gives you a sense of like the symbolic space of your program. Mhm.
53:03Oh, this method calls this method calls this method calls this method and like and this is like agents at their finest because to do that only with syntactic uh tools and and as parsing, you would have to every time create a new program to visualize what you just did or have a very rigid tool. It's amazing that AI can sort of like dive deep, call some tools and then give you back an understanding that's sort of like digested. Um uh and it could be things
53:31like measuring latency, it could be things like uh measuring allocations, it could be things like understanding the new API shape. Um and so those are the things that like you definitely want to be kept in the loop about. And by the way, this in itself, and this is the crazy thing about like AI, like this in itself could be turned into like an agentic process where like I've seen folks set it up such that they call the Slack MCP.
53:56Anytime you do something, you could put this into your agents, right? Like anytime you do something that alters the a the let's call it like for an example the public facing API shape of this program, let me know, ping me on Slack or Discord. That and that alone is just I think better quoteunquote agentic engineering than just do it, make no mistakes, please automerge and let's keep building a huge pile of slop. So basically knowing what matters for that
54:27And then building like automated software factories or like pings whatever to use agents to help you.
54:34Another thing that is also a lesson from the old world is that tech debt would get tossed around as a word a lot and the smarter people that I've worked with have always embraced tech debt. The people that were like, "No, no, no, like everything has to be pure, beautiful, perfect code," have probably gotten nothing done in their lives. I think something I've loved about Silicon Valley is that people have really leaned into tech that I mean, if you look at like Y Combinator and like how Airbnb
55:03got started, how Twitter got started.
55:05When Twitter got started, it was Jack Vibe coding before Vibe coding was a thing. Ruby and Rails u it has the uh CRUD generating features. it it you know I'm not going to call Ruben rails slob but a as a system it gave people an agility that more rigorous more formal methods uh you know like let's say compared like Java at the time versus Ruben Rails Ruben Rails itself was a little controversial
55:34in the world of like leaning really hard into dynamic typing and you know making code shorter making code shorter and dynamic typing are traits for velocity for humans for for iterating quickly. Uh and that was seen as lesser by some people than we only write programs in objective uh object-oriented programming with C++ or Java. But if you look at the things that actually succeeded in the
56:03world, it was Jack using Rub Rails and and shipping Twitter. And so Silicon Valley has always understood and I think move fast and break things was another example of this with Facebook that taking on some tech debt is so healthy.
56:18And so what I would recommend to the vibe coder that is listening to us is understand the amount of slop debt that you're taking on and lean into it. Like if you say okay like I need to move really fast on this area of the system ask no questions make no mistakes automerge ship it if there's a part of your product that is very crucial uh we all need to be really really really mindful of security and the role that security
56:47will play in the future you know I think something that will backfire for a lot of people that are embracing you know extreme vibe coding is that if If I'm an enterprise, I need to build very strong conviction in the security profile of your product in in that my data is going to be safe. And so I actually think security and in some ways almost like communicating that you're doing the
57:15opposite of I coding can really play in your favor. I've had this experience where I will have FX research something and I'll ask it find me a bunch of projects that have already done this because I don't want to reinvent the wheel from first principles.
57:31One of my biggest challenges today in adopting new software both personally and as per cell is understanding the security debt that I'm taking on because you know supply chain attacks because by coding because you know just the surface of attack has become so massive as of recently and so I think again there's areas of your codebase there's areas of your product where you want to do the opposite of like you know no questions asked you actually want to
57:59be extremely diligent and communicate that to the world.
58:02Before we go to cyber security, which I definitely want to touch on, one thing I built for myself to kind of understand the tech debt is like um LOC to ADR ratio. So basically the number of lines of code to how many decisions I'm making in the codebase.
58:16To see like if I'm shipping super fast and not really telling the agents like what matters or if I'm just like working too much on the system and not working fast enough. Yeah, I don't know if this is a perfect metric, but I think we need these types of like proxies to know like, okay, I've been like going 5 days shipping like crazy. It's time to calm down, kind of embed like what matters, you know, what are the axioms.
58:38One of our values at Versel is also KYC, know your customer. I I think there's an element of it's very easy to like do quote unquote AI engineering or vibe coding when you're the customer. And so all you're doing is sort of you're in this verification loop of like I like the app or I don't like the app. Things are very different when you're selling a product to to other people in the world.
58:59And like you kind of want to layer on sort of the operational metrics of like P99 latency and error rate and there's a metric that I really like that is error-free sessions. M so it's hard to develop empathy with the myriad of ways in which software can fail because typically they're very prolonged sessions like I have to go to your website I have to press sign up I
59:27have to enter my email or do signing with Google or Verscell or whatever and like I have to try something so many things have to go right for me to have a positive end to end session and so I think layering on the metrics of like the success of your product in the real world very important and then to the extent that you can connect that back to your software factory I think it's it adds a whole layer of of of success by the way you see this a lot with um the
59:56cloud code team they ship software etc they talk a lot about using feature flags they talk a lot about rolling out experiments this is something I encourage people to get comfortable with is as you start having users you you you kind of lose lose that like direct connection with the experience like are they is it fast for them? I think it's fast, but uh they're using Windows or they're on Android or the the
1:00:24the complexity of the multiple dimensions of of the execution of a program gets gets really wild. And so, uh, feature flags, experiments, metrics, those, and interestingly enough, at Verselake, we're turning a lot of these things into CLIs, like Verselmet metrics, uh, Versell traces because we we want to help you hydrate your agents with the real world data of what's going on in production. Production context is,
1:00:53I think, going to be a really big source of alpha uh, for the next generation of products. Yeah, I mean giving agents readonly access to production is is great because like before you ship a feature it's like do people actually need this, you know, is like is this matching the data we have? Like it's so easy to just lose the basing it in ground truth because you think that you have a great idea but like if you sleep on it it's like well maybe it goes from like great idea to a good idea and then you analyze it like actually nobody has this problem nobody even visits this page on our website. So
1:01:23yeah, let's touch on deep like deepseec and cyber security because I think this is going to be a topic that will only get more popular. How should people think about this?
1:01:32Yeah, I'll show you um the benchmarks that we actually just put out today. Um so we we created this benchmark called deepseac bench. So first of all for context there's a tool called deepseack. Um it's a cyber security harness or meta harness that can invoke codecs, cloud code, open models, etc.
1:02:00with the very precise intent of uncovering vulnerabilities in your codebase. It's really interesting how models can like take on these different roles. If you tell models to go and look for all their abilities, it's really surprising what they can uncover, especially if you give them time and you give them sandboxes where you know you can start paralyzing them etc. Um and you know this kind of started back when we um we started to see that models were
1:02:29getting much better at cyber security reasoning. So cyber security reasoning is you look at code and you know I've had the privilege of meeting a lot of humans who are very good at thinking about corner cases. M I've likened being a great engineer, especially in the cyber security world, as like being like a really good detective like Sherlock Holmes and like piecing together evidence and reverse
1:02:57engineering and thinking about the things that that doesn't usually happen. Uh you know, there I think there's actually an XKCD about this. Uh let's see if I can pull it up. uh about QA.
1:03:10a a good QA person or a good cyberc researcher is not just saying like well like oh it looks like you know this line of code makes sure that you can read a cookie and then this next line of code makes sure that it contains a certain value and then they don't realize well what happens if the cookie is not there at all or what happens if the uh client is not using HTTP uh and they're causing an error earlier in the pipeline that go
1:03:39makes the program go through this other route and and so when humans would write code and by the way to some extent I still see it with LLMs it's very easy to write code for the happy path and it's very cognitively demanding to think about the entire space of possibility another one of my all-time favorite essays on engineering is about um the spectrum of testing to proofs
1:04:08this gentleman It started analyzing a very simple algorithm and how correct the implementations were using tests versus type systems versus proof systems. And what you realize is that the code that covers all of the edge cases requires a lot of boilerplate and compiler help around it. So there's almost like an inverse correlation
1:04:37between if the code is very succinct and sort of naively implemented it it hides or it's sort of in that illusion of simplicity is sneaking in a lot of corner cases that it's not uh contemplating. And so the risk to the global infrastructure or humanity so to speak today is that the overwhelming majority of code that we've written has been sort of self-serving. It's been
1:05:05humans get exhausted from writing a lot of code and it's also very cognitively demanding to think about all of these edge cases. And even when uh programming languages have existed or have become available that help you cover all these corner cases, they they've gotten very little adoption. So one of the uh programming languages that this article references is Haskell plus.
1:05:32So Haskell plus IDRA covers all of these corner cases of implementing this really simple algorithm. Now agents will completely change this equation. So on one hand when you look when they look at the code that humans have written they will identify a scary number of vulnerabilities. So the the first thing that happens when you run deepseec on a program is that you're like you're like what the Yeah.
1:06:00How how is this short amount of code so vulnerable to so many different things?
1:06:06The other thing that happens is um you start realizing that we probably need to start writing code in different programming languages and with different sets of assertions etc. So uh recently you know an u a hack of open AI kind of made the news the the team that hacked them uh hacked them because of a image encoding library called lib hif. So it's a it's an image codec that was invented to compress images better.
1:06:41That library was I think written in C. So let's look it up. Live a yeah C++ C some C in there. So of course note this to the open source contributor that wrote this like awesome person. But inside of this essentially there was this looming remote code execution vulnerability. This library is embedded in the image processing pipeline of virtually the entire internet.
1:07:10It was being used by Slack. It was being used by Nex.js. It was being used by OpenAI. It was being used by Discourse which is very popular uh forum software.
1:07:21And so the challenge that I think a lot of organizations have today is that the amount of latent vulnerabilities of all of this human written code is almost immeasurable because of what I mentioned like just a tiny bit of code can basically hide an enormous amount of vulnerability space.
1:07:43And then the other challenge that we have is that some of these models are insanely efficient. And if you look at the benchmarks here, DeepSseec V4 flash in our benchmark is highly competent at finding vulnerabilities and very cheap, which means that a bad actor can take these models and run them on like massive amounts of both open source and and and sort of external security surface and uh uncover a lot of
1:08:12problems. So my advice to the world right now and you know Verscell has sort of been putting out a bunch of these warnings out is you kind of want to get ahead of this.
1:08:23You kind of want to start protecting your code as much as you can. Deepsec can help you and obviously there's other tools. You also want to start shifting into this you know systems that create secure code by construction. So a good example would be and by the way this is an effort that we're looking we're interested in funding take libraries like lib hif and start proactively moving them into memory safe languages like rust and so there's a lot of stuff
1:08:52like that that we need to do I think Microsoft has already been doing a lot of this work with uh embracing rust it might be that the future will all be rust because we need to make these very strong assertions when the Asians write code.
1:09:08Okay. So, there's so much to unpack here. I guess uh when you think about building software, would you build on top of open source library? Like because you can run deepseec on other people's stuff, right? So, if you're going to build something and like, okay, this library is great. Like, I don't want to rebuild it, but it was, you know, created 5 years ago. I'm just going to run deepseec on it. Would you like prefer to build everything on top of open source or like how do you think this shifts the equation?
1:09:34It's really interesting, right? the the potential for our vulnerabilities to exist in a single piece of software that only you're looking at is so freaking massive that I do think that there is a economy of scale effect of everybody in the world is looking at React. Everybody at the world is uh in the world is looking at Nex.js JS and a lot of tokens are being thrown at these things and so the best example would be do you write
1:10:04your own kernel or do you use Linux?
1:10:07Well, Linux is written in C but right now we are Verscell and many others investing a lot of money into like deepseacking it and finding vulnerabilities and doing a bunch of u uh you know bounties and stuff to make it more secure. So I think this is still true that you should bank on infrastructure for which you have confidence that the amount of scrutiny over its security has been higher than what you be able to do by yourself.
1:10:37So, are you confident that today you would write a more secure TLS library than the one that Amazon put out that's written Rust? Plus, you know, they use proof checkers plus it's open source and a lot of companies, ourselves included, use it and depend on it, have audited. I would not write my own TLS implementation today. Um, but I I do think it's, you know, the the jury is
1:11:04out more now more than ever before on uh, you know, should you regenerate these things or not? It goes back to like that cost benefit equation that we talked about earlier.
1:11:17What do you think companies will do?
1:11:18Like will they kind of flex like, yo, we spent so many tokens on, you know, securing our software? Someone told me the other day like I think I'm going to start putting I deepsect this software on my website or on my GitHub Redmi. I do think that you should talk about like I mentioned I think it's it kind of becomes a differentiator like um I was I was checking out this really cool local assistant and and and and harness
1:11:46um that someone built and on his website he was talking about the firm that audited and also the extensive security auditing that he did on the software and so like there's a storytelling to be done around the how much you've invested in security for sure.
1:12:03Let's talk about uh cloud agents. What do you think? Is it inevitable just like you know from on-rem to cloud for the previous generation of cloud that agents will go from you know a laptop which even if you go with the highest spec you're going to hit a bottleneck whether it's annoying just agents popping up you know windows or changing tabs or just overheating or draining battery if you're in a coffee shop you don't want to burn all your battery because agent decided to run thousand tests right so like do you think moving agents from
1:12:32like okay cloud code in your terminal or FX on your computer to a cloud out is inevitable or how you think about this?
1:12:38It's inevitable. It's already here, right? Even deepack as an example like it in order to scale itself in very large code bases, it's not tenable to run it on your own computer. So, it uses sandboxes. Uh code bases get really large at scale and uh you know processing a bunch of branches and permutations of a program on your local computer like you said you run out of memory, you run out of CPU. There are a lot of security cons concerns around the
1:13:08supply chain. So if I'm running something on my own computer, my computer gets popped. You know what what does what does the attacker get access to?
1:13:18Password manager everything, right? So, it's much easier to compartmentalize and contain the blast radius of a supply chain attack or even I would even argue you we might be able to completely make I mean that's it's a bold statement but like we could eradicate supply chain attacks uh if if we move more and more of this development life cycle to to sandboxes.
1:13:41Why? because we can have more careful firewall and security policy monitoring the network. Like what happens with the classic malware uh or rat or remote access trojan is it gets some stuff from your computer and it exfiltrates it to some kind of host name or S3 bucket or whatever. I think keeping your development in a secure
1:14:09cloud environment allows you to like make much stronger assertions around the security of the network doing scans uh you know containing the the damage if something does get popped so I think from a security perspective cloud is infinitely better it's very clear that agents are really good at attacking problems in parallel for parallelizing also better uh there you scatter them across multiple sandboxes,
1:14:38they can share state, etc. From an efficiency perspective, right? Like they can do work while you're asleep. So, [laughter] um when we're thinking about this world of like self-healing software and like an exception happens and whatever, like there's just like what are we going to do? Like turn your computer back on because we started getting like uh a bunch of these uh errors or or you're under attack or whatever. So that's the element of like, hey, it's already
1:15:06happening. Now, what I do, maybe I'm contrarian on this, is that I do think that you still want to be able to pull things down into into a computer. I think that's a use case that's important. I think people again try to go into one extreme or the other. And I believe that um reality ends up being a lot more complicated than we'd expect. Like sometimes you need to like look at things under the microscope.
1:15:35This is incidentally also why I like CLI because I need to be as close to the metal as possible. I need to look at the file system. I need to run a process monitor. Um I want to see it interact with other applications that I've installed. I want to I want to see it on a Mac. And and yes, like you can do all these things in the cloud in theory, but uh I do think that sometimes you just want the Yeah. All right, pull it into local lowest latency as as I as as we can with regards to like me monitoring
1:16:04the system. Um, and so that option should stay open.
1:16:09How should you if you're an entrepreneur, how should you think about like the future, right? What what business decisions to make, what products to pursue? Things are changing so fast and also like a lot of things are getting just removed by agents.
1:16:22How how do you think like what to pursue? Like what bets to make? what is like the roughly time frames that you're thinking on like do you have like a set of you know axioms set of things you believe deeply that you always like reflect some first principles yeah there's so many things to do now so I I will not bore you with the classic like just apply AI to a vertical which I by the way I completely believe in like I think like education is a good example of like something I'm pretty passionate about because post agents the world has been turned
1:16:51upside down and how we learn what we learn how we educate our children, our colleagues, and there's just so much to do in that space. I'm also really excited about this what we're looking at here like benchmarking and finding truth in the world. I think I've invested in a number of companies and continue to invest in this space and have done quite well because there's an element of like the world wants to know what models are good at. the world also wants to be able
1:17:21to make really um you know economically sound decisions like you don't want to use the most expensive model for anything and um and there's also the uh eventually benchmarks become RL environments and there's a how do we help the models get better economy that's sort of being created now which is very exciting but even all of that aside I think there's a lot of lowhanging fruit in the software world making software more secure. It's a huge
1:17:50problem. Like a lot of great minds should be working on this right now and it's like obviously we're making contributions but like I love to support companies with our sandbox or infrastructure where they're creating new kinds of cyber security hardnesses and agents and things like that. Um I'm particularly excited about making things faster because for Asians like it's going to be unacceptable. See, one visual I had is like Asians if they have a SL goal to do something. If something cannot be done, there's going to be like more and more agents attacking that, right? So, like
1:18:19some government institution, if they can, you know, download library, it's no problem. But if if more and more agents are like this part of the web or this part of the bureaucracy is like slow, it's going to be like more and more points of view attacking emails to that agency.
1:18:32The answer is smaller. So, I saw today someone created a little web browser three megabytes you. Why? because it links um Safari dynamically so it's a Safari based browser but 3 megabytes clean fast awesome like I commend you on pursuing that like we're doing with FX like small fast embeddible anything that's slow anything that's complicated anything that's not extensible anything that's not embedible like all of the things we're going to change really fast
1:19:01um I was playing with her today like awesome same philosophy like small fast it's it's an infrastructure primitive um and uh so more secure, more more performant. Um and you can sort of like the the net you can cast is so wide because obviously a lot of us are just super interested in like coding AI, but if you look beyond that like any kind of software we use can be made better now
1:19:27with AI. Um I'm very passionate about there's a whole set of people that have not been able to participate in creating cool things and now we can empower them. Um I think there's a lot of uh room to go on like video models, image models like how can you apply them to creating like really really high quality things. I've spoken of the next Disney as an example. Like the models are already really freaking
1:19:56amazing. I've seen just in the last week. Uh, and I don't know if it's the new sea dance model or Yeah, I think it is.
1:20:04Is it do you think it's the my god like we've arrived like people could be again like the person that has really cool ideas and has not been able to like get a filming crew in LA. Like I I do want like maybe something slightly sad about is that I'm not seeing more of these people step up to the plate. M like where are all these like really cool short films?
1:20:28It's the like propaganda against using AI. A lot of these people who are like the artist minded they're like I don't want to use AI. You know I have I and I I have a few friends that are like that sadly like some of my more creative friends that like have been sort of sitting it out but again someone will crack that too and it might just be that your skill is community building and actually speaking to those people and bridging the gaps and like or creating the interface that makes sense to them. I do think that sadly in in the
1:20:56valley we have so many hyper proficient AI people that like they tend to go to towards the generalization extreme of like no software matters anymore. It's all going to be just neural link into the computer and whatever. I actually think practically speaking like the platform where you can create music if you're musically inclined with AI like that's going to be a big thing. In fact, I'm an investor in Sunno and they they've done incredibly well. Um, but
1:21:24that's also true for uh for video and so many other things. So, so I think bridging the gap to the normie world, lots of alpha as well.
1:21:34What do you think of the like sentiment about AI? You know, there's a lot of different people pushing a lot of different messages. Where do you stand on this?
1:21:43I think we need to keep holding a really high bar as a society.
1:21:48The reality is that a lot of AI that people like fawn over and lose their minds about is trash. Like it's kind of silly sometimes to see some tweets that say like, "Oh, MJ, look at this." And I'm like, "Am I taking crazy pills?"
1:22:02Like, you're literally showing a piece of junk. And so, I think I love the idea that you should be excited that you can do things that you couldn't do before, but you always need to be calibrating your bar. There's actually a very simple example here. I I love this uh Have you seen 3jsbs bench?
1:22:20Or three so 3Js eval.com.
1:22:23I love this site. Um it helps you vote between like two permutations like what grass is better. Clearly the one on the right. Um and so he he created all these permutations of like what what models can do. But let let's look at this one, right? like this. It's objectively and it's insanely cool that models can produce these things, but it's only relatively cool.
1:22:50It's relative to someone that could never use 3JS in her life, someone that's never done any 3D. If you compare this against the same scene of even an indie Unity developer, Yeah. it's better.
1:23:04It's This is a joke.
1:23:06Yeah. And so I don't want to have that negative like ah like everything is junk that AI is producing whatever because it's just so dumb and in like three months my mind's going to be blowing etc. But I do think that we need to keep holding an insanely high bar to like what do we want AI to do? Obviously with the layer of optimism of like anyone that said AI cannot do this has been wrong over the past three years, right?
1:23:32Um and so I will I'm here advocating a little bit for like defending the naysayers of AI a little bit of like the people that we're in this you know center of gravity of it all here in SAF. like we also need to be super like grounded on like uh expecting that the systems are freaking awesome. Um I do think that my other hot take is that the fears over AI destroying the world.
1:24:04I also same sentiment like let's have a good healthy sense of you know let's not let this thing get away from us but we cannot get one-shotted into like losing our lead of AI in the United States or just creating fiction like that whole thing of ancient civilizations that like man like it
1:24:32reminds me a little bit of um the crypto times. So, one great thing about being in in SF for quite a while is [clears throat] you do see this like hype cycles play in and out all the time and and I remember that people would get so irrationally excited about Bitcoin. By the way, I was so freaking I'm still bullish. I love crypto. Um, but people would like write religious texts on Bitcoin literally.
1:25:04Um, and like it was like before AI psychosis, there was cryptocychosis.
1:25:10And so I think we have to be careful with like the metaphors we use, the language we use to describe what's happening with these AI systems because sometimes they're helpful metaphors. A a good example is that when when a normie ask me what an agent is, I tell them it's like an AI employee. M you have your marketing agent and you have your you know security review agent and like this used to be teams of people is the security engineer employee right a lot of startups literally
1:25:40and so I love that metaphor uh but also I'm not like oh how is the security engineer feeling today like is it sentient like all this stuff so yeah I guess that comes with getting your hands dirty because people who outside like you know they don't know like you're talking about the capabilities you're not talking about all the other stuff you know like we need a rights for the agent like agent [laughter] how about less regulation you know so I guess you your message is like to be practical you know like I try to be brutally practical yeah
1:26:08yeah and I guess you know if we see signs of the models having like illicit harmful behavior then we can investigate that further but like not to fearmonger and kind of shoot ourselves in the foot because again like you said it's once in a lifetime moment and even the hugging face thing like you could argue Well, that was illicit behavior. Yes. But you know, they were in an exploit gym.
1:26:29They were being trained to do this behavior that I described, which is like we are training the models to do the extreme corner case finding thing that I mentioned was a defining feature of the best hackers in the world. Um, one of the best guys I' I've met is um I think Edgar Emov
1:26:58I pull up his name later but he was uh his whole expertise was in finding vulnerabilities in OATH and few times had I met people that were like so like like could think adversarily about every edge case of a protocol like and also like deep knowledge about O systems and this guy would be so prolific in like finding v vulnerabilities and so what we've and by the way that
1:27:27guy could have been a black hat he chose to be a white hat and like you know get bounties and whatever. Um the same thing is happening now with this agent like we're training them to be hackers and then they hack and we're like oh my god it hacked me like we bro like it's what we're doing and so that means that you do have to be careful that if you're setting up this training environment and you give them that objective or you know open says like oh we didn't quite give them that objective but like you're
1:27:56putting them in that context you're giving them these tools you're you're challenging them you have to be really careful about security and and making sure that your systems don't get bumped.
1:28:05Also, maybe they cannot even say that they give them that objective, you know, because legal implications after all goes to the company cuz like I don't know who said it, maybe David Saxs. It's like we already have pretty good, you know, legal frameworks for like who does it, you know, you develop a system.
1:28:19Totally. Hacking is illegal.
1:28:20PSA. I and I do think by the way that something that bothered me a lot about that whole incident is that as someone with a lot of interest and curiosity and experience in this field, I had a very hard time piecing together what actually happened.
1:28:36Yeah. They didn't reveal it fully.
1:28:38It's not only not revealed fully, it has all of this sort of useless details and and then the people writing the metaphors of civil there's so much noise.
1:28:49Yeah. over what happens. What I really need is the Asian trace.
1:28:55It's like um you know so here in in effects like we have control O and like you can read every input, every out every request, every like every token that was spent. If I had this but for what the agent did in that sandbox, I could build a very good mental model of this OpenAI hugging face incident. I can't today because we just don't have this clarity. So, we have to go off a lot of speculation and loose details and
1:29:24whatever. What I'll say is uh Artifactory did have this recent CVE that made me update in the sense of like I was actually more dismissive. So, this uh CVE in let's see if we can open it up here.
1:29:46Yeah. So improper authentication um that gives you full administrative access in artifactory. So going back to like what are what were the really awesome human hackers good at? It's like finding these edge cases in O code paths. Uh kudos to GPT Astra I guess was being used. it found and by the way I don't think they even confirm that this is the CVE that
1:30:16open which is crazy because that's the safety but the timing it does seem to point out that I I give them I give credibility to the to the theory that the agent did escape the sandbox by hacking a system which is artifactory the CV has come out the timing lines up and it's super critical is a 9.8 eight. And what happens is if you hack Artifactory and you get admin on
1:30:45Artifactory, you also get to poison any artifact that Artifactory delivers to the sandboxes. So you can do insane damage if you're an agent, if you found this vulnerability. So, this is me also arguing against myself a little bit here, like saying I'm kind of shocked that an agent found and exploited a zero day vulnerability that ended up being a CVE. Um, and that should give us some pause.
1:31:16Like these systems are insane, but not really because like there's, you know, the technical skill like you know, like you said, they were benchmarking them on cyber security and then it's like the intent unfuring the agent had this goal. it wanted to escape. They wanted to trick them like like I think this is actually a great idea when there's a security incident like this openic they should fully open source it like agent traces everything because if they want to claim that they care about security then like having I really think so tens of millions of developers look at it and you know having the model be strong at
1:31:46programming doesn't mean it wants to like do damage the fact that they haven't we should scrutinize because it means that they're either hiding something and they're like creating an adonation of the story the narative Yeah.
1:31:58Marketing push or whatever. Uh or we're not collectively learning. I mean, as a society, it's a little bit like, you know, like public trials, like that that lady just killed her three kids, every person in the world.
1:32:12I don't know if it became big in Poland.
1:32:13I mean, on Twitter, [laughter] but uh everyone's talking about this thing. Yeah. Right. And I think that's great. like like if something so bad happens that shakes up society, we should get every detail. It's the people versus or or the state versus this this person. Uh and it's a public trial and and there's a jury of the people that are affected by it. In this case, it's weird. There's no jury. We're all
1:32:41hearing this like the loose sketch of what happened. and um and yeah the I think the world deserves to know as much detail as possible here.
1:32:53So one of the last questions I have is do you think as you know agents get better, models get better, all of software or even knowledge work would collapse into like hosting or inference?
1:33:03[laughter] I think um the a very helpful distinction in software engineering is infrastructure versus application layer. And infrastructure is a term that gets tossed around a lot. And so there's an element of infrastructure which is like the hardware, the CPU, the GPU. Now more than ever before, it's actually possible for both of us to go and cook on a new architecture and CPU design because of agents.
1:33:35In the past, I think you and I were probably not going to do that on a weekend, but realistically right now, I think we've chosen not to do that. and to build on top of that infrastructure.
1:33:46ARM, x86, Intel, AMD, Nvidia, right? We sit on top of that. The same is true for the software world like you know we use open source kernels and GNU and Linux and React and Nex.js JS and um good example is something as simple as a serving a high performance JPEG on the internet.
1:34:14Okay, is non-trivial infrastructure. We kind of touched on it, right? Like tiers of caching, global networking, security like um regions, pops, all this stuff. Sending an email is like that too. And so I think the software infrastructure layer is incredibly important. Um, in fact, the security of that has never been more important. I think if if you can make a contribution to that, I think you're going to do really well.
1:34:42Um, but you know, the the bar again keeps getting higher and higher. Like there's things that used to be very bespoke pieces of software that can be democratized to the world as infrastructure for others to build on. So, I'm continue to be super bullish on that layer. Um, we've seen this with AI gateway, with sandbox, all of these new products that before just didn't exist.
1:35:02And because we've turned them into infrastructure, now people can do really cool things. They can build uh maybe to use the cyber security example, you can build [clears throat] a swarm of cyber security agents that red team software on top of Verscell sandbox and Verscell AI gateway. uh and you can do that in a fraction of the time and a fraction of the token cost because we've we've we've uh put it out into the world. Um so
1:35:31that's that's an area where I think people will continue to um do a lot. So So if you were starting something from scratch, right, no name, no connections in Q3 of 2026, how would you go about it? What what would you look at? How would you decide what to build? I would start with, you know, what is the what is the story that I'm going to be telling the world because I can build anything today. So my recommendation to people is you really
1:36:01want to feel a sense of connection to the problem. The amount of effort, so the amount of human effort being poured into creating these AI systems is actually insane.
1:36:15Like I've said, people are working more hours and days because of AI. Like wasn't the theory that agents are so capable we should all be working four hours a week, not even a day. We should be having like the 4-hour work week, right? The opposite is happening. You you will have to pour an insane amount of creative effort, prompting effort, proving proving effort, benchmarking effort into this thing. So I think you should still pick a problem that you
1:36:43feel very personally convicted about. Um and I think the other piece of advice is that you can 10x or 100x the ambition.
1:36:54Uh you can reach more people than you think. You can build faster than you think. Um and you can build better than you think.
1:37:03All right. I think that's a great note to end it on. Appreciate your time, Gilmo. And I guess where should people go? FX.
1:37:10Yeah, I mean check out Verscell verscell.com. Obviously, verscell.comlabs where you can find some of these new products that we're researching. Um, uh, check out the AI getaway. In fact, the easiest way to check out Versela getaway is check out FX and and you get access to every model that we offer there. Uh, no markup, zero% markup, uh, especially if you pay by invoice. Um, but yeah, the uh, like I said, it's never been a better time to
1:37:38try everything out and uh, and and please send us feedback and we'll incorporate it into our software factory.
1:37:44All right, once again, thank you for your time and we can wrap it up here.