0:00and just prompt claw. I don't need you.
0:01When you work for me and I'm paying your tokens, you better be like actually producing thoughtful stuff and I have two or three employees right now who are basically under performance review because they are just giving me cloud slop. If your videos went up like 14.7% week on week and I say why and he doesn't know. He's actually never watched the videos. He just used Claude to like analyze the videos and claude doesn't have video analysis.
0:23We're spending over $100,000 a day um on token. the the agents themselves have been able to completely do closed loop uh iteration and testing on on the token spend side. Um it's become you know our single largest uh non manufacturing line item very very quickly. We didn't tell anyone to cut their spending or do anything even though we were seeing this exponential growth in in our our token spend.
0:47First the news the curve debate drifts toward runaway AI. A frontier executive entertains caps on intelligence and Anthropic's original restraint pledge comes under scrutiny. Short pause. Then Thomas Sr of Posatron AI on what a chip design team actually buys with its agents and Swixs on what an employer still pays people and vendors for once anyone can prompt the model. Long pause.
1:15To close, the hosts on Claude's advance on the threesome problem and whether AI audits could trigger mass retractions in science.
1:23Light Haven had a conference called the curve uh over the last few days. What was your takeaway from the curve?
1:28It's a really interesting and pretty unique event in the AI space started by a couple people um who had the idea that we need to bring more contrasting opinions together. So can you speak to some of the divides that you saw uh over over the weekend?
1:47Pro- open-source uh because it's going to save us from concentration of power issues, anti-open source because it's going to create you know bio risk. But I do think the big thing that people have kind of come together on at least to a significant degree is it does look like the AIs are going to get really powerful and this is something we're really going to have to reckon with and can't just sort of um hope goes away or you know feel like maybe uh if we just wait you know the bubble will burst or the fever will break or what have you. And the
2:16question now on the capabilities front has gone to like is RSI going to fume?
2:20Is it going to you know be fast for a while and then kind of level off? Yeah, we've got these super powerful things, but is it really going to kind of run away from us? Those were still kind of the capabilities questions that were getting asked. But there has been, I think, you know, some healthy updating across the community just based on what we've seen from the Frontier Labs. Were there any kind of um limits or constraints that they say that
2:46they said you know should be enforced on uh players in the space including themselves and others. Yeah, that was a huge topic and so actually had a chance to talk to and hear from um in some cases in sessions, in some cases just in you know kind of passing conversations all Chattam House rules who all kind of abstract away but there were a number of things that I actually thought were like kind of newsworthy.
3:13um one let's say senior executive at a frontier lab um first of all was was saying that just pre-training continues to deliver uh that the models are just at the base level of capability getting stronger and stronger and you know there doesn't seem to be an obvious stopping point to that scaling laws are holding or maybe even bending because the data quality is getting better data you know seems to be definitely really
3:42delivering give the model a bunch of data and think like what what could I transform this data into that if I then pre-trained on it would make the model smarter and you're sort of converting these reasoning tokens converting back into uh pre-training data which then you know goes in of of course to the next generation of models. So this is like one way in which recursive self-improvement is happening. you know, RL is important, too. But his take actually went as far as saying he
4:10believes that the latest frontier models do have the ineffable research taste that would be needed for them to start making paradigm level breakthroughs. He believes it's currently an elicitation problem where the RL is not bringing it out well enough. They don't know quite how to bring it out well enough and so it's kind of rare. They'll figure out how to elicit it. And then he said a couple things that I thought were just like legitimately newsworthy. One was there likely is, let's say, a level of
4:40intelligence that we just shouldn't go past.
4:46Yeah. first time I ever heard that from a Frontier Lab leader, but like this was, you know, person who um again like everybody would recognize, you know, their name and and position um saying I think it is likely that there's a level that we shouldn't go past. Now, was that not ever or just, you know, not right now? Um I think that was also a little bit ambiguous, but it was a it was a pretty firm statement. Definitely
5:14the the strongest statement I've heard about imposing kind of a hard cap on capabilities even for a time. How do we operationalize that? Would you be prepared to sign on to something like a limit on the number of flops that go into the next pre-training run? No more than 10 to the 27 flops in the next pre-train or something like that. But what actually happened was he basically said, "Yeah, I think that could be reasonable. Again, the details will be like incredibly important if we go down
5:42that path. Um because the more data you enrich, you know, the more flops you put in enriching the data, you know, you can kind of you can definitely play a sort of shell game of hide the compute. I think these days, what about Elon? What about Zuck? But the the take was kind of the two to three companies are really pushing the frontier and even the other two that we would usually think of in the top five, Elon and Zuck, are actually distilling a lot whether they
6:13And the way that they are distilling these days, it's not that they're, you know, calling the cloud API and getting traces and feeding them in, but rather the cottage industry of RL environment uh makers and sellers is in in fact functioning as a distillation conduit because what they're all doing is having Claude create these RL environments, selling the RL environments to other frontier frontier model makers, and then
6:42they do their internal RL on an environment that only exists because cloud was smart enough to make it. The the kind of synthesis of all this is if they hadn't been releasing the models and if people weren't able to do these like various sort of direct and indirect distillation techniques, they actually think that probably would have happened.
7:03uh that the other companies are not really keeping up but for the the sort of leak of intelligence in all these different ways including the RL industry um that is gradually finding ways to transfer from frontier uh models to their competitor anthropic and opening high are both structuring for more like the two to three year mark you know um and they're and they're looking at a 2028 2029 AGI RSI honestly even sooner I would say was
7:31kind of the vibe and this is probably somewhat biased to be a short timelines crowd, but again like the short timelines crowd includes the, you know, executives, top researchers, founders of these companies. They were talking more about just next year. Definitely not full agreement on how fast it goes or if or when it tops out or whatever, but a shared sense that like the decisions that get made over the coming months and and the way that compute is used over the next year could be really critical
8:00in terms of, you know, will they still have need for compute or or does a cap on pre-training scale or other, you know, compute input cap um that might constitute a pacing mechanism? Does that mean the end of the compute buildout? I would say no. Uh, another question I had the chance to ask was, of course, everybody's heard Jensen uh, on Ezra Klein, I think at this point. He said, "At NVIDIA, we spend like 20% of our effort on designing the chip and we
8:30spend 80% of the effort on verifying and validating and testing all the edge cases and making sure it's going to last a long time and be reliable and, you know, all these other things that are not just the original design.
8:42Congratulations, you've succeeded. now you're going to enter the era where you're going to have to make them safe, reliable, trustworthy, etc. And so I had a chance to ask, you know, one of these people like, "What do you think about that uh take from Jensen?" And the response was basically like, "Yeah, that's that's kind of reasonable." you know, without again knowing exactly where the numbers will shake out. U the possibility that the majority of could compute in the future
9:11could be going into all sorts of different safety measures, you know, which could be monitoring, could be chain of thought monitoring.
9:19People are getting kind of bearish on chain of thought monitoring, but then there's internal, you know, activation monitoring. They're already spending a significant percent of compute on monitoring and it sounds like that might go up a lot. Another thing that they're spending compute on in a big way right now is fixing the RL environment. A model was trained on a bunch of hackable RL environments where cheating was rewarded. It's a bit of a cheater.
9:45That's a big problem. They are now applying these models to the RL environments themselves and saying hack this environment. And now that is the direct task. They're doing it. They're finding, you know, all these places and and ways in which the environments can be hacked. They're fixing those. Um, and gradually they're driving down the rate of flaws, which means that they're also driving down the rate at which cheating would be rewarded, which in turn means you're going to get less cheating from the models as they actually come online.
10:15I think this is still like, as they say, an open research question, but it seems like they're probably going to be able to get comfortable releasing the next couple things even as they do consider themselves bottlenecked on alignment, but then past that they're like, "Yeah, all bets are off." There was definitely some interesting discussion around
10:35how hard will anthropic press their advantage to the degree that they have an advantage right now. And the sense that the AI safety community had when Anthropic was founded is that they were committed to not advancing the frontier.
10:50They wanted to be at the frontier so they could be relevant and do all the research, but they weren't going to push it because they didn't want to make the race condition worse. And people challenged some anthropic leadership on that. And their response was like, well, what we said was we would never publish methods that would advance the frontier.
11:04And people were like, I don't think that's really what you were saying at the time. And now they've been asked like, "Okay, you know, OpenAI's trying to get their house in order. You guys definitely have your, you know, work to do too, but you seem like you have maybe a little bit more in order than they do.
11:18Are you going to chill and like let them get their house in order and not try to like run away with the competition or not?" And we didn't really get an answer on that. So, a chip startup's AI bill now rivals its payroll. And this episode asks what that money buys inside a chip design team and when the cheaper model becomes the costlier choice.
11:43Thomas Sr's co-founded Posatron AI whose engineers are taping out their first custom silicon with AI agents working the tools.
11:52Our guest for today is Thomas Summers.
11:55Thomas previously founded Rex Computing, a processor startup and was a 2013 Theel fellow. He later became principal hardware architect at Lambda, which provides access to graphics processors for AI developers, and worked on its early cloud infrastructure. He then served as director of technology strategy at Grock, another AI chip company. Posetron was founded in April 2023 and shipped its first Atlas server in August 2024. Atlas uses programmable chips. The company's next generation
12:23Asimov is custom silicon intended for Titan servers. Positon reports that more than 50 Atlas racks have been deployed at Oracle. It's a pleasure to have you uh on the show.
12:34Thank you. Thanks for having me.
12:35One of the things that you've done is you've speeded up the um the process of designing and deploying the chips. Can you tell us a little bit about how you did that?
12:43I would say like a core founding tenant was that we wanted a real hardware in customers hands as quickly as possible.
12:49Was about 15 months from cold started company to uh uh shipping a first uh customer server. Um and then you know about uh 3 years uh total time uh to to win Oracle as a customer. But for our second gen which is full custom silicon uh it's really really amazing the advancements in in AI over the past year specifically in our ability to close uh the iteration loop in in terms of uh everything from RTL development. So all the front end work related to chip
13:17design doing the the verification which is the the biggest time limiter I would say um for historically for chip designs uh all the way to doing the full back AI has been able to to drastically uh accelerate our our own development of chips to make AI cheaper and and uh uh uh more energy efficient.
13:36How um how does this translate today to like token spend versus salary spend? So on on the token spend side um it's become you know our single largest uh non manufacturing line item very very quickly. Um bigger than bigger than human salaries.
13:54Yeah. Recently it eclipsed human salaries. It came back down um I'll explain in a moment. So 6 months ago was like equivalent of a single employee got um claude plus plus GPT and a little bit of cursor and other stuff all for the cost of a single employee.
14:10So our token spend has massively increased when in uh June it started to now be multiple employees and now it's more than like we were peing you know in the in the days after GPT success and we're really trying to to push its boundaries etc. we were we were spending over $100,000 a day um on tokens. Um that tempered down a little bit in the weeks after as we optimized a bit and and just weren't trying to run as many
14:38parallel experiments. Um but I would say the big thing that was like you know a saving grace was you know Opus 5.5 came out and in a lot of our tasks not everything it was doing better than than Astra and was you know a quarter the price. You mentioned how much AI is helping you on the verification side.
14:58Um, in a uh notable interview Jensen Wong recently gave with Ezra Klein, he talked about how Nvidia spends something like 20% of its time and energy designing and then 80% verifying, validating, you know, ensuring long lifespan, reliability, etc., etc. Um, what do your ratios look like on that?
15:21And you know what more color can you give us about like specifically where AI is really helping?
15:29So for us I would say it's a little bit different cuz we're starting from scratch. It'll end up being somewhere around maybe 60% you know uh design focus for us to 40% verification. Uh, I would say like the probably most amazing thing that I I don't think would have been possible 3 or 6 months ago and and only really became possible with GPT6 Astra and and open now Opus 5.5 that that we're we're leveraging a lot um in
15:57in this task is um we we uh for for doing our verification work we use Cadence Palladium um emulator systems. So these are big giant racks that are full of custom AS6 that are specifically designed to u just gate level emulation of of the silicon agents themselves have been able to completely do closed loop uh iteration and testing.
16:20you know, I I think people thought it was crazy, but uh you know, handing uh you know, these very powerful agents, you know, full keys to our our internal infrastructure and saying, "Here's Palladium piece." And I highly highly doubt that these models had um uh you know, Palladium documentation in in their uh uh pre-training thing, especially because a lot of documents are new as as software updates. But um the fact that it went and looked up the documentation when we're looking at
16:50these agent traces uh the reasoning traces and tool calls etc. it's reading the full documentation effectively compressing that itself by by you know reading the PDFs generating its own you know markdown files of of cheat sheets of what to do and then building out all of the testing infrastructure and its own harnesses for being able to access this. If you're just doing regular software RTL simulation, that's running on the order of like 10 10 hertz for us
17:17in in in those cases.
17:21If if you remove a lot of the debug pieces that slow things down, etc., we can run the full chip emulation on the order of 500 kHz. And right now that's really just focused on implementing test programs, finding cases where they thought fail and then um uh writing out reports which you know get reviewed by other agents and and by humans in the loop. But um I I would say when this really started
17:49working, you know, 6 weeks ago, we roughly, it was beginning of September when when uh Astra 6 came out, it it was a huge like step function improvement from, you know, GPT 5.6 where it it could not do that full closed loop. It still required humans at different stages. So one of the commentaries on um AI in general is that it's a lot of like process following work especially like
18:19for example taking a a manual for a particular piece of software or hardware and um implementing it or following the rules behind it and the criticism has been that it is not really true new innovation. Have you have you seen any uh signs of really some unexpected innovation rather than just rule following in that sense?
18:47I have not yet I I'm I'm somewhat proud of the fact that it has not figured out some of our cleverness of what we've done. it has like questioned and like this seems like a bad design decision until it gets explained to it why something's done or it actually runs test and sees and understands oh that's why you're not doing the traditional way a systolic array is done and so it it does make me feel a little proud even how just how unbelievably amazing the models are today that it won't come up
19:17with the I would say the ingenuity that we had as as human designers do I think that'll hold for another 6 months uh not not sure part why I actually think it's not going to be like the agent will come up with a better idea immediately. I think the amazing turn of events that we've just had in the past few weeks in terms of us having access to these models um is is the fact that just because it can iterate so fast and do experiments and things by itself,
19:46I wouldn't be surprised if if it would come to the same conclusions we did or maybe you know make a better things than what we would do by ourselves just because it's going to be able to iterate so fast and and go through so many different design possibilities. what kind of trade-offs does it involve or what kind of you know engineering challenges do you take on in order to get the advantage of being able to use commodity memory versus the you know the very high-end memory.
20:15My fundamental belief was that um there isn't going to be a stop anytime in the near future in terms of scaling laws and and getting increased value from from making models larger. So that eats into memory on one side. Um, but then the other side, I think the main limiter today of uh AI being applied to to most applications is context length and uh uh being able to hold more context uh uh you know per user and then scaling that out to drastically more users and and
20:44really today I would say that the driver of quote unquote users or individual sessions is actually just having more agents. I've gone from, you know, three, four months ago where I had probably on average two to four agents like running constantly in the in the background to where I'm now running 15 to 20. So, um, you know, if you then multiply how many people are using AI and how many uh that that concurrent completely separate contexts uh uh adds up very very quickly. That memory driver had us say,
21:13okay, we have to use the commodity memory because that's the only thing that's going to be able to scale and be cost effective. And when we set that as sort of the constraint in our architectural uh design um we had to come up with with uh you know very very clever innovative solutions and and you know the the two main pieces of that and you know the thing that we showed with our first generation product is being able to achieve extremely high memory bandwidth utilization from a uh uh compute architecture perspective. So
21:40while Nvidia GPUs, you know, I I would say on average are getting between 30 and 40% um memory bandwidth utilization in the actual decode forward pass of of a transformer model. So basically that means even though they advertise you know 8 terabytes per second of theoretical memory bandwidth with uh you know a B300 um you only actually are seeing you know something in the ballpark of like 2 terabytes per second of actual realized memory bandwidth down
22:09to the reuse patterns that really don't exist in transformers that they designed their hardware architecture around for for training and other workloads. Um and so with our firstg product, we were able to hit and sustain 93% of theoretical memory bandwidth. We're we're um you know partners with Credo um semiconductor that uh um and and have developed a uh memory chiplet solution that allows us to go from you know the
22:36maximum number of LPDDR same type of memory that is in phones, laptops etc. The max number of channels that you can you know have find in any other products is on the order of uh uh you know 12 12 to 16 channels and we go up to 72 channels of of uh uh LPDR5X uh with this uh uh decoupled uh memory chip solution.
23:03So um walk me through uh a model that fits locally like which which parts of the networking or coordination start to disappear with Ozimov the the our chip we have up to 2.3 terabytes of memory capacity per chip. So when you compare you know the the B300 shipping today caps out at 288 GB. So that means on a per chip basis you now can actually scale and what you would have needed from a memory perspective eight GPUs for we can do on
23:31a single device and it's not just oh you've got that cost silicon savings whenever you have to scale to having more than one device there's overheads involved with that there's uh you know your all gathers all reduces that you have to do uh within for each layer and each really each multiply uh that you're doing of of uh uh sharding across those devices comes down to real real performance in the There there have been increases in prices of all kinds of memory modules in the last year or so and pretty pretty significant increases.
23:59Has that affected your uh projections on the TCO uh going forward?
24:04I was looking at a quote that we had a year and a week ago and it's gone up four and a half times um since then. If if um if all all other things being the same TCO comparative to other um compute providers, everyone's in the same boat with having to increase their their cogs and and then thus prices. So that doesn't affect us that much from a competitive standpoint. I would say the the fear would be oh if if the cost of the products goes up and you're
24:34otherwise delivering the same value then then uh there's a problem. But uh uh but I think the the reality there is compared to a year ago, the capability of a model using x amount of gigabytes of memory is way more than 5x than than it was this time last year.
24:51So obviously the nature of the hardware that people can access determines in part what kind of models they can deploy. I'd love to get your kind of perspective and breakdown on how these hardware considerations relate to looping. maybe start with like why do we loop? You know, why is that even a tempting thing to do?
25:11It's it's been very interesting from the loop transformers paper to that being the rumor of, you know, one of the big advancements with the GPT6 Astra. Um that uh uh simply repeating the forward pass uh you know gets you that improvement. It's it's crazy that there was really work that I think was only being done in sort of the fringes of the open- source uh you know transformer community you know 2 3 years ago of taking like a llama 7B and just
25:40duplicating layers within and you got better results from it even though it was repeating the exact same uh map imagine what you can do when you train it to work that way. C can you go just a little bit further into the hardware connection of that? I'm kind of coming in with assumptions along the lines of like, you know, just very fundamental terms like why loop at all?
26:01I think it has to do with this kind of memory bandwidth, right? Because now I can keep the same weights on the chip and I don't have to like shuffle things in and out as much. No, not not really. Because basically the assumption is if you're doing inference that you're going to have all of your weights locally in in DM. you're just sharting that over some number of devices, you do get some amorization, but just with the size of experts and the size of of caches on chips, there's
26:30not much of a reuse opportunity or you know, by the time you're already done with the layer, you've already gone through the u amount of memory many times that of of what is in like the onchip SRAMs. uh from from like a hardware perspective like looping really is more memory capacity savings than memory bandwidth savings because like let's say that you trained two models with the same base set and you trained one to do with looping and so it only has 10 layers instead of 20 layers and
26:59and you know model model B is is 20 layers um that 20 layer model it's doing very very simple is going to be double the size of the the 10 layer version and if you find that the looped 10 layer version gets 95% of the same, you know, quality of results and it's half the size, you're probably going to deploy that one. You're you're saving more capacity than bandwidth because doing the second loop, you still have to do the same memory fetches and do all of the same like same number of maples.
27:28There's the exact same number of bytes that need to be moved and the exact same number of uh um flops that have to be done in that model A and model B scenario. when when you design a chip and there is a certain point in the design where you kind of freeze the design and you know hand it over to start uh the rest of the process. What are the decisions that you delay or defer until the very last moment before you decide to hand over?
27:58There there are a lot of features at different points of the chip development stages that we left on on the cutting room floor. Most of them are things that we still want to do in in future chips, but it's what can you actually accomplish, especially with our goal of iterating quickly. And it's better to get a product out quickly. Um even if it doesn't have all the features that you want, um to get that out there, get all the the learnings, product feedback, revenue, everything that you get from actually having shipped something. It's
28:25it's not just the the design time, but it's the verification of that component and you know figuring out how that integrates with the rest of the system. There are a bunch of optimizations that we could have made and done differently if if you know a year ago we thought that uh um linear attention was was going to take off the way it has at least in the open source model space.
28:47I'm still not convinced that that's being done at the scale that we see in you know the open source and specifically Chinese open source models in like the big model labs. Um but uh it makes sense in the constraints that uh the Chinese ecosystem has uh in terms of hardware supply etc that they're effectively being forced to go in in that direction and so they're innovating in in that way. As far as I'm aware,
29:15none of the major US model labs uh are doing anything like MLA or you know um uh DSA or or any of these other techniques that the the Chinese labs have done. Primary customer focus is on like the the big big model um companies. So I don't necessarily think we made the wrong choice there to to focus on the things that we did. Posatron closed an $875 million series C on September 10th.
29:49Money earmarked for taping out Azimov and ramping its Titan servers.
29:55As the token bill kept climbing past payroll, did anyone tell the engineers to pull back?
30:02We didn't tell anyone to cut their spending or do anything even though we were seeing this exponential growth in in our our token spend. Um it it was really just the fact that um uh we we both felt that we were getting good ROI on it. So we weren't going to to uh taper it. Um but also the fact that like we have very high confidence belief that uh uh costs are going to come down and yeah we're we're using um you know the
30:30the major model labs uh uh you know their their models for everything today.
30:34you know, we also have faith that we're going to be able to move things to, you know, local models running on our own hardware. I would say like we we do a little tiny bit of this with GLM 5.3 today. Um, but personally, like my philosophy is I really don't want anyone to be running like a real software or hardware development task on anything that is less than the best model. like I I don't care if it's, you know, a tenth of the
31:03cost, you know, per token. Um it's uh just not worth the the um the the expense.
31:12Shaun Wang Swix employs engineers and non-engineers, buys software to run a conference business, and decides who speaks at AI engineer. That puts him on every side of one question. When anyone can prompt the model, what is an employer still paying people and vendors for?
31:34Next guest is Swix or Shaun Wang. Uh he is the co-founder and CEO of AI Engineer, which runs conferences and workshops for people building software with AI. He's also co-founder and editor of Latin Space, a publication and podcast about AI and the people developing it. and he's also an adviser to cognition, the company behind the coding agent Devon. Um he also founded Small AI, whose AI news project combines automated reading and summarization with
32:04human selection. The next AI engineer New York conference is scheduled for October 12th through 14th with a focus on applications across banking, investing, insurance, and financial technology. Let's bring up Sean.
32:18Thanks for having me back on. It's so effortless to create so many things that I really am bottlenecked on my ideas.
32:26What is the vibe that you're getting from people who have been software engineers building software applications? Um are they still feeling like super empowered? Are they starting to feel threatened? For a while, people were like, "Oh, well, there's going to be, you know, so much demand for more software that like it'll be fine. You know, there will be even more software engineers."
32:44I think college kids are a bit worried. Uh but other than that, like if you're relatively plugged in and um very capable with AI engineering tooling, uh you're in more demand than you've ever been because your expected value is higher than it's ever been.
33:03Therefore, the demand has increased a lot for a very specific kind of demand for people who can manage coding agents uh productively instead of producing a whole bunch of slop. Um and I have been in that situation. So uh I am both an engineer but also an employer of uh engineers and people who are not engineers who are vibe coding. When you work for me and I'm paying your tokens, you better be like actually producing thoughtful stuff. And I have two or three employees right now in their performance review because they are just
33:30giving me cloud slop and that's really bad for them and they don't understand it. You're not like producing any value from Claude. I need you.
33:38Like I'm probably also mostly producing claude slop. What's the delta?
33:43Go for depth of insight rather than breath of coverage, you know, of like, you know, my my my YouTube guy will be like, "Oh, yeah, your your videos went up like 14.7% week on week." And I say, "Why?" And he doesn't know. He's actually never watched the videos. He just used Claude to like analyze the videos. And Claude doesn't have video analysis, contextual understanding, or anything like that.
34:07You're a human using these tools. You cannot let these tools think for you. You got to actually think to to use these tools. What did I do right? Can I do better? There's lots of areas for human domain expertise, but then they send they like substitute it for the like the lowest energy, lowest cost means of production, which is uh chuck something in the cloud and hope that I don't notice.
34:28He is a buyer as well as a boss. AI engineer stages conferences several times a year and the videos design and marketing come from outside agencies.
34:40When a vendor's work comes back, how quickly can you tell whether a human thought about it?
34:47At an agency actually, you know how these like launch videos come up, agency goes like, "Oh yeah, we made that." I like reached out to one of those agencies. I want to work with new requirements and they came back to me with like very clear cloud stuff. It's like really doesn't get what we're going for. And I think like to some extent cloud stuff is fine as long as it's like useful and good. And they they were just like, "Oh yeah, you caught us. Uh we're going to move on." And so like they're just hunting for for people with low standards or like who just don't care or who who have enough money whereas like I
35:15think the people with taste um just care about the product the end products and will just have like an im immediate reaction to like yeah clearly you have no relationship whatsoever with this core content uh and you're just doing stuff. So this whole taste question to what degree is it coupled or correlated with a background in software engineering though because I'm not reading any code. Most people I'm talking to aren't reading any code. But do you care about the part of the resume
35:43that's like oh I've been coding for X years at this point or is it not so relevant? I I don't I don't care so much uh except for let's call it uh uh security roles um and for um anything sort of backends uh scalability and I really need you to know the differences between GCP and AWS and like what I can do on each each one of those things. um uh you know like uh I just recorded an interview with the
36:11superb base founders who who you know are scaling Postgress to uh a scale that we've never seen before like yeah good luck trying to hire somebody to scale Postgress uh like uh you know with with a complete uh self-managed sharding solution uh that works at YouTube scale um there's only one person in the world that can do that and they hired the guy is about taste uh and it's not so much about length of experience actually sometimes experience works works against you because you just have a set way of
36:39doing things and you don't understand how to do more than one agent at a time. You should be relatively comfortable juggling like five to like 10 ongoing multiple things. Definitely true that you're much more of a manager than a individual contributor. Uh I do think that people who uh look at data rather than code are more valuable these days.
37:00Um so uh being able to say like here's the logs here's the traces here's the uh schema um here's the input output capturing it turning into an eval uh all these are basically the sort of the merging of AI engineering and ML engineering that is that is happening that people need to upskill on um and then like managing scaled up runs um uh of that every coding agent works inside a
37:27context window a bounded working memory what falls out of it, including what another agent did last week, is simply gone.
37:37When several of them build one product, where does the human's oversight have to sit?
37:44Like, right, like uh I think that the the person that has, let's say, 10, 20 years of experience in software engineering really reviews every line and tries to make every line make sense, whereas now we just need to make the the modules make sense. and I I I can allow slop in there because it helps me to go faster as long as I contain the slop in things where I completely understand the whole system. Uh where you go wrong is you have too many modules, too many
38:11black boxes where you don't even know and actually the code also gets confused. Making my own uh sort of Slack competitor, I saw a bug where like the the messages weren't load loading. I refresh and the messages were loading and I refreshed again and the messages were not loading and it's like well it's the same exact code. what the hell's going on? Turns out there's two paths and there's a race condition. Why?
38:31Because two different coding agents worked on it at different times and because they they just like made their own thing and agents would potentially sometimes do it because like sometimes just things fall out of the context window. But like you need to have the oversight of the module inside of the module can be a black box.
38:46You did uh a kill my SAS uh project.
38:50So so why don't you tell us a little bit about it? It came about because you had a product that you were using for AI engineer and you found that they were kind of over billing you or you you you were not happy with how much you were paying for the amount of you know work that they were doing. Basically uh you can call this like you know a bounty on uh mid-tier SAS that should not exist and this definitely was one of them. uh basically like you know half a salary or
39:18one third of a salary for a for a piece of software that I don't own that nobody enjoys using. Let's just throw that out. If I'm going to spend $40,000 on this SAS subscription, I can spend that in in amount of tokens and that buys me a heck of a lot of tokens.
39:33The software he wanted replaced runs the back office of a conference its organizers expect to fill with more than a thousand people and over a 100 speakers next week.
39:45How do you tell a replacement that works from one that merely looks finished?
39:51This is the problem. Uh especially with like launching it without planning that just did it on a vibe uh to my entire audience. Uh we have so many submissions that we then have to eval them. Um and so emails are the problem. So like we're no longer checking like here's like two or three requirements that you fit.
40:07We're checking the the entire UX of the entire flow from three different perspectives of like um you know organizer, attendee and and sponsor or a speaker and like uh then you need different login then you need like all the different workflows and like uh to submit different uh applications across uh the uh the different formats that we do and I actually shipped evals for to help my participants on on uh the second but um for really frontier things the
40:34human just has to play play test all of it uh on top of the regular data. Uh just be aware that like if you offer like a large bounty like we offer $10,000 as the prize. Um you'll get a lot of sub submissions because people want to vibe code for fun and a lot of them will be low quality. They will just be like Claude, make no mistakes, go do this. Make mistakes and they'll just submit.
40:55And so like the the sort of verification load is very imbalanced because they spent zero thought on this thing. They just threw it in there and then and then you have no idea if it's like low quality or not.
41:08Devon, the coding agent from Cognition, takes requests in plain language from a Slack channel. In a talk earlier this year, he described a designer on his team sending it annotated screenshots with no training.
41:23What finally moves staff who distrust anything vibe coded?
41:28But overall, very successful. So I have one of the most adversarial nonAI teams in AI. I hire event professionals who are like super old school. They do everything in spreadsheets. Like these are not a glamorous like high-tech jobs.
41:40Then they're very super suspicious of anything that is new and vioded and techy. Initially they were like we'll never use this vibe coded thing. Like this is like like I want to use the tried and tested stuff that like has has been uh working for for Microsoft and all all these these other guys. Um and then they saw the quality of the submissions and they they looked at their existing platform. They were like, "Yeah, okay. We're going to switch."
42:00This is one of the the the benefits of me partnering with Cognition on uh is I just all of them access to Devon to modify the code. So, anything they don't like, they can just request a change and they have it uh pretty much in one or two hours. Um which they've never had before. Just to give you an idea of the the the the way we work with these SAS companies right now. When we request a change, they'll say, "Okay, that sounds pretty cool. It's on our Q3 road map."
42:24And like we don't have confidence that it'll actually be done in Q3. We requested for a thing in Q1. It just landed. So like SAS is quite cooked if you're if you're mostly a CRUD app.
42:33So how can you synthesize this experience with the first statement you made around if you're decent with vibe coding never been in higher demand?
42:45There is a lot of latent demand for software that hasn't been met yet. But I have a hard time seeing how we don't kind of pick a lot of this lowhanging fruit and then end up in a sort of a lot of companies go out of business. Are we going to see mass layoffs from big tech?
43:02How long can the party go on?
43:03I mean the simple answer is like if you look at it as a fixed pile of software engineering then that is the conclusion that you will arrive at. Uh but uh if you look at it as the point of like my my competition or my TAM is spreadsheets. All spreadsheets made by anyone for with any sort of productivity tooling can be turned into custom software for them to put on rails and with a beautiful UX and with like um you know automations demand for custom
43:32software becomes a lot larger. uh and more beyond that um if it's all if it's phone calls, if it's emails that go back and forth um and and you know eventually physical meetings between uh people um all of that can be incrementally turned into more and more and more and more custom software and hardware um and and models by the way like why is Salesforce so damn big uh is because people can customize Salesforce but at some point
44:01they need to stop paying the 300k K a year baseline subscription and spend 30k building their own personal CRM and they're good. In fact, they're more than good. They're a happy year because they can now modify it to whatever they want.
44:14So, you should have an idea of like the sheer diversity of human needs and desires. We don't even have the same amount of customization in our software as we do our handbags. What the hell? We just need to get the infra there and need to scale infrared. Limit on all this is the the the sum total of humanity's total token bandwidth, right?
44:34Which which is like literally the amount of silicon we can produce. It's going to dictate pricing and availability and and rate limits and all these things.
44:40You've already said that you've had issues with um I would say poorly spent tokens, right? tokens which were um uh you know parts of your organization that perhaps uh didn't have the uh critical faculty of uh evaluating um the tokens that were being you know emitted with sufficient um you know uh thinking. So how does that work as you kind of start
45:08to scale organizations and you get less ability for a human manager to kind of evaluate whether their subordinates are spending tokens proper?
45:19I mean yes you need the ability but like um are you asking if I have a solution?
45:25Yeah I mean how how how do you think that would work?
45:28So so I think this is this is coming down to just management philosophy in general and I don't have any sort of core competitive advantage about that. So on one hand uh yes um you know I would say that uh the leading companies and the names that you all heard of have a practical hiring limit because they just don't have enough managers to manage those people that they're hiring.
45:48It's roughly let's call it um something to the order of like a 100 to like 400 people a year. But like if you ask them to onboard 4,000 they would they would struggle. at the large large like more than 10,000 person company level nobody cares about the individual humans inside of a team anyway and actually you just manage by numbers what are your three top goals per team and are you delivering those three top goals to me so that I can plug you into the rest of
46:17my org inside of that black box I don't care like um uh I'll hold you to some cost efficiency metric but hire the humans you want spend the agents that you want you actually don't need to know every single person in your like actually you just need to know your your team leads and your team leads need to know them but like I'm just going to fire the entire team if the team team lead is not performing like it's not very humane actually are very like this you just have business units and every every business has a GM which is basically a CEO they're plugged into an
46:46overall strategy um and I think it works well as we bench the score coding agents advertise tests a model on fixing real bug reports from open source repositories, one issue at a time. A whole product is never the unit of measurement. What is missing when a model can clone the software you pay for?
47:09I would say that uh there's a lot of UX issues left. You know, if you look at Sweetbench and you're like, "Oh, wow. We're like 90 on Sweetbench." You're looking at the wrong thing, my guy. Like um you know, if you haven't actually tried to completely vibe code a SAS that you use, um you don't understand how models are still very bad at all this.
47:28It sounds as though you had to set up a SAS in order to evaluate your kill my SAS submission.
47:32But yeah, then we just like completely evaluate against like a sideby-side playrough of each each of those things.
47:36So the beauty of like computer use being good now uh is that you can basically just do a point andclick and screen capture and then compare compare compare. We have a 200page document that that documents every flow and then just make sure that each each of that is represented in your in your clone. We asked them to also put in some human taste into like well what what would you change if like you could do things differently? So you are allowed to go off script and like why why would you not?
47:59Over the weekend we've had a uh a series of math discoveries. I think one which perhaps is worth maybe uh discussing is uh three sum. If you take three numbers and you sum them up can you get to a zero or not? And this is basically a search operation at the end of the day.
48:18And so 2 - 10 and + 8= 0 is a threesome solution. Can you get the number of operations that you do below the n squed number? For 20 30 years now, it's been no. And uh these guys that published yesterday or today, they managed to get it to n equals 1.99992 by finding a combination of numbers in the number space that you know uh you
48:44could get the um particular like search time to under n squ. Um and this is pretty significant because it breaks the barrier and um for the moment no real uh impact because these are extraordinary large numbers that these guys found. So it's not not a big deal. An anthropic employee gave a cryptography problem to um you know Claude and ask Claude to solve it. And that cryptography pro problem relied on this n squ being true.
49:19That the that the limit of how quickly you could do something was n squ. And claude while doing this cryptography problem which relied on this n squ being true broke broke it and came back and said well you asked me to you know solve the cryptography problem and I kind of just solved it by undermining the entire basis for that set of problems. The other the other interesting thing about this problem is that after Claude solved it, Anthropic approached two
49:47mathematicians who were um leaders in this field and asked them to publish uh paid them to publish, gave them the solution and then they kind of reformulated in a readable manner and then they published under their name. So they are listed as the first and second author. Claude is listed and the fact that Anthropic gave them the solution is listed in page 63.
50:17And uh this is being claimed as the way that things are going to be in the future that um we we we would have models solve these things and we'd hand the hand the solutions off to mathematicians to kind of verify and the mathematicians who verify or explain it would end up taking the credit. There's still a lot of unease in the um in the in in the uh math community right now on what is actually the right thing to do.
50:46The guidance I got from AIS was like it's not likely to be immediately super consequential but that barrier busting effect can at times be super important. You know the classic 4-minute mile right once it's demonstrated that somebody can do it all of a sudden lots of people start doing it. So, it will be really interesting to see if all of a sudden there's like a wave of kind of follow-on optimizations that bring the the number down substantially further. This also seems
51:14to be going on right now in the nano GPT speedrun department. Just headlinewise, it seems like we've gone from something like 70 seconds on this um this classic benchm the idea is to train a small model to get to a certain loss uh as fast as you can in terms of wall clock time. This result here took almost half the time off which was more than I think the last 45
51:43improvements combined. And this one has been accepted. One thing I did think was quite interesting about this particular one that has been accepted and kind of seems to have kicked off this wave of uh you know significant drops in in the time was that the author here said AI didn't play a big part in it. This company is doing also like largecale training. So they did also make a point in their publication about this that the
52:13optimizer that accounted for a significant but I think not majority of the time savings that they did share is actually not as good as the one that they are using internally. Uh so they have an even better optimizer improvement inhouse but they're keeping that to themselves. This company is prestition. I think the nano GPT's uh record was a form of you know performance marketing intended to show
52:40the capabilities of their optimizer and their and their team in general.
52:44Here's a question for you. This this question actually I uh heard from aa Katra in May. So in May she asked me and others in an audience um how many copies of yourself would you need to have to do as much stuff as you are doing today with AI assistance. At that time I said
53:132x two clones or two or twice the speed.
53:18And this weekend, thinking about it again, I was like, I think it might be more like four or five now. Software in particular, it's like way higher. Uh, so you know what I did last night with like 12 prompts would be easily weeks worth of work. But even subtracting out the software development and just imagining sort of, you know, the content, the sense making, writing, you know, whatever, giving talks. Um, I still
53:47think I'm probably right now in the kind of four to five range. I haven't measured it, but what's your intuition for you?
53:55Um, I don't know, like maybe 10.
54:01So, so, so the the way I the way I look at it is that a lot of what we build for ourselves tends to be uh a couple a couple of different things that we tend to build. We tend to build like sensory organs. For example, Swix uh created AI news and AI news. He set up pipelines of um you know unstructured data from discord channels on AI from archive from all of these like different places and they kind of got filtered and you know
54:29ranked etc etc until he came to uh be able to send out these emails which were you know very uh compressed kind of you know sensing or and those became kind of sensory organs for the entire community to kind of know what's going on in the community. If you were the kind of person to wake up in the morning and look through like Y cominator hacker news and like four or five or six different pages to kind of keep up on AI that used to take like you know 45 minutes of your time in the morning and all of a sudden that's compressed by AI
54:57news to like you know 5 minutes that's a permanent compression. So when you say like let's take away the AI agents you also have to say like let's take away all of these affordances that we've built for ourselves over the last year or so. And once you start taking away those affordances, then you need to replace that with like an actual person kind of. So I think the numbers are actually much higher than we think. It's not just the number of tokens that you spend on a daily basis, but the fact
55:26that you've spent those tokens and that becomes kind of a capex and you have this capital asset that you've built for yourself. One other big takeaway from the weekend at the curve in terms of expectations and this was definitely shared again across your frontier company AI for science is very much where they think we are headed in the notistant future. One sentiment was are AI is going to be superhuman at just some things like the verifiable tasks or are they going to be superhuman at
55:55everything? The sort of middle position which I think is very credible is we are seeing superhuman performance at anything we care about and are you know really committed to investing in. Um, so that doesn't mean it, you know, it's full generalization to every domain, but even in domains that are thought of as like not super inherently verifiable, they feel like when they put their
56:25focus on it and they, you know, go license what data they may need to license and they, you know, of course, apply lots of processing to that data to augment and, you know, create synthetic versions and what have you. um and you know put some RL on it too. The whole package basically is I think broadly understood to work on essentially any problem that they really choose to focus
56:50on. And so AI for science is coming next and the expectation is we will see one person went as far as to say in no more than a year it will not be possible to do the very best science without AI playing a big role in it. And you know, we'll we'll see.
57:1730% to 50% of all science papers are probably, you know, non-replicable and are probably trash. And that's what we're training the AIS on.
57:25There's no doubt that you're right that there's a lot of junk science out there.
57:29And if these AIs are going to get that good at science, which I think seems very likely, there's no no doubt that they're going to be able to be able to separate the good from the bad with presumably imperfect reliability. I would bet it will be accurate enough that people will just kind of generally believe it. At a minimum, we are going to need to be graceful about that process. Um because who knows, you know, why all these different junk things happen. It's not fraud. There's all kinds of ways people
57:57can go wrong in the in the scientific process that's not fraud. Watch out for mega retraction would be my expectation. Can't be that much more than a year.
58:06there is going to be an enormous kind of social push back on um in the research community on this this kind of thing because it's going to affect a lot of careers and livelihoods. I think one other thing that I heard at the curve this weekend was on the topic of medicine in particular, but this is sort of a not not confidential thing, but just somebody who works on the uh AI for medical advice domain said doctors have been
58:34surprisingly warm and receptive to AI. Uh much less threatened by it than you might have thought. much more inclined to embrace it. Um, even on the level of like the AMA, uh, which I had long assumed would come in and try to shut things down that, you know, we're starting to encroach on its turf. Um, the point of view that I got from somebody who, you know, has been
59:03doing some of that work with with those kinds of um, organizations was like, not really. I thought that was really exciting and I would hope to see the same thing from scientists. It's a little more personal for the scientist. One thing, the AIS can come in and just be super useful in general in medicine.
59:17And now it's like, okay, this is a whole new reality, but they didn't attack you specifically as as a doctor, you know, who is uh who did wrong. The mega retraction watch is going to be something that's going to touch, you know, a lot of individual names and that is u inherently going to be just a lot more contentious. I wondered to what extent uh it's because the expectations that people have for doctors to be current on um all of this new medical
59:44technology is so high that they've been feeling like just overwhelmed you know for the last few years um you know post Google basically and just not able to keep up with the questions that are being asked by by patients who go in Google and and now they have something that that works for them, right? That that that works to keep them up to speed uh very quickly and accurately, too.
1:00:13Thomas is CTO and co-founder of Posatron AI, where engineers are taping out custom silicon with agents working the tools. Swix Sean Wang curates the AI engineer conference. His work is at a.engineer.
1:00:31That's the episode.