0:00Last week was pretty crazy for new AI models. We got two new heavy hitters that are truly changing the game. Gemini 38 Flash and Muse Spark 1.3. How can you ever decide between those two? They're both so incredible. Obviously, I'm joking. We're actually here to talk about Fable 5.1 versus GPD6 Astra, the two actual models that matter. And I'm sure my attention guy already wants to kill me for doing that joke and confusing people. So, sorry. Well, let me have some fun. Okay, everyone's going to watch this video anyways cuz they all want to know what model to choose. And
0:29these two models are unbelievable. I went in with pretty high hopes on Fable 5.1, but didn't expect too much and I still managed to be absolutely blown away. It's an incredible model and I use it every day. GBD6 Astra is the future.
0:45It is such a crazy leap in so many different ways. I love using it. I'm doing thousands of dollars of inference a day with it. These models are both incredible. If you want to just skip to the end to see which model's better, I guess you could do that, but you'd miss all of the different ways we can compare them because these models are very different in their capabilities and the things you can do with them. And there's a lot to be learned from the various skills these models do and don't have.
1:10I'm doing nearly a billion tokens a day across both of these models consistently. And I've seen their strengths and weaknesses in all sorts of different places, many of which have surprised me a ton. So, if you want to understand what these are good at, what they are bad at, how efficient they are with token usage, how hard the subscription limits will hit you, all of those types of things and more, and of course, the most important question, which sub should you buy today? I'll do my best to answer all of that after a real quick break for today's sponsor.
1:37Writing code has never been easier, but as more agents spam more slop at projects, it's ever been harder to maintain a cohesive, coherent project.
1:45Your actual repos are slowly falling apart as more and more agents spam code at them that isn't really verifiable with things as simple as lint rules and static checks. How do you know that the code coming in is actually following your architecture or doesn't have some small annoying quirk that a bunch of other code did in the past or that it's following your UI system and guidance the way that you want it to. You could add all of these things into your agent MD, but that only guarantees that your agent knows about them, not that it's actually going to follow. How do you make sure at code review time that all
2:13of these things are done properly?
2:15Wouldn't it be nice if you could control the agents reviewing your code and maybe spawn specific ones that will only approve if certain conditions are met?
2:23Wouldn't it be nice if you could just throw a markdown file in your repo that describes a specific behavior or pattern that you want to make sure exists in all code coming in and won't approve it if it doesn't match? Maybe you even choose which paths in the codebase need this review. Well, spoiler, this is a real file in the T3 Code codebase, and it's powered by today's sponsor, Macroscope.
2:42Not only do they have some of the fastest and most reliable code reviews of any of the review agents available, they also have an approval system that actually makes sense for projects that are moving as fast as something like T3 Code. We're about to break a thousand open PRs, which is just unfathomable.
2:57I'd be a lot more stressed about that number if we didn't have these custom agents checking through Macroscope whenever a PR comes in. The team was added one a few days ago because a lot of my slop was breaking the UI by making things like the tool tip patterns inconsistent across different elements.
3:11So now we have a UI consistency agent that will actually check during code review if any of the files for UI are touched to make sure we're being consistent with how we actually style things in our app. Now all pull requests get checked by the effect service conventions and the UI consistency runtime agents that we've introduced.
3:27And as you see, they're hilariously fast, taking 12 seconds and 21 seconds, respectively, roughly the same length as our also very fast CI. Once all the agents come back with their thoughts, the top level macroscope agent will give its approval rating if it thinks the PR is good to merge or not. And I'm going to be real with you guys, it's incredibly rare I bother merging PRs that haven't been approved by Macroscope. The slop is coming and it's trying to destroy your codebase. The best way to stop it is atvive.link/microscope.
3:53So, we have all the different ways we want to compare these models. Let's start with what I'm sure everyone is here for, science, because this is a science channel, right? Okay. Seriously though, this one of my favorite things to cover, not just cuz I think the 3D stuff is really cool, but there were some funny things in this launch. I'm going to start with the official Fable 5.1 launch notes because they were very excited to show off their progress in terminal bench science 0.1, a very difficult realworld science bench using
4:21tools to solve real world problems. And they had a huge jump here both in the efficiency for cost as well as the success rate overall where their best case previously was on high with Fable 5. It cost $34 and it got a 25%. Now 5.1 on XH high is even cheaper and got a 50%. Huge improvement. I see why they put this here. These numbers look awesome. At least they looked awesome until you went to the GBT6 Astra launch
4:48because Astra on low gets a 54.3 and only costs $11, putting low Astra higher than Max Fable 5.1 at under a third the price for the first benchmark that Anthropic had listed. What I'm trying to say here is if your job is science, you should probably go get a codec subscription. You can do some pretty crazy things with it, like mindblowing ones. Both of these models have been massive leaps in science. My assumption
5:16is there's either some new training data or some new RL tooling that has been given out or sold to the different labs and if it is something useful in post- trainining then open AAI got it and is using it heavily there and since OpenAI are the goats of post training they were able to apply that better. On the note of overlapping data, 3D rendering, both Fable 5.1 and GPD6 Astra made massive leaps in their 3D rendering capabilities. Astr was way bigger,
5:44though, like hilariously bigger. Where better to start than Fish Slop? This is Fable 5.1's version of fish slop. It is a 3D game where you have a submarine and some fish that you can feed and a mini economy as well as a little bit of combat. And this was the most impressive version of fish slop at the time because this version of fish slop had real models that were surprisingly good. Like the fish actually kind of looks like a fish has eyes placed in the right place.
6:12That is a lot harder than it sounds.
6:14There are like coral and rocks and lights that work. And most importantly though, the movement is actually very nice. It feels good to move around in this version of the game. like flying around or not really flying. Floating around with the sub feels awesome. It controls great. It plays great. It's surprisingly decent overall, but there's definitely room to improve. I say this with confidence because of the version Astromeade.
6:40I'd say this looks slightly better, just a little.
6:48It's a generational gap here. This is nextG and the other version was not.
6:55What's the hit button to shoot? Okay, left click.
7:00There we go. Killed the guy. This does have its flaws. The movement doesn't feel anywhere near as good, and I had to like go back and forth to refine it a bit. The core loop isn't quite as well refined overall, and the performance was bad until I told to fix it. But it is [ __ ] stunning.
7:16Like it's actually decent looking to the point where I might have to drop slop from fish slop in the not too distant future. It is a massive leap. And if you think I'm just showing this and saying this now because I'm trying to glaze Astra, go watch my video on Kimmy K3 where I was for the first time ever genuinely really impressed with the 3D modeling capabilities of an LLM. This is a world of a difference. This is like multiple generational difference. So, while I definitely want to give this
7:44without any question to Astra, I actually think I have in my notes here a bunch of examples of crazy 3D stuff people did with the model. Here, Dra made a copy of the Amazon spheres that he was working in during his internship, all with Blender using Astra, which is just insane. or this demo home that Thomas built using Astra. All also with Blender where I can make a real environment with good furniture, a nice
8:13backdrop for the window and like it's good. It's like actually usable. You could use this to make real world mocks for real world 3D stuff. It's not a hypothetical anymore. This is in my opinion the equivalent of the jump from autocomplete agents to actually using your agent to complete tasks. And this all happened from Fable 5.1 to GBD6 Astra. The 3D capabilities in Astra are unbelievable. Which is why you might be confused when I move the arrow over here
8:41to gamedev for a sec. The reason I'm doing this is that as much as I am truly genuinely blown away with the 3D rendering capabilities of Astra. It is a like generational gap. It's the biggest gap of anything here between the models.
8:563D with Astra is comically better than 3D with Fable until you start interacting with it. And this is where Fable still just absolutely moss. Fable is so much better at interaction in general, at things like handling animations moving the right way when your cursor goes in a certain place or making the camera move the right amount when you move your mouse in a 3D world.
9:23For my silly 3D demos that I've been working on, I find that the best way to build them is to make the first prototype with Astra to get everything roughly looking and feeling how you wantish and then have Fable come in and clean it up and do all of the like detailoriented work. Astra, in my experience, does not handle the delicacy of that type of stuff anywhere near as well. It really does feel like a sledgehammer. Babel can apply things more gently in a way that makes things
9:51that work better. Back to non-code though, because there are layers to this one. Computer use. Realistically speaking, a lot of these categories are going to see a sway one way or the other where like certain things are better with Fable, certain things are better with Astra. The biggest gaps by far are the 3D rendering where Astra clears and computer use where Astra also clears. It is so much faster than 5.6 was for computer use. It's crazy. It just flies
10:20through things on your machine. Computer use of 56 soul was a big enough jump that I started to actually use it dayto-day for various different things.
10:27With GBD6 Astra, I've been using it so much that I got another Mac Mini just to let it run with computer use 247 to do all sorts of different tasks. It's so good at realworld computer use. It can do things much faster. understands them much better and the benchmarks don't in my opinion accurately measure the gap here. And the gap isn't just the model either. A lot of the improvements have been through the changes they've been making to codeex, especially on Mac OS to make the computer use spin up faster,
10:54get context more easily, move around your computer faster, and do things in the background better. 5.6 Soul would do tedious things I didn't feel like doing for me. Astra can often do tasks faster than I would have, which is very convenient because I'm currently down a hand. Astro computer use is so far ahead, it's kind of hilarious. Fable can do it decently when given the right tools, but it's just not even comparable to computer use. Funny enough, the weird misses that I'm used to OpenAI models having for code, it feels like Fable has
11:24those when it's doing computer use.
11:26Speaking of just getting it, let's talk about copyrightiting for a sec. Both of these models are massive leaps in the quality of pros that they spit out. If you have no Agent MD or Claude MD and you send the same prompt to Fable 5 and then Fable 5.1, the output from 5.1 is way more readable. If you do the same with 5.6 Soul and Astra, Astra is way more readable. Both of these models are huge jumps in how much less painful it
11:55is to read the text they put out. They almost have like unslop baked in finally, which is great. It's a meaningful improvement and it means I don't hate the outputs that I'm reading anywhere near as much. It's awesome.
12:04Which one's better at copy? I got the hot take that they both still suck and nothing has topped Kimmy K2. Not even K3, not even K2.5. The original K2 non-thinking model still writes the best of any model I've used. But it's also stupid as hell and very quick to be incredibly rude. It's not a model you should actually use for much. But 5.1 is a huge improvement here. Astra is notably better, though. I've had Astra come in and make suggestions for cleaning up copy from Fable. Astra is
12:33much better at recognizing the [ __ ] copy from Fable, but its proposals still just aren't as good as I would like. I still find myself rewriting most of the copy that these models create for my web pages. And while I do think Astra's copywriting in general, like pros and readability is slightly better than Fable, it does have one really, really, really [ __ ] annoying edge, which is that it loves to stuff all cap subtitles into everything it builds in UI. It just
13:02throws these subtitles everywhere and it is the worst. Possibly the silliest place to see this failure is the fish slop 2D version that I had it build. I would ask you to count the unnecessary subtitles on this page, but it would not be worth your time. Before I had it clean up a bit of it, there were 21 21 unnecessary subtitles. It's just spammed everywhere. A little tank, a lot of life. Coral Coast. Three little lives,
13:30all yours to look after. All systems go.
13:33A tiny world worth looking after. The first version, it also had an online badge at the bottom that said online and ready. I didn't even see that one. Chad was spamming captain, you're home. It's in It's so bad. Like, I I don't even care that Astra is better at copy because it's trying to show it off by being shitty at copy everywhere constantly. It's It's unacceptably bad.
13:57The more I stare at this, the more I hate this model. I I just, for what it's worth, straight up do not trust Astra with UI anything at this point. Well, we should move on from this ocean of possibility to other categories because while I do give Astra a slight lead in copyrightiting capability, the way it spams it pisses me off too much to frankly want to ever use it for copy.
14:19Last thing in the non-code section, then we get to what you guys are all actually here for. code agents and cost audio and video work. I'll be real with you guys.
14:27From my experience, both have disappointed me here. I am admittedly very picky, like annoyingly picky about these things. So, I should not be the only voice you hear when you listen to opinions on audio and video work with AI and AI agents. A lot of people have told me that they had Fable or Astraedit videos. Every time I watch the videos, they suck. I do have one exception here, though. This is a demo that Ben Davis did showing how he uses computer use
14:56with Astra, not to edit the video, but to get everything set up to start the edit for the video. You can think of this in like code terms back pre-AI.
15:08Imagine if AI was so good at like using VS Code and understanding your codebase and like what you needed that when you were about to start working on a thing, it could open up VS Code. It could open up all the files you specifically need to edit. It could open up your terminal in the right place, GitHub in the right place, and a dev environment showing exactly what you're about to work on.
15:28So, it sets up everything you need to start working. It doesn't do the work.
15:32It just prepares you to do the work.
15:33It's like getting the room ready and cleared with everything you need. Benis had a surprising amount of success getting Astra to do this type of work, to get it to actually set up his editor in Final Cut to be ready to go to edit video. If Ben says it's good for getting his editor prepared, I'll take his word for it. I have not had the pleasure of editing a video in a while. I do really miss it. I cannot wait till I can edit a video again once my hand is back. Soon TM, but this very promising, its
16:01understanding of computer use means it is more useful for real professional audio and video work. July 2020 when nobody knew who I was back in the day.
16:11Ready for a crazy throwback, guys? This is a GPT3 demo that an OpenAI employee made showing how you can have GPT3 make a functioning React app. Crazy. Describe your app. A button that says add $3 and a button that says withdraw $5, then show me my balance. So complex. Look at that. It made the React code and it works. It didn't even use hooks. It used
16:41class components. This was a huge deal.
16:44I know that my joke six years later doesn't hit quite the same. I think this might be able to pass an interview. But like at the time this seemed unbelievable, but also like if you're a real engineer, you look at this and you're just like, "Yeah, I could do this in my first week of class." The hilarious way this looks to us in terms of capability, we're like, "Yeah, that's cool. That's not actually useful for real world work." This is how I feel when people post videos that are edited with AI right now. If you let AI edit your content for you, if you let AI do
17:14your audio for you, if you let AI manage your AV pipeline for you in these ways, you are doing the equivalent of shipping an app using GPT3. We are that far behind right now. Do I think we'll have an astro moment for video editing tools?
17:29Perhaps we've made real improvements here, but it's so far from usable right now that I I'll be frank. I just don't get how people think that they can actually automate video editing with AI right now. Not even close. Not even kind of there yet. So yeah, I don't think either are good enough that it's even worth ranking here. But due to Codex's incredible computer use, I think Astra just gets a free win here. So this whole non-code section, Astra wins. The only
17:58parts that Astra wins massively though are the 3D rendering and the computer use. Although I do guess the science progress is pretty meaningful, too. So like these three Astromogs much harder in 3D rendering than the others, but it does meaningfully improve in all of them. So now we are done with this section front end. I'm sure you guys can guess where things landed here by now. I do want to make a few things clear before I go on the utter brutal roast session I'm about to do. GBD6 Astra has
18:26made massive improvements overall with front-end stuff. It is meaningfully better. It follows instructions better.
18:33It can make landing pages more effectively and less cringey. And it's so much smarter and better at understanding things in general that you can force a good design out of Astro with enough effort and iteration. For reference, here is GPT 5.6 with some basic designs on Witch AI by Dra. It's a nice little demo here. This is 5.6 Soul.
18:54Right now, this is fine. The site's funny if like a little bit laggy which I don't love obviously but it has all these subtitles notes that think beside you loved by 18,000 curious minds an unnecessary m dash built for remembering.
19:15It's not the worst, but it's also not the best. And some of these are just such boring Tailwind templatey stuff.
19:22Pretty cringe. This one's okay. This one's awful. And this one is another boring tailwind template. So that's 56 soul. We bump over to GBD6 Astra. We can see how much better or worse it is here.
19:35Still too many of these subtitles. A second brain. A lighter mind. Free to start. Yours to make your own. I hate these so much. But open my space. Your mind a little more organized. Six little pieces of your mind. A little space just for you. The word little appears on this page 15 times. All of your little things. Still looks better. It does like this little arrow thing. That's kind of cute.
20:03It's fine. It's meaningfully better.
20:07That had a weird font pop in, but those are easy to fix. These little animations are nice. Looks fine.
20:15This one I hate. Just like a personal hate, but I hate it. Especially with the subtitles. This one's actually kind of cute. I don't know why that got cut off.
20:22Like there's a lot of cut offs here that shouldn't be happening that are but where it's going for here I I see what it wanted to do and it's not the worst like decent starting point and this is very boring and old school. So yeah improvements it is slightly better to meaningfully better in all of these categories. I can turn on the claude code design skill from anthropic and it does get slightly better. It's just a little too quick to experiment with
20:52fonts and things, but like these are passable. These are starting points you could reasonably use for something. But I need to be realistic with you guys.
21:00Moving the windows, I don't block this one as much.
21:04Do you see how nice that looks? The animation of all these paths coming in.
21:09The structure of the page is great. None of those unnecessary subtitles.
21:15It's a much much better starting point for a one shot. It's really good at these types of animations and having like a distinct style. This one is so much better than previous models. Like if we switched just to Fable 5 or a similar design, gross. 5.1 same mock, same design.
21:33Actually genuinely compelling. Like if I had seen this on the internet, I would never have guessed this was AI generated, much less a one shot. This screams like custommade to me. It's really good. That said, I have had my problems with both with design, and I have found that both actually kind of pissed me off. I've been quite annoyed at how much both have pissed me off, especially recently as I've been trying to iterate more on some existing design work. Fable takes the cake here easy, significantly better, but both still
22:01frustrate me constantly, and good designs still require you sit there with a hammer and beat the [ __ ] out of the model. So, while there is a gap here, I think the gap between a two and a five out of 10 is much less notable than the gap between a five and a 9 out of 10. And when it comes to like 3D capability, I would say Astra is a nine and Fable's a five. But with front end, I would say neither of them are more than a five out of 10. But at least Fable can make something that looks decent without having to like actually beat the [ __ ] out of it. Astra needs a
22:30lot of help if you want a good design out of it. So, next full stack. I'm going to be so real here. If there is a difference in how well these models grasp your stack, I'm impressed because for me, they both get it. We are finally now at the point where both Frontier models from both Frontier Labs can look at a code base and the backended front end and clients and servers and how they relate and make good, reasonable decisions operating against it. That was not the case before. Fable was the first
22:59model that could do that. 5.6 Soul could act like it was doing that, but often just missed things. didn't get and fully like understand the end to end story.
23:09They both do now. I would say it feels like Fable has better intuition with it.
23:14Like it understands the consequences of edges by default slightly better, but Astra is more willing to just go at the problem forever and test every single edge and validate every single assumption and force itself to come to the right answer. And I've seen this type of thing happen so many times where I ask Fable and Astro to do the same task on a big open- source full stack project something like T3 Code. Fable looks at the code, comes to some conclusions, has a few concerns, and then addresses them through its design.
23:43Astra has all of the things after it finds them, decides there are three valid options, stress tests all of them, eventually finds the thing Fable already knew from the start, and then makes a similarish solution addressing the same things. So in the instances where Fable can find and understand things, it is a nicer experience because it's so quick to work and actually apply changes because it builds its understanding more effectively because it almost feels like
24:11it understands better. Astra really has that groundhog day feeling to it where every time it starts it has lost track of everything in the codebase and it really is starting again. But because of that it will find things other models miss. Fable's coming in with the assumption that it knows and can figure out everything with it super fancy smart brain. Astra is the smartest model that acts like it's dumb. Astra quadruple confirms everything it's trying to do, especially in X high mode. Which means
24:39if the bug is outside of what Fable can grasp just from reading code and it's something deeper or harder to find, Astra will find it. It'll just take four times longer and burn way more tokens.
24:51So for me, it almost feels like a tie.
24:54But this is also going to be the start of a theme that we'll be touching on throughout. The theme is that Fable's performance is generally more steady, where Astra is much spikier, where Astra has moments that leave me in awe and moments that make me question why I'm paying OpenAI$600 plus dollars a month.
25:10We'll come back to that theme in a bit.
25:12But on the prior theme, which is Astra's relentlessness, its desire to grab a problem and strangle it. Giant project rewrites. Astra has made meaningful strides here. I have thrown it at some crazy [ __ ] and I'm blown away with it.
25:29It is mostly done with the rewrite of TypeScript in Rust. It's porting TSGO to Rust and it's making real progress. That actually reminds me how much am I [ __ ] my usage as I do all of this right now. Not too bad considering that I have that just like running in a death loop right now. It paused the goal again. You have my permission to do whatever you need to do. I have you on full access for a reason. Keep going. I am very deep in the TypeScript Rust Rewrite. With 56 soul, I was able to get
25:56about 30% test accuracy going against the giant test suite that they use for testing the TypeScript code bases. I was able to get over 80% with Astra. And it didn't even take that long. It was able to do that in I don't know, I would say about 3 days roughly, not even. It was able to jump from that 30% range to 80 plus. But then it stalled out super super hard at 82.6%.
26:22And I still don't fully understand why.
26:25I've been trying to go back and forth with it to figure it out, but it got comically further than anything else I've thrown at a task like this. And I want to be clear, I'm not in the loop on this one. I just set a goal and gave it 40 sub agents and said, "Do as you please." Remember those 40 sub agents, though, because we'll be talking about that in the agent section. I will also admit to a bias here, which is that when I'm doing early access testing with the new models with OpenAI, they don't
26:53heavily restrict my account. It does not count towards my usage, which is at sometimes annoying because it means I don't have a good feel for how quickly the new model will burn your account, but on the other hand is so useful for actually testing the capabilities of the model and pushing it to its absolute limits. I did 130 billion tokens with this model during the early access. That means I can yolo it at these types of things and not have to worry too much about the bill because I'm not paying it. I've never had free unlimited access
27:20to Fable before. Which means while I have seen and know the capabilities that Astra has here and have even shipped some of them like the T3 code mobile rewrite in Swift, I even have a T3 code build here with GPUi that I almost forgot about where I had it rebuild all of T3 code with GPUi, the Rust UI framework, just to see if it could and to test the capabilities, performance, and stuff like that out of curiosity.
27:46And it did it. It made it work great.
27:48But here's where one of the bigger catches with Astra comes in, and I've seen this in pretty much every one of these big rewrite attempts I have done.
27:58One of my old benches I used to do a lot in my videos was having the new models rewrite my legacy ping.gg codebase, the one that is used for my video collab tool that I got into Y Combinator with.
28:09I gave it the original codebase, told it to write a plan, and then I had it and Astra review each other's plans separately out of curiosity. But once it made his plan, I told it to make a branch for the work and commit the plan to it. Do that in a work tree. I had to do the commit so I could access it in other places. Then I told it to build the whole thing. 1 hour and 20 minutes later, it did. Since I hit yet another obnoxious case with codeex where since I was in work mode instead of in codeex mode, I can't actually open the
28:38terminal. I have to wait for this to do the things I want. So while I wait for that, I'm going to show you what ping normally looks like. This is the app that I use when I bring collaborators on for streams and all sorts of other things. So, if I hop in here, you will see me. Hi. You might recognize him.
28:56He's in all my videos. And the point of paying is to make it easy to do a call like this and bring guests in and most importantly be able to embed them in a program like OBS. I'll mute so I don't double audio. This is a direct link to just me in my call in ping that you can embed in something like OBS to put someone as a layer in your video production software. This is still used by a ton of the biggest Twitch streamers when they do collabs. And it needs an
29:24update pretty badly. So I asked the model to do that update to move it from the beta early version of T3 Stack and try to get it onto something a little more modern. The original page was customuilt by our designer and original co-founder Brin, who put a lot of time into it and made a genuinely awesome looking page. And this is what it replaced it with.
29:48A bunch of tech slop.
29:50But once you go in, it gets way worse.
29:52The dashboard looks like this. No real info, impossible to know what's going on. this giant pile of announcements at the bottom versus this page that actually shows you what's going on.
30:06And then once you try to join, the layout shifts a ton, which is a really big deal for content creators because it breaks a ton of [ __ ] for them. And when I want to actually select the device, I have to unfold device settings and manually pick them from here and then hit start preview. And now I can see this is atrocious. This is layers upon layers of sins. But do you know what the worst part is? Because the worst part isn't even in the app. It's here. It's
30:35in the prompting. Note this particular sentence in my prompt. Port features over one at a time. Reuse as much UI code as possible. Reuse as much UI code as possible. Do you see any UI code shared between these two versions of this app? Do you think there's even a line of tailwind in common between
31:01these? There is nothing. I cannot fathom. I just assumed I pasted the wrong prompt when I did this. I couldn't fathom that the model would so egregiously ignore that detail in my prompt. especially the model that supposedly follows your prompt obsessively. And it just straight up threw away all the great UI work that we had already done and paid a lot of money for in favor of a bunch of text slop in
31:28a shitty black and whitened app. So bad.
31:32So [ __ ] bad. And to those saying, "Oh, it got lost in compaction or something." No, it didn't. I had it on ultra with one mil token context windows. It just does this. And I've had it do this since even on the official non-preview version for those saying, "Oh, they changed this in the snapshot."
31:49No, I have randomly had this model ignore my request to rebuild or update something while maintaining existing UI and it would just blanket over destroy the existing UI in the process. And I've never had a model do this ever, much less this aggressively. It's almost like in their attempts to force the model to be better at UI, they accidentally gave it the side effect capability where it might just come in with a sledgehammer and destroy all your good UI and ship something atrocious instead. This is one
32:17of the most serious examples of that.
32:19And we'll talk more about understanding intent later. Don't you worry. But for this reason alone, I would make a bigger gap in the rewriting capabilities that we were just discussing. Because if the thing that you're doing this bulk rewrite for has UI, you cannot trust that UI is going to be carried over. I had even told it with the T3 code GPUi version to make it look and act as much like the existing T3 code as possible.
32:42And it didn't even use the sidebar properly. It made the sidebar the old school one because it just wanted to make something work. I'm sure if I added a bunch of stuff at the bottom of my prompt that was like to be explicitly super clear, I expect the UI to be pixel perfect identical. If any of the UI has changed, you have failed at your goal, so keep going until the UI looks identical. Then maybe it might have done a better job here. But I shouldn't have to add a paragraph to the end of every prompt to this model to get it to do what I asked it to do, which is reuse
33:11the UI code. You could say that I'm being too picky, that I am expecting my incredible god in my computer that is charging 200 bucks a month to do exactly what I intend every time, or you could just use Fable, which doesn't have any of these [ __ ] problems. Anyways, this is one of the parts that I think is most important. Code mergeability. This one is tough. And it's not tough because the decision is hard. The decision is easy.
33:34It's Fable 5.1. Fable 5.1 writes way more mergeable code. It just does. I've done the numbers here. Fable 5.1 from PR filed to merged averages two additional follow-ups from when the filed PR exists to when the merge happens. Meanwhile, Aster's been averaging six for me. What this is measuring is how many of my AI bots are catching mistakes. How much of the verification loops are finding things and then fixing them after. But to be frank, what it's measuring is from
34:04when the PR is filed to when it is merged, how much [ __ ] has to happen. And I found that with Fable, the PRs are filed in a pretty much ready to go state. And its ability to make the necessary changes without accidentally scope creeping and then actually give you something shippable is just it's higher. It just is. I did say I was struggling with this decision though.
34:24Why am I struggling so much with it? I'm not struggling to pick between the two.
34:27That's easy. Fable is better at this.
34:29Fable PRs are much easier for me to hit merge on. Generally speaking, the reason I'm struggling with this is because when I say this, it is perceived by everybody, including my friends at OpenAI, as me saying that Astra is constantly throwing up on mergeable slop. Notice how I didn't say that. If I was to do yet another arbitrary ranking here, let's make a beautiful, super, super accurate chart that perfectly describes real numbers. This is the most vibe based chart you're going to see in a while. I promise. If we were to rate
34:59the mergeability of code for Soul, Fable 5, Astra, and Fable 5.1, we would have one of the world's greatest benchmarks because it is really hard to measure mergeability. But if I was being realistic here with how I felt and how I still feel, low is bad. high is good.
35:15Obviously, Iota said 5.6 soul was like in the two to three out of 10 range for mergeability by default. You can do things to improve it and if you give it the right review bots to give feedback and iterate in the tooling it needs to check its changes, it could make working code incredibly well. If we were just measuring the ability for Soul and Fable 5 to get working code, they were neck andneck. But code that I'm willing to hit the merge button on in my real world projects that shipped to hundreds of thousands of users, Soul was cleared by
35:42Fable 5. Fable 5 shipped significantly more mergeable code, like three to four times more mergeable. And by that, what I mean is three to four times fewer things you have to fix after the PR is up. So, where are things now with Astra?
35:58This is why I've been struggling. I think Astra is roughly at, if not slightly ahead of where Fable 5 was. So, we were comparing Fable 5 to Astra.
36:08Astra's winning now, but we're not comparing Fable 5 to Astra. I was during my testing, but we're comparing Fable 5.1 to Astra now because we live in the real world where these are the models we have access to and 5.1 is like right on the edge of 10 out of 10 for mergeability. It is rare that Fable files a PR, especially for like known quantity changes like bug fixes or feature ads or performance improvements or all these types of things. The code
36:34Fable 5.1 puts up is just better. It is.
36:39and even Benjamin Ben Davis, friend and manager of the channel, co-host of the podcast, OpenAI's number one defender, who showed up in the GBD6 Astro Launch videos. By the way, we'll have a fun reveal in our next podcast episode, which by the way, if you're not subscribed, Nerd Snipe on YouTube and pretty much every other platform. It's where we get to chat more. And if you want to hear somebody disagree with me instead of me just yapping constantly, that's the place to do it. So Ben, whose job is literally to disagree with me and also the lover of OpenAI and the
37:08Defender of Astra, has admitted that he thinks Fable is way better than he expected for real world code and he finds himself using it much more than expected. It has, I think, three accounts with cloud code now. It just ships more mergeable code. It is what it is. But again, the thing I'm trying to show here with this diagram is I'm not saying what I said before with Fable 5 versus 5.6 Soul. This gap was comical.
37:35This gap was big enough that I would really only trust soul for exploratory work or things that could be verified programmatically where the code didn't matter that much in terms of its quality and maintainability. Fable 5 was actually shipping things that were mergeable. So if you perceived this gap before where you found 5.6 Soul is not good enough to merge code from, but Fable 5 often cleared your bar here, you'll be totally fine with Astra. Astra is unbelievable. I was totally fine with
38:03Fable 5 and in a lot of ways I still probably would be minus the slop of its outputs. And if you were to think of this purely in terms of the mergeability of code for fixed focused changes, I would say that Astra is a notable upgrade from Fable 5 because it is slightly more mergeable with its outputs, but also when you're using it, the output that it gives you is much more readable and it's more thorough. So it will verify the changes. So it won't make like like Fable's code is beautiful and mergeable, but it sometimes misses
38:32details, especially Fable 5. Asteris code can be elegant and it can be the right subset of changes to make the thing happen. It does still have the habit of letting scope creep just eat it up and destroy the scope of what it's trying to do pretty often even, but it talks so much better than Fable 5 did. I don't hate it the same way I hated Fable 5 for just reading its outputs. I don't feel like I have to go to Helen back to unslop it just to make the output usable. So from Fable 5 to Astra obvious
39:02win easy. Astra versus Fable 5.1 is where things get more complex. But the thing I wanted to emphasize here, the reason I drew this diagram is that if you are thinking the gap between Astra and Fable is as big as between Soul and Fable, you're just wrong. This gap was massive. This was like a 3x difference that made it hard for me to justify using soul if the code would ever hit users. GB6 Astra is much much closer to
39:30where Fable 5.1 is, but there is still a gap. So, I do still find myself defaulting to Fable for things like a quick bug fix or UI changes that matter or feature improvements or my favorite thing to use Fable 5.1 for to take something that Astra has in a death loop that it just cannot get through. Stop it. Switch over to Fable 5.1, hand it the PR and say, "Hey, make this actually land. Clean it up. Throw away whatever doesn't belong. Make a new branch if that's easier. Make the changes we care
40:00about here land." Bale 5.1 lands code better, but Astra lands it well enough that you're totally fine with it. So, there you go. Code mergeability. Babel still wins, but it's not as big of a gap as it was before at all. The speed of catch up from OpenAI is genuinely impressive, and I'm very excited for Astra 6.1, which will hopefully make these things even better. Speaking of the problems, you heard me mention this earlier, managing scope creep. I want to be clear, Fable can still fail here. If it gets the wrong review comments at the
40:28wrong time, Fable can bloat things pretty badly. But god damn, Aster does it by default. If you are very explicit with Astra to make the smallest possible changes, it will try to, but it will quickly have its context get bloated a bit by its reminders from all these review agents that will send it off course and then you end up with a thousand line of code PR that should have been 50. If you're in the loop enough, you can work around this. I'm not trying to say that there is some
40:56unique unbelievable thing Fable does here that Astra doesn't, but Fable takes less effort to prevent these things with than Astra does. I kind of want to take this phrasing and skip right over to understanding intent, but I'll wrap up code super quick with gamedev here.
41:13Astra is so good at 3D that it almost gets the win by default. But as I said earlier, Fable is so much better at the like edges for interaction, like the animation curves and the speed that your character moves and that your mouse affects the camera, all these things.
41:29Astra makes games that look better on Twitter. Fable makes games that actually feel nice to play. So, personally, I think Fable catches up to Astra's unbelievable 3D capabilities just due to the smoothness of its outputs. only for geo in chat just said none of them can make a good game and I agree they cannot but if you use both of them carefully enough you can combine them in a way that allows you to iterate effectively and potentially make a decent game I do
41:58think we are now at the point where we're going to start seeing real games where the vast majority of everything was created via AI like I would guess by the end of the year we'll have our first top 20 indie game on Steam that was built entirely with AI so now we're out of the tradition additional code section. We're done talking about code.
42:17Don't worry. Talking about the agency stuff. I imagine I wanted to skip to the understanding intent thing. So, I will think this is really important and it is still one of the things that anthropic just clears on. I have so many examples of this that I could do a whole video and I am honestly tempted to do a video on just how bad Astra is at this sometimes. It is an improvement over Soul, but it's failure. somehow feel
42:44more egregious. I mentioned earlier that I was able to merge like 150 PRs fully yolo merged with these models and had only two regressions. The first thing I feel obligated to say is that both of those regressions came from Astra. The other thing I feel obligated to say is that Astra was [ __ ] miserable to try and fix those things with. I will also admit I sent these prompts pretty late in the evening drinking with some friends on a weekend because I was making some changes and noticed that on
43:13the marketing site for T3 Code, it had nuked my beloved section for all of our testimonials from our users. The nice little elegant auto scroll here, it's not a big deal, but it is a thing I care about. And it replaced this with a fixed grid that you had to horizontally scroll with an ugly ass scroll bar in the middle of the page. Horrible. And it
43:41did this in pursuit of performance improvements that were simple changes.
43:45It removed the animation and turned this into a shitty manual scroll section. And I noticed this too late because we weren't deploying the marketing site actively. We just did it when we made certain changes manually. And I made other changes and deployed. And then this regressed and I was pissed. I was really pissed. So I asked, "What happened to my beloved autoscroll on the marketing site? It was beautiful. I need you to revert whatever change broke that and bring it back." It might have been part of the marketing something. I don't even know what I said there. It might
44:15have been part of the performance overhauls, but that was not a necessary change. Please revert. There's one particular word I used in this prompt, and I actually used that word twice. You know what word it was? I'll give you a hint. It was the specific thing I wanted the model to do.
44:33It's the word revert.
44:36It's a pretty important word in this prompt. I would argue that the word revert being in this prompt twice would imply that what I want it to do is revert something, which is exactly what it didn't do. Do you know what it did instead? It deleted this user controlled variable and then a bunch of listeners.
44:59Note what none of this does. None of this brings back the autoscroll. I don't think this change actually did anything at all. But what's even worse is I asked if I could test the change. And I was doing this change on another computer because I wanted it to have computer use and not affect my laptop while I was using it. So I needed this to be hosted via Tailscale because I had this other computer on on tail scale. So it spun it up as a tail scale dev server because I told it to spin up a tail dev server so
45:28I can try out the changes as well when you get a chance. I sent that as a steering thing which we'll talk about steering in a bit. Actually steering is actually one of the cooler differences in these models. So I'll sneak steering in here thing I want to cover in a minute. Before that, back to this thread, I asked it when it made the changes to file a PR, but also to spin up a tail scale dev server so I could test it. And what it gave me was a dev
45:54server to T3 code itself, not to the marketing site where these changes were happening. It gave me a dev server to check how T3 code was doing when I had made changes to the marketing site. And here is where things really fall apart.
46:09This is one of the worst runs I've had with a model in a long ass time. No, I wanted it for the marketing site, not for the actual T3 Code instance. Come on. You're right. I started T3 Code app launcher, which is the wrong target. I'm stopping it now and will expose the marketing Astro site itself, which it did. And I clicked and it failed.
46:30Blocked request. The host BB1 corpus micros is my tail scale address is not allowed because it's not a server allowed host.
46:37There are a lot of ways to fix this. The easiest admittedly is to make a change to a config file. But since BB1 corpus micro.ts.net is something specific to me and my setup and a lot of people use tail scale over IP addresses, probably wouldn't want to add this as an allowed host. Which is why it's particularly funny that the next commit was allow Tailnet host and dev server. Yes, it committed and pushed this
47:05and then it did the restore continuous endorsement rows which I thought would be the fix, but it also after that fix killed the dev server. I left tail scale pointing at Astro's old port and sent the link without checking. My mistake.
47:19And then finally it was working, but it hadn't actually made any of the changes that I request. It didn't find original PR or do the reverts. It didn't even fix the bugs I was reporting. Do you know what it did there? I'm trying to find it. But what it did is it merged the PR.
47:37It merged it for me after I told it you didn't make the changes. I sent the exact same prompt to Fable right after because I was so annoyed. This is an exact copy paste. The only difference being I added the third prompt in the spin up a tail scale dev server. I just copy pasted and added it to the prompt.
47:58Five minutes later, the auto scroll is back. The marquee is live with a link that worked for anyone on your tail.
48:05What broke it? Exact PR that had to be reverted. It replaced the two counters scrolling marquee rows with a static horizontal grid. This PR then bolted an 8-second page by one viewport timer on the grid. Neither was needed for the Perf goal since the marquee is a single GPU composited transform. It is worth noting that it did get tripped up a little. This shouldn't have taken 5 minutes. It got tripped up because it noticed where is it in here? Might have lost in the history. It might have been later. It noticed the other PR merge and then get reverted and it got concerned
48:34that maybe there was something more important here that I had to understand and I had to interrupt and say, "No, that was another agent doing things wrong. Ignore that. Everything you've done so far is right. Just keep going."
48:44And it did. The only difference here between the prompts is that I sent the tail scale part as a third prompt and then I had to correct the model five additional times when I did it with Astra. Here I included that as part of the first prompt. Exact copy paste and it got it perfect first try. Exactly what I had intended and exactly what I had in my mind which is find where this went wrong, revert it and then give me a link I can click that actually lets me verify the changes. This shouldn't be
49:13that hard. I'll be real. If 56 Soul screwed this up, I would have been insulted. Astra blowing this one egregiously. One of the legitimate worst AI code experiences I've had this year without question. In fact, I'll say something bad. Discounting things like Flash 38 and like obviously meme tier models. Astra has had the most bad model experiences I have had of any model this year without question. Without question.
49:42It's just random [ __ ] like this. And it's not all the time. It's actually quite rare. But god damn, this model can just suck sometimes. I don't think Astra failed to understand my intent here. I just think it failed to do basic work as an agent. I see chat 50/50 on this here.
50:01Binary Shokan said that they haven't had to do as much feedback for any agent as they have for Astra so far by far. Yep.
50:08It's not always, but when it does happen, it is so ho stupid. I have an important thought that I need to pencil till later because it's more for like the summary at the end. So, we'll get there when we get there. Let's blast through the rest of these. Quick orchestration. It's Astra. Astra has this unbelievable new capability, best referred to as swarms. The way Fable does large numbers of sub agents is it plans up front. It decides, I want these four sub aents for this. I want this
50:38type of sub aent next. And if these four find things, they can spin up the next type of sub aent where it goes through steps one, two, three, four, five.
50:46Astra, however, can just spin up a bunch of sub aents that are doing whatever and then pass messages to them, let them pass messages to each other, and most importantly, it can receive updates from those sub agents and use that to fan things out to keep the whole swarm moving well in the right direction. I've never seen anything else like it. Astra has unlocked novel orchestration capabilities that given a model that wasn't as expensive would legitimately
51:15be the path to something like curing cancer. I genuinely believe that what OpenAI unlocked here is the start of the next era of what agents can potentially do in the size of problems that agents can legitimately solve. And I've seen it in action. It's a little harder to see now because tokens cost money again, but you can see some of it in this run I currently have going for the Rust rewrite of TypeScript. You'll see all the time these interacted with root/independent reviewer interacted
51:45with root/object members regression cause. It's able to manage the context back and forth pass messages and build with all of these agents. And I have 40 running here in parallel. And it can actually keep track of all of them and work with all of them. Unbelievable.
52:04Especially because workflows in cloud code are so good that I did a whole video about why I like cloud code largely around workflows. I never would have guessed that the more rudimentary implementation of sub aents in codeex was because they wanted the model to scale up and work through it. That's exactly what they wanted. It's exactly what it did and it does a great job. One of the capabilities that makes this possible is the steering side here.
52:27Astra is so good at getting random [ __ ] inserted in its context while it's working and not losing track of what it's doing. I cannot tell you how many times I had a model working on tasks one, two, three, and five and then I send, "Oh, I'm sorry. I forgot number four and then it does number four immediately and then never finishes the other one, two, three, and five tasks."
52:48That was just the default. I was used to that. Fable's better about this. It's or Fable 5.1 just doesn't do this too too much, which is nice. Astra eats this. It loves this. It begs you for more. One of the really cool things Astra does now when you're in a harness that is set up properly for it or an app with the right harness, for example, T3 Code, which funny enough, T3 Code actually does this feature better than Codeex, even though it is an Astra feature. It can ask
53:16questions while it works. So, it could be doing a thing and notice at some point while it's working, huh, I would like to know if the user's okay with me doing this. Or, huh, I wonder which of these three things the user would prefer. And it doesn't block itself. It keeps going. But at any point, you can answer. And if the model's already done, it will spin back up and address your answers. Or if it is still going, it can steer it in the right direction without interrupting the work it's doing. It's so cool. And it's like a meaningful
53:45behavioral change. we haven't had as many of recently. Like models work the way they work. We just have to prompt better. This is a change in how the model actually operates. That's really nice. When I was first testing, they didn't have this in codec and obviously I didn't have T3 code either. So I would just see it ask a question in its reasoning and then not be able to answer it and it would just keep going. Now with this, it's actually quite nice. So yeah, the steering stuff way better with Astro now. Fable was far ahead before.
54:13Soul would just get lost if you tried to steer it. Fable had a pretty solid default here. Astra has actual new capability unlocks for steering, which I think are cool as [ __ ] especially when you combine that with the orchestration because now it can have context injected from 40 other agents and handle that fine, which is awesome. And when you combine that with the self-prompting, which I will also say Astra is much better at here. Astra still writes worse prompts than I do, and I would say most devs who do this a lot do, but it can write a decent prompt. I've seen it
54:43write some pretty decent prompts. I'm much happier with Astra's ability to prompt itself and other agents. Nice.
54:49Not as big of a gap as the other things, but it is a gap. Cool to see. I'm excited for a future where agents actually understand things like a claude MD or a skill file well enough to write good verbiage for agents because right now it sucks and you quickly end up in a slop loop. If you let the agent write the skill, you let the skill be involved when it writes another skill, you end up in slop hell very fast. This helps avoid that. Astra still falls into it. Fable falls into it even more. Please write your skills by hand or at least audit them and make some nice changes.
55:19Self-prompting honestly fits under the skill writing stuff as well. Same gap there.
55:25Like none of them are good at it, but Astra is better at it slightly. And then we have skill using. This one is interesting because all models should be able to use skills totally fine, right?
55:36As I've crashed out about many times now, sadly, Astra has some skill issues.
55:43I don't know what's going on. I have learned recently that there is a system prompt section around how skill should be applied in codecs and part of it specifies that if a skill is used in one turn, that skill should not apply for future turns unless requested. In the example I gave in my previous videos and in the podcast where I told it to babysit a PR and then it stopped and then I asked it if there's things worth addressing and it said yes and then didn't do it. In that one, it had pulled
56:11in my babysitting skill and then just pretended it didn't exist. Even though it was still in context, it just ignored it for the next five follow-ups. It sucks. I never ever ever have to think about this with Fable. If the skill has a reasonable description and it's useful for the work going on, Fable correctly pulls it in and applies it. It just does. I don't know what the [ __ ] is wrong with Astra for this. And I honestly think a lot of it is codeex.
56:34Astra doesn't feel like it understands my skills anymore in a lot of ways.
56:38Fable 5.1. It does. It absolutely does.
56:42So yeah, skill writing. Astra has a slight lead. Skill using. Fable has a huge lead. Now we have honoring refusals and boundaries. And as you have already seen, Astra is forgetful. It's okay at honoring things when you tell it to, but it it might forget. If you tell it to keep it in context, it usually will, but it still is quite forgetful. And it is still quite frustrating when it just does a thing it's not supposed to. Fable
57:11isn't as good at strictly following instructions, but that also means that when it doesn't follow them, it doesn't feel as egregious. Both are still not where I want them to be here, especially for the levels of intelligence. I would say Fable is ahead. Astra looks and generally acts better about this, but its failures are more egregious, which is why the hugging face hack happened.
57:29So, take that as you will. I have more to say about this and also all the code stuff. We need to talk about cost. I've seen some very dumb takes on cost here, like exceptionally dumb ones. First and foremost, token costs. They are the same except for one important exception, which is that Fable 5.1 has a huge drop in cash read costs. They went from a dollar per mill in for cash read to 25 cents per mill cash read. And that's
57:58awesome. That genuinely is. And it would be a lot more awesome if cash reads were more than 10% of my costs. They are closer to 3%. And now they are 1%.
58:10Awesome. Do you know what is much more than that? Cash rights. Cash write costs are over 60% of my costs with Fable. So while it has made cash reads hilariously cheap, it has also made cash rights hilariously expensive. I am spending more money asking Anthropic to save the state of the model on their servers for five minutes than I'm spending actually running the GPUs. I am paying Anthropic
58:38more money to manage RAM for me than I am paying them to run compute for me.
58:43And if you don't think that's absurd, I don't know what to tell you. But I'm very excited for the cash write revolution that has to happen. And paying for RAM storage effectively is so stupid. And I am sure we will fix this soon. We're going to get worse before we get better because OpenAI didn't used to charge for cash rights. They used to be free. Now they do. And they are not cheap. It is 25% more expensive than a normal read. So if reading is $10 per
59:12million, cash reads are a dollar per million for OpenAI and they are 25 cents per million for Fable 5.1. Cash rights are 1250. You're going to spend a lot of money writing cash on this model. I promise you on both of them even. So know that going in. The cash read is it turned a 3% cost down to one. But cash rights are where the money actually matters. So if either of these labs make cash rights cheap or free, they win cost by default by far. But the actual token
59:40cost barely even matters in a world driven by token efficiency. God, you guys, I I literally just explained this in chat. is doing. Guys, guys, guys, Bill, you're better than this. You're one of like the more important devs we have on T3 code. What did I just say?
59:55The difference in price for cash reads looks really big because a dollar is four times more than 25, but when it adds up to less than 5% of your spend, it doesn't matter. It's like saying that Mac OS is a 100 times faster than Windows because it gets through the splash screen at boot 30 milliseconds faster. You're you're shaving a percentage off a percentage. It doesn't matter. Four times cheaper for 3% of
1:00:25your cost is a less than 1% deal. It just doesn't matter. It's fine, Bill.
1:00:31You weren't here for it. But I really need to emphasize this point cuz I've seen some incredibly smart people say some incredibly stupid [ __ ] about this because they just haven't looked at the numbers. And if you need proof, it's pretty easy to find. Cost per task on artificial analysis intelligence index.
1:00:47GB6 Astro was $3.26 per task. Opus 5 almost $6. Fable 5.1 $7.60.
1:00:56That is a comical gap even though they are priced the same. And theoretically speaking, the input token cost is four times higher on Astra. It is still a fourth the price in most real world work because it is so much more token efficient and cash reads are such a small percentage of your costs. That said, Astra is 326 and Soul was 199. So Astra is a meaningful increase in price compared to Soul, but it's still way cheaper than all of what Anthropic has
1:01:25been doing. Also, funny enough, previously they had said Fable 5.1 was more expensive than Fable 5, but when they redid the artificial analysis like index because they were getting cooked because it just was not measuring things well at all anymore. When they switched to the new index, 5.1 actually got cheaper than Fable 5 because the cash read difference finally actually mattered. I made the points I want to make here. D6 Astra is still way cheaper even though the prices are the same and the cash cost is higher because in the end there's a lot of other things that matter and that number is not where the
1:01:54cost is happening. This number is where the cost is happening. Token efficiency.
1:01:58GBD6 Astra is one of the most token efficient models that they've ever benched. I removed pretty much everything here that inarguably does not matter. And Astra is the most token efficient by far. Meanwhile, Fable 5.1 is the least by far. Astra did in 27K tokens what Fable 5.1 did in almost 80K tokens. Take it as you will. It's also worth noting that OpenAI offers a flex option which cuts the prices in half but kills all the guarantees for throughput
1:02:26because it's flex so it doesn't start responding immediately. It just happens eventually. So if you're willing to let your jobs take an unknown amount of time, you can put on the flex end points and let it run when they have compute around and it will cut your costs even further in half. Now that I think about it, it would actually be nice if they could add that as an option for using with codecs because I don't mind if my threads take a while and I also am aware that a lot of the time I'm using by agents is at bad times for businesses.
1:02:54So since I'm working at like 9:00 p.m.
1:02:57and not like 10 a.m. when businesses are, it would probably be cheaper for OpenAI and me. Just saying it is what it is. But that means we're talking about subscription limits. And I think it is fair to say I am uh pretty experienced with the limits on all of these things.
1:03:17Being that I have five accounts with Claude and that I have four with Codeex, think I have a pretty good gut feel here. And to be completely frank, if you are concerned about maximizing how much return you get for the dollars you're putting in and you're doing anything other than the codec subscription, I would love to meet the person who convinced you Claude was a better deal so I can hire them. I always need good sales people and I could really use that one because they sold you a [ __ ]
1:03:47bridge, man. Seriously, the gap in what you get is insane. You might be a bit confused because I've covered the numbers before. And with Anthropic and Claude Code plans, you get around 8 grand of inference for 200 bucks. And with codecs, you get around 12 grand for 200 bucks, which is better, but that's only like a 50% improvement. Well, there's a few factors you have to consider. First, we're have to take a huge cut in our quad code limits because they're going to drop off that 50% boost they're giving us. They're going to keep at 25%. So, we're losing around 17% off
1:04:17the top of our limits. But much worse, the fable limit. You only get to use half your usage for fable. You might notice in my UI here that there's a third column for the claude section that is faded out a bit. I call this the opus pot. It's a whole steaming smelly pot of opus. And that's all it is. And when I am heavily coding, what ends up happening is the column on the left all
1:04:45go to zero and the column on the right all get stuck at 50 because I'm allowed to use half my weekly limit for Fable and once I'm out of Fable, the account is useless to me because Opus sucks and Sonnet's hilariously bad. So I end up every month losing those 50%s at the end of each week, at the end of each reset.
1:05:05And since I get eight grand a month of allocation per account, I'm only getting four grand of that with Fable, which means it's actually a third of what you get with Astra in the $200 CEX plan. You get 50% of 8 grand with Claude Code and you get 100% of the 12 grand with Codeex. And they made further improvements to how Astra is being hosted and how it uses your limits. Tibo claimed is up to a 3 to 4x difference. And I don't
1:05:34necessarily believe that, but I've seen a meaningful difference. Like I have some hellish health threads going right now. I have two ult running with unbounded sub agents. It's 40 sub agents each can run. And it's going down. I think it's hitting this particular account right now. And it's at 82%. It's been going for the last 8 minutes since I refreshed. Not even. And it's down one more percent. Cool. if you weren't doing the absurd, sinful, wasteful background
1:06:02jobs I am running, trying to dick around with my new fork slash uh decomp of Super Smash Bros. Melee while also trying to port the entirety of TypeScript to Rust. It's not that bad.
1:06:14You can hit the limits and you should be very careful of the alter button and in particular the fast button. Straight up, fast is not the way you should use this model. If you're using fast with Astro, you're probably using Astro wrong or you're just burning tokens for fun, which I understand. But like if you're actually finding yourself reaching for fast all the time, you're not using it for its strengths. You're not using Astra for what it's good at. But you can get a pretty good amount of usage out of these plans. Not great, but a good amount. There is a nice powerful catch
1:06:43to this we'll get to in a sec. But I have one last thing I need to say about the Claude stuff that has begun to piss me off a lot. Openai does not do five hour limits on the pro plan. So the $100 and $200 tiers do not get fivehour limits. They only have weekly limits.
1:06:58All Claude plans still have 5-hour limits. Each 5-hour limit gets you about 40% of your Fable 5 limit, which is 20% of your weekly. So, you can clear the 5 hour on your Claude plan five times before being out of weekly limits, unless you're using Fable and then you'll run out in two and a half. The reason I set up CLI proxy is because one cla account just isn't enough at all for most work. If you're using it a few hours a day, two to three days a week,
1:07:28and you're not going too hard, and you're not using a lot of sub agents, you can probably get by on a single $200 plan with Claude. The only way a $200 Codeex plan isn't enough is if you are like using higher reasoning limits than necessary and not paying much attention.
1:07:42Most devs, when used responsibly, could absolutely get away with the $200 plan with Codex without issue. It would be genuinely hard to get away with it with the Fable plan. I even call it the Fable plan because that's what it is to me.
1:07:54It'd be much harder to get away with it on the Claude code plan because of the additional limits on Fable, the much less efficient usage of Fable and everything else. So, Astra is more token efficient. It costs less, so it uses your usage less. It has way more generous limits in general with Codeex.
1:08:12It's just like I would say it's like a 4xish gap between the two is how it feels to me. feel like I can get four times more done with Codeex on Astra than I can with Fable and Claude Code on the same $200 a month tiers. But then things get even more ridiculous because of the resets. It's become a meme. Tibo just throws them out for fun now. In fact, when they were delayed trying to get Astra out, they sent us two bananked resets. I actually opened this account that I don't even need yet because I
1:08:42wanted to collect those bananked resets so I'd have them ready to go when I did need them. So, your limits just get reset all the time. They got reset while I was filming earlier. So, even if you do hit the limits on Codeex, you might just randomly get a reset, which is quite nice. So, yeah, it does happen with Claude. It's just so much rarer that it you can't really count on it at all. But with Codex, the resets giving you your limits back much nicer.
1:09:04According to Raphael, there were eight free resets in the last 30 days. So, the weekly limit is more like a 3-day limit on average. Pretty nuts. Somebody accused him of maliciously resetting 46 hours after the previous global reset in the time frame around where or in the time frame around where a banked reset was landed. This is not true at all.
1:09:24I've actually seen Tibo intentionally delay a reset in order to prevent that.
1:09:29And even just now, he announced the reset many hours before so we got to go spin up our furnaces and burn a bunch of tokens. He does not maliciously align the times there at all. He actually does the opposite. So yeah, if cost is a concern, wait a few months because this level of intelligence will be accessible at a much cheaper price. But if cost is a concern and you want the best class stuff right now, you can get a reasonable deal with a $200 plan with codec still. If you're curious about the $20 plans and the $100 plans, my
1:09:59suggestion would be wait for this level of intelligence to get cheaper because this level of intelligence is expensive.
1:10:05I'm legitimately doing 30 to $40,000 of inference a month on all of my plans.
1:10:10consistently now, even without early access testing that's free. They are not cheap. Astra is cheaper than Fable, but neither are cheap, and you need to know that going in because if you go in with different expectations, you will be upset. They are both expensive. They will both burn limits fast, but Astra is more efficient when you compare the two directly. So, with all of this said, let's answer the question. Which model is better? I could only pick one. Which
1:10:38would I have? Let's say orange is fable and I don't know, we'll say blue for Astra. Cool. This is meant to be the quality of the outputs over time, over attempts, whatever. This is generally speaking, what's the quality I get out when I send a prompt with Fable 5.1.
1:10:59It's a relatively stable line at a relatively high bar. It has its moments where it impresses me and it has its moments where it disappoints me a bit.
1:11:06Occasionally it gets pretty rough, but for the most part, it stays in this general range. Quality I'm getting out of Fable pretty damn good, and I would feel bad complaining about it. It does have its weird spikes, and I'll have plenty of videos where I show the times Fable piss me off. Don't worry. But generally speaking, it does what I expected to do when I tell it what I want. Astra is why I'm drawing this chart, though, because Astra can do things Fable literally could never and blow me away. But then it does the
1:11:36stupidest [ __ ] I've ever seen a model do and it's so goddamn spiky. This is the thing I really wanted to try and communicate. The quality bar with Fable is relatively consistent. It goes up and down a bit, but never but more than like two points in either direction. When I send a prompt to Astra, it's equal chances it drops my jaw because I'm so blown away with how unbelievable it is.
1:12:01or it drops my jaw because I cannot fathom that I just spent a thousand dollars for it to run in a loop and not ship anything and then break my website.
1:12:08And Ben in chat said that each line also represents your blood pressure while using each model. Yes, I would honestly say Astra stresses me out more, a lot more. This is the best I can explain the difference here. If you have low tolerance for [ __ ] if you leave your computer when the model does something stupid because it pissed you off so much. If you are easily agitated to the point where it affects your work, you should pay the extra money for Fable. You just should. It's a way more
1:12:37consistent, reliable model. Astra, however, is way cheaper. It can do things no other model can, and it impresses the absolute [ __ ] out of me when it does hit those peaks.
1:12:49Personally, I like having both, and I find myself rotating between the two quite a lot. If I had to pick one, I would pick Astra and have it write DMs to me to Julia so that he could send off the prompts using Fable for me. And I know it puts me in a weird place because I will always find some way to get good code out of the anthropic models even if I can't do it directly. Astra has more novel uses and especially now that I'm down a hand, Astra's ability to like use my computer and get work done and
1:13:19benefit my life outside of code. I almost feel like as a heavy user of computers, the $200 C codeex plan almost feels essential because it if you use computers professionally and you make over $100,000 a year, you or your boss should be paying for the $200 a month plan just because it makes your ability to use your computer more effective. If you don't have that and you're saving up in order to pick the thing that writes the best code and you're really easily stressed, I would still pick Astra, but I would use Astra to make Astra better.
1:13:46Find everything you can to smooth out the rough edges. put together the best set of skills in the world and make benchmarks that prove why they're the best. Blast it on Twitter, get a job at OpenAI, and never have to worry about token spend again. What I'm trying to say is I don't think that person actually exists. A lot of you act like that person does. They don't. If you're concerned about costs and that's your main motivation, go get the open source plan. See if you can convince either lab to give it to you or just wait for these things to be accessible in cheaper
1:14:14formats. They will be. GLM53 Flash is a more pleasant model to use arguably than Astra. I've seen GLM53 Flash go off the rails less than Astra for similar work.
1:14:26So, I would expect this level of capability to be cheaper and more accessible in things in the near future.
1:14:30So, just be patient if you're feeling like the price is too high. I get it.
1:14:34That said, Fable is what I reach for when I'm trying to land code and Astra is what I reach for when I'm trying to use my computer. The gap for code is not as huge as it might sound. The gap is more in this chart here, but the harsh reality is just this gap in capability.
1:14:52This is the post I made when the models both came out to try and explain why I like both. Astra is world class at a shitload of things. It is genuinely the best model in the world at all of this stuff, and half of it it is far ahead for, but Fable writes code that is mergeable 20ish% more often, and its stupid spikes are much less stupid. Its smart spikes are not quite as smart, but its stupid spikes are way less stupid.
1:15:17And for that reason alone, I think Fable 5.1 is a better choice for day-to-day code stuff when cost is not a factor.
1:15:24But if you're willing to put the effort in with Astra, you will get a lot for it. Both models are awesome. I will be defaulting to Fable for code and I'll be defaulting to Astra for every other thing that I do. Actually, one last fun side thing that I should have mentioned before. This will be a really silly one to end on. Remember everybody being really upset about Fable's security refusals? Astra refuses more often. Just wanted to put that one out there.
1:15:48Anyways, both these models are great. We are in an unbelievable time to be engineers. The fact that we have technology like this coming out every few months that massively changes how we work is so godamn cool. It is hard to pick wrong here, but if you don't see the benefit of all of Astra's incredible capabilities, I am confused because they are useful to everyone. And if you don't think payable is worth the extra money for the quality code difference where it's not that big a difference, I totally understand, but you're not shipping enough if you don't see why
1:16:16that matters. Either model is great.
1:16:18They're both unbelievable. They both have changed how I work. And when you look at the numbers for how we're using them, both models have massively improved our productivity within T3 Code. And they don't replace each other, they compound each other. These models are great and you shouldn't feel bad using one or the other. You should experiment with both because I think you'll be blown away with what's possible. And pretty much everyone who can have the codec sub probably should because it's such an insane deal. Have fun burning tokens and using all these models. I think they're awesome and I
1:16:48bet you will too. Until next time, peace nerds.