0:00Before we get started with this one, I have a question for you. How many agent threads do you have running right now? I don't mean have you run today. I don't mean are you going to run later. I mean in the exact moment where you clicked this video, did you already have things running in the background? If the answer is less than five, you really really need to watch this video because you are not taking advantage of the genuinely awesome and unique opportunity we live in today. And if the answer is more than five, you should also probably watch this because you'll get a lot of nice edges and solutions to problems you might not have even known about. and
0:29you'll have a great resource to share with your friends who are not taking advantage of where we're at today. I'll be real with you guys though. This video is not my norm and as such, this doesn't feel quite right. I don't think we need caffeine for this one. We're going full rodeo.
0:48Much better. This is the video that Anthropic does not want you to watch, and I'm going to make it very hard for them to do anything about it. So, before you get too worried about bans, I promise I've done everything in my power to prevent it. And of the bans that have occurred across me and my friends who do all of these things, of the 30 plus claw accounts we're hopping between, I've had one friend get banned twice because he was on a VPN when he registered an account with a credit card for a business that was not his. It was mine.
1:15And those two accounts got temporarily banned and reinstated almost immediately after. So, while this is probably not something the companies want you to do, it is a thing that they are not stopping us from right now. In fact, the only thing that can really stop us is a quick break for today's sponsor. I have two quick questions. First, have you ever used an agent without search? If you have, you know how painful and miserable it is. It basically can't do anything.
1:38My second question is the opposite. Have you ever used an agent with insanely fast and accurate search? I personally hadn't until I started using today's sponsor, Parallel, because they have the best and fastest search results for pretty much every single thing you can measure. Historically, I've really liked the search built into OpenAI, but when you compare it to Parallel, it's just night and day. Watch and see just how fast Parallel can get you good results.
1:58Going now, all real time, of course. It took just over a second for Parallel and OpenAI still going. Still going.
2:06Meanwhile, OpenAI's endpoint took over 6 seconds long. If search was all they did, they'd be one of the best options available. But they also have everything else you would need on the web. From a proper monitor that will send your agents info when things change on different pages to a traditional response API for when you want to actually get synthesized results from your search queries to their extract endpoints that let you send a URL and get back the data that your agents actually want and need. Even parsing JSheavy pages, by the way, there's even an MCP so you can expose parallel to your existing agents in whatever tool you're using. Now, every month you'll
2:36get 5,000 requests for free. And on top of that, if you sign up today, you'll get $80 in credit. What are you waiting for? Join now at swordv.link/parallel.
2:43This is a real real rough order of events that I have planned here. We're going to start with how do you actually get enough tokens to do the terrible things I'm about to show you guys. Right after that, we're going to talk about actually accessing them once you've taken advantage of the ways to get more.
2:56Then I'm going to talk about how you actually use them practically well. And of course, right after poorly, which I'd argue is more fun and much more important. And then at the end, the thing that we all want. How do you use your tokens when both you and your laptop are asleep? While I will do my best to make this a tople guide on like what things you should do, the thing that will matter more is the little pieces that I show. Not because you should copy all of them, but they should help you understand how I think and operate with agents and how I've been able to do like 100 plus PRs a day for a
3:26bit now part-time. The point here is not to give you the exact formula to copy.
3:32is to get these ideas in your head and these examples floating around in your brain so that you'll apply some of these lessons in hopefully your own ways. Like the best possible outcome of this video would be you guys going off building more cool [ __ ] with agents and then coming to me with solutions for problems I hadn't even thought about yet that make my life and my agent maxing better.
3:52I can already tell this video is going to cause huge blowback on Twitter. So, I'm going to just list the questions they're all going to bring up to try and dunk on me for making this ahead of time. So, let's go through the common questions and [ __ ] comments that I know are going to happen the moment this video goes live. This only works for side projects, not real apps. I have a good feeling if you're saying that my apps are more real than yours. I also happen to know that a lot of the strategies I'm showing here have worked at companies scaling as big as I don't know AWS itself where I have friends who
4:22are learning these lessons from me as well as at Microsoft where I have friends who are learning these lessons from me and even some Fortune 500 companies that are doing some of the subscription abuse stuff that I'm going to show y'all. We'll talk about that in a bit. Not all of us are rich. Cool. I'm sorry. This is a great way to save a shitload of money though. These strategies work at small scale and large scale. And the crazy numbers you've seen me spending in the 300k plus range are not spent at all. I'm putting in like $1,000 for every 50,000 of tokens I'm
4:50taking out. Like worst case. And if you're doing worse than that, you're paying API prices. We'll get there in a bit. Of course, a JSE thinks this. Oh yeah, the devs who understand the importance of having software people can use instead of software that people can gawk at and not use. Yeah, you're right.
5:07I do know what people want, which means I am pre-qualified to talk about what people want in building real software.
5:13Wow, OpenAI has paid you off. I have paid OpenAI more money than you're worth. Wow, Anthropic has paid you off.
5:19I've paid Anthropic more money than you're worth. Now that we've got this out of the way, let's actually be useful. First and foremost, getting more tokens. Let's do some math here. Who here knows how much money in tokens you get when you put in $200 on the Cloud API? give you a hint. It's $200.
5:38What about the codeex API? I'll give you a hint. The answer is the same. Which is why what we will be doing today does not involve the API. You might think this is silly or dumb or not viable for real businesses. If I was to list the big companies I know that have their employees subscribing to personal tier accounts in just turning off the setting for data sharing, they'd all be really mad at me because these companies are very private and they don't want these things known. But they're all [ __ ]
6:07doing it. You have no idea how many companies are just paying for the subscriptions. And when you do the very basic math, you'll understand why.
6:16Because a $200 Claude sub gets you around $8,000 in tokens.
6:22That's 200 bucks a month for 8K in tokens. There is a catch for this one, though. Only 4,000 of that is Fable because they have the 50% limit on Fable separately. That still means you're putting in $200 and taking out 4,000 to Fable. That's a very good deal. The $100 tier is weird because you do actually get half of this total amount, but you get a fourth of the hourly limit. So the 5 hour limit is much more strict, which means it's harder to burn that all on
6:51the $100 account. So personally, I don't really recommend the $100 account. I think you should save a little more, do the 200 and burn it, find a way to make money off how you burnt it, and then reinvest and keep going. And the $200 codec sub is even more interesting because it's roughly $12,000 of inference. And there is no distinction between what you're doing with Astra and with other cheaper models like Soul. They don't split it that way.
7:15You just get the tokens. And if you guys think these subsidies are brutal, I got an even crazier one. You all know how much Fable costs is $10 per mill in and $50 per mill out. That sounds expensive, especially when you compare to other cheaper models, especially in the open weight world where there are things that are doing real work for a tenth or less that price. What if I told you this number was also more than 50% off? Fun fact, when Anthropic first announced
7:45Claude Mythos preview in April, they actually quoted a price even though the model wasn't out. This was just for internal use at companies for Project Glass Wing for securing their stuff.
7:56Those companies were offered to continue using the model after at $25 per million in and $125 per million out. That means the price we're paying is like almost 70% cheaper than what they intended to charge. Why would they ever make it so much cheaper? It's because their margins are insane. The prior rumored margins for how much money they made selling tokens at API price compared to like the energy costs and the eventual usage of the compute causing it to fail over time
8:25roughly calculated from various napkin math afficionados I know to around a 95% profit margin. So if they charge you a h 100red bucks, they made 95 and they spent five on replacing dead compute and paying for electricity or in the case of Anthropic renting the compute from SpaceX. So why would they knock this price down so much? Well, Anthropic knocked the price down because they know everyone has the margins, but they were already in the lead and they wanted to maintain their lead. So they decided to
8:54make the price lower in order to make it harder and harder to justify using competitors. I would argue that at 25 in, 125 out, this model was, especially the time, worth it. But at $50 out, it's a bargain. And at $50 per mill out, combined with the 200 to8,000, so the 1 to40 ratio of subsidization means you're getting some really good deals. So you
9:23can kind of mentally double the numbers for Claude as well. and also codecs because there is no world in which OpenAI actually plan to charge the literal exact same rate as mythos.
9:36OpenAI does numbers very simply. They keep them the same. They lower them 20 to 80% or they double them. Those are the only things OpenAI knows how to do except for after Fable destroys them and after that all of a sudden they're going to nudge the price a little bit. So the reason that Astra is as cheap as it is is because of Fable. And the reason that Fable's so cheap is they decided to eat their margins a little bit in order to guarantee a win, at least at the time.
10:02And this is all without even talking about the resets. Last I saw someone do the numbers, Tibo gave out 14 resets in 30 days. That means you're averaging a reset every 2 to 3 days with your codec subs. That means that the weekly number, which is 4 grand, is more like a 3-day number, which means you can double the codec sub here to 24 grand if you're really maxing it out. It's so bad right now that they've temporarily paused new subscriptions in the $200 tier for
10:30Codeex, which sucks due to the nature of this video. And I have a bad feeling this video is going to make the problem worse quickly. Sorry in advance. Feel a little bad exposing these things to the world, but my job and my loyalty is as a journalist and informant, not as somebody who's trying to protect and shelter the few hundred people who know how abusable this [ __ ] is. I think everyone should know and I'm going to share this info with the world and how things proceed after is how things proceed after. Don't be mad that I'm sharing. Be mad that I'm right. Okay.
11:00When you set up your account, the first thing you should do in codeex is hop over to settings, go to data controls, go to improve the model for everyone, and turn it off. As soon as you do that, the terms of service for your personal sub are effectively identical to the terms of service for a team account.
11:17There is no real difference. So if you are paying for API prices because you think you're not going to have your data trained on, you're not saving anything.
11:26You're just wasting money. Every dollar you spend on API is 80 to 90 cents you've thrown away. And if you're not thinking of it that way, that's on you.
11:36There's a similar setting in the cloud code settings. I don't feel like going to find it. You get what I'm talking about. So now with all of this established, the best way to get more tokens is obviously to get more accounts. And since we've established the accounts aren't meaningfully different from API pricing, you get a good deal. You should still use the API for a couple things though. One is very important. Anything facing users, if you are rating your prompts or your agents are rating your prompts or the prompts are running on your system on your codebase from things that are in your
12:05GitHub issues or whatever, all that is fine. That is awesome. But if you are trying to serve public traffic where users can submit a request and it runs on your inference that is entirely against the rules with the accounts. If you're using this as an alternative to the API to serve user traffic, you deserve your ban. That's not what these accounts are for at all. They are being very generous letting us experiment with coding with these huge subsidies. You should abuse the generosity by building cool [ __ ] not by reselling things to
12:33squeeze margin that they aren't. I'm giving this advice up front because the things I'm about to show would make it hypothetically easier to do that, but I think it's time to get there. It's time to talk about how we actually access the tokens. In my first token maxing video, I gave the example of using two cloud code subs. And I showed that if you have cloud code running in the terminal and it's completing a task and you use another tab to change the off for your cloud code instance, it will correctly recover and keep going. That is great
13:02when you have one or two computers and one or two threads with two accounts.
13:06That does not work when you are doing what we are going to be doing here. You need systems that manage this. Not just for your local cloud code with your one or two accounts. You need a solution that will handle the insanity that is how often the cloud code off breaks in a single shared bucket on a residential IP address that you can route all your other traffic through. So let's go through some pro tips here. So tip number one, I love VPS's. We'll talk a
13:34lot about VPS's. I highly advise you do not sign in to cla or codeex, but especially cloud code on a VPS.
13:42Anthropic in particular is very very nervous around people reselling their cloud code subs in VPS's serving traffic to other people with cloud code because it lets them sell it for cheap and also collect a bunch of data they can use for distillation. So they're being extra aggressive with signins and off and requests coming from IP addresses that are known to be servers like anything on AWS or HNER or any of these other providers. So I highly recommend you do
14:11not do this over a VPN or on a cloud that you're serving the traffic from. I would also advise against having that traffic come from multiple different computers at the same time because that's how they know you're doing some sketchy stuff. So, how do we avoid that?
14:28How can we do our best to make sure most traffic comes from one place, ideally one residential IP address, regardless of what you're doing, what account you're using, and where you're using it from. Especially with something like cloud code where the O breaks every two to three days. Leo from the future because I forgot an important disclosure before I go any further. This might get you banned. This might be against terms of service. I am just speaking from my experience. Do not hold me accountable.
14:50I am literally drinking and filming as I talk about all these nerdy things. If you get banned, and I didn't, that's unfortunate. I'm sorry. I can't be held liable. You'll probably be fine, though.
15:00Just don't do this with your Google accounts because when your Google account gets banned, you're [ __ ] Thankfully, there's no use for Gemini models anyways, so there's no reason to deal with that. Cheers.
15:11As we were saying, it's time to talk about the proxy. CLI proxy API is a thing that's come up in a lot of my videos, but not really been showcased in most of them. The way that this works is you sign into it using your O for codec and cloud. And once you have signed in, it will maintain the off for you and give you an API that you can use in things like cloud code and codecs where you point it to this instead of pointing it to the traditional ooth servers. And
15:40now everything is routed accordingly. It is very very convenient. We'll get back to this in a sec cuz I realize I forgot one other thing we need to talk about in the getting more token section. So one last thing here so I can actually wrap it up. You might have noticed the subscriptions I was talking about. I was talking about claude subs and codec subs. I didn't include other subscriptions here. Why did I not include other things like the open code sub or the subscription to T3 chat or the
16:09subscription to cursor? Even the subsidies are comically comically less. Curser does probably get a deal with anthropic. The deal is probably 30 to 50 at best% off. That is very different from 95 to 98% off. You get a better deal than cursor does. You get a better deal than open code does.
16:31You get a better deal than pretty much everyone does. Did a reset just get announced for codecs? Really kind of crazy that even in a compute crunch like they have right now at OpenAI where they don't have enough compute for all the astros. They had to cancel the subs.
16:44They're still giving us resets. I will point out it is currently Friday evening which makes it a lot easier for them to do this because they are hoping the power users like us, yes I'm including you now viewer, you're joining us on this journey. The power users like us will burn it all to the ground before Monday when the enterprises wants to pay API prices again. So yeah, the only subscriptions that you can really massively benefit from in terms of you put in money, you get more tokens out
17:13are Claw and Codeex. If Grock 47 comes out good, it might make sense, but I just it is hard to token max with Grock because the Grock models just cannot stay coherent for as long yet. So back to where we were with accessing the tokens. I gave you the prerex. You want it on your home network. You want it in one place ideally, and you want everything to route through. CLI proxy is a very cool place to start, but CLI proxy not necessarily the best UI. The guy who got me into CLI proxy didn't
17:43know it had a dashboard because he has avoided it to the best of his ability.
17:47Technically, I think the version he uses and that I started from was a fork called vibe proxy, but I have extensively forked this project. I have made a lot of changes and when you set it up and look at the dashboard, it won't look anything like this. And if you want to change that, then congrats.
18:02you have your first token maxing task.
18:04Give it a screenshot of this and tell it you want it to look like this and it'll probably figure it out fine, especially if you use Fable instead of Astra. There is one other change you almost have to make and I would argue these changes combined would be worthwhile somebody creating a like real fork with an easier setup like flow because the other change you need to make is account prioritization. By default, the way CLI proxy works is it routes traffic to all
18:33of your accounts evenly and it distributes across them. Which sounds great until you look at my usage here where you see I have one account that expires at 1 p.m. tomorrow and I have another that expires in 4 days.
18:45Obviously, I want to burn as much as possible of the one that expires tomorrow and not touch the ones that have another few days. The only reason you wouldn't want to do this is because you're constantly hitting 5 hour limits even with multiple accounts. I can't believe I'm saying this in the token maxing video. What I would say there is slow down a little because a 5 hour limit gets you through 40% of your 7-day fable. I've done the math. I know this one for a [ __ ] fact. So, please do not give me [ __ ] I know for sure that 100% of the 5 hour on the 20X plan is
19:15exactly 40% on the 7-day. It's only 20% of your full 7-day limit. But there's a reason I have this column grayed out.
19:22This is the opus bucket from hell.
19:24Ignore that right column. It doesn't matter. We only care about the 5 hour and the 7-day, and we really only care about the fable 7-day. I often accidentally call my claude accounts my fable accounts because opus is a trash model and you should avoid it to the best of your ability. Sup nerds, less drunk future Theo here, wanted to call out something that didn't exist when I filmed this a few days ago. Opus 5.5.
19:43I'll be clear, all the advice in this video still applies and Opus 55 behaves very similarly to Fable 51. So, anything I say about Fable for the most part applies here, too. The big difference is you might not need 5 to 10 Claude subs anymore because Opus 55 on just one sub I have found pretty hard to like hit your limits with. So, take that as you will. Maybe you don't need as many subs to take advantage of these patterns now.
20:07You might not even need to set up something like the proxy I'm about to teach you about. But yeah, thought it was worth calling this out cuz 55 is really solid overall and has changed my way of using things. Opus 5 was a useless dumpster fire that I had to work around a lot. 5.5 pretty good model.
20:23Back to whatever drunk Theo was rambling about. Once you've gotten all of these signed in and you've had one of your models go through and make these changes, changing the routing to optimize for which accounts have the next reset up soonest. And the other one I think is important is by default your codec subs won't go through websocket.
20:44So there's a few changes you can make to fix that which hugely increases performance. Just like the time passing messages back and forth goes down a ton.
20:51So make sure you have the web saga stuff working. It's a little annoying but you can. One last piece that's important and thankfully my agents have been smart enough to do it right. It's called session affinity and account affinity.
21:01What this means is if I have a thread going, ideally the next prompt in that thread goes to the same account because caches are account specific. So if you are 800k tokens into a thread and you switch to a different account quietly behind the scenes with CLI proxy cloud code doesn't give a [ __ ] It doesn't know the difference. But the API then has to go recreate those cache entities.
21:25You only have to do it once per switch.
21:27But if your stuff is switching per prompt or worse per tool call, then you're going to need a lot of cache rates that you probably don't need to.
21:33Thankfully, the Affinity stuff works very well by default in CLI proxy and the caches only last five minutes anyways. So, it's not too too bad. Just thought you should know because if you're using a dumber model to set this up and it gets it wrong, like if you use Opus, it might get it wrong. So, keep an eye on that. Make sure that you are looking at how much cash reading and writing you're doing and that those numbers aren't getting crazy. You'll get a good feel for it soon. You can always just ask the agent, hey, how are we handling such an affinity? Are we writing cash more than we should be? and
22:01it can probably get an answer. For this next piece, I will request that you look up the top of the screen. You'll see the URL I'm on. BB1, which is my framework, my desktop in the other room.
22:13Micro.ts.net. This is my tail scale.
22:16This is my tailet 318 is the default. And then I'm on the management html page. This URL is being accessed this way because I am using it over tail scale. Tails scale is very very good to have for a setup like this because it makes it easy for other machines to access this proxy endpoint and it means you don't even need O. I was experimenting with some things yesterday and I set up an API key for a different thing and it broke my
22:44inference on everything because by default it just works with no API key.
22:49This sounds horrible until you realize it's only exposed to other things on my tailet. So, if you don't have my Google account that I use to sign into Tailscale, you cannot hit this end point. So, it doesn't [ __ ] matter.
23:01So, you're fine. And this also makes it really, really, really convenient to set up other machines. For this next part, I'm going to have to do a thing you probably didn't expect from me in 2026.
23:11We're have to look at a real repo. This is a project that I have on GitHub and on my machine. And despite the fact that most of my projects are on most of my machines, this one's only on this machine. That's because this is my fleet management repo. This is the repo that explains how all of the other computers I do my work on function, including the one that hosts my proxy and manages the tail scale for it so that everything else can connect. One of the files in
23:39here is how to set up a box. This describes all the things I want set up in my boxes and all the random tools and things like I want RGFD, jq, t-mxb, top, node, python, and build tools. all of this [ __ ] Codex and cloud code installed and I want them to use the proxy and all these other things. All the details of all the machines are in here. It gets updated whenever I add a new one. So when I get a new machine, step one, I set up SSH. Step two, I go to this repo. I open up Cloud Codeex or
24:08in my case usually T3 code and I say, "Look, here's a new box. Here's the SSH key. Go set it up the way I like." And in not very much time, I'll all of a sudden have a page open in Helium for a tail scale approval. I'm like, "Oh, that was probably my agent running." I blindly hit accept and now I have the machine set up and it can do inference on these other boxes easily. So now I have what I needed. I have one machine on my home network going through a residential IP address that all five of
24:37my cloud accounts and all four of my codeex accounts go through and all my machines all around the [ __ ] world connect to that overtail scale. And it is impossible for Anthropic or OpenAI to know that I'm using them for servers.
24:50And setting this up in Cloud Code and Codeex is trivial. It's so trivial that I'm not going to tell you how. I'm going to tell you to tell your agent to do it.
24:56It will swap two variables in the configs and you're done. Thankfully, both Cloud Code and Codeex have to work with API endpoints because they're being used so heavily with companies that are doing all their inference through bedrock on AWS, which means they need to put a URL in. So, this will work indefinitely. It would be bad for them if it didn't. In fact, the Codex desktop app handled this poorly. And enough people on AWS complained when Bedrock happened that they fixed it. And now, if you're using Codeex with this setup, not only does it work great, it'll even show you in the corner the name of your
25:25proxy. It's pretty legit. And if you're skeptical or you really really want to have your normal install with a normal O, you can still do that. And you can make another instance of codeex or cloud code with a different home directory that has this configured. So, you have a different command for each. The other real cool benefit of this is if you do set it up, you get access to all the models that are in that setup inside of your cloud code. Notice what I said there inside of cloud code. Do not under
25:55any circumstance use this to use your cloud code sub in something other than cloud code unless you really really like dealing with anthropic support and getting accounts banned all the time.
26:05You will be banned. You will be upset. I highly recommend you only ever use Claude models through your Claude subs in Claude code over the proxy. One last thing I think is worth knowing about the setup because it catches me off guard every time. Anthropic resets work different than Codeex ones. The difference is the dates. When Codeex does a reset, it resets your weekly entirely. So if your weekly was going to renew in a day or 10 minutes or whatever else anyways, your reset is now 7 days
26:35off. So, when the reset T-Mo has promised hits at midnight, all my accounts that are currently 3 days away from a reset will now be 7 days away from a reset. And this means you end up with a weird pacing issue where you can quickly get to a point where you're 6 days away from a reset across all your accounts, which really sucks. It gets kind of balanced out with the banked resets, which across my accounts I have three, six, seven, eight, I have two more in my other. So, I have 10 resets across my codec subs. And the two types
27:04of resets are banked and immediate.
27:06Banked resets are the ones you can click whenever and those cost them more money because the reason they can do these resets so freely is they do them in hours where they're not competing for the inference with their enterprise customers. So they can freely give them on Fridays and Saturdays, especially Friday nights. They can't give them so freely on a Tuesday morning when they're about to have all their customers trying to use Astra at work that are paying API prices. So the banked resets cost them a lot more literally because it's opportunity cost. I do see a future where they start pushing us to only use
27:34these subsidized subscriptions at off hours. I'm legitimately at a point where I would shift my sleep schedule in order to maintain this level of subsidization because it's so good and I'm [ __ ] addicted. Let's be real, and we all will be soon. So, now that this is all established, the difference with Claude is there's only one type of reset. It's a limit reset. And not a timer for the limit, just the percentage for the limit. So, if Claude was to do a reset right now, the only thing that would change is all of these numbers become
28:03100% again, the dates and times all stay the same. So, if I was to get a reset from Claude right now and my next actual one was in a day, that means I need to turn on the furnace immediately. I need to burn all the tokens I can because they'll vanish at that point. And this is the mindset shift I need y'all to get in when your normal reset hits. So, at 1:00 p.m. tomorrow for me with this account, whatever I have left here is money I lost. Don't think of this as you
28:33spent $200 and you got more. Think of this as you've been generously gifted $4,000, but every week where you don't spend a,000 of it, it disappears forever. This is the opposite of the mindset you're in when you're paying API prices where you're trying to get as much as possible for as few tokens. You need to have the depressed mindset where you feel yourself losing money because you are and we don't want to be stuck at the permanent underclass. Burn your tokens.
28:58Cool. That's most of what I need to show in this dashboard. I'll probably be back here. I know I will be back here. Let's be real. I guess this is most of the accessing your tokens bit. Which means next you talk about using the tokens.
29:11Well, this one is going to be full of all sorts of different layers. What I'll say for the core of this one is I have another video that's probably already out by now that's about how I code without one of my hands. It's a video all about my life after accepting that I can't really type. Even after my surgery and I get my hand back, I might not be able to type very well, which means I have to use my computer different. Not just like voice to text, but I don't want to switch between apps as much cuz I can't command tab. I don't want to be
29:40staring at the thread as it generates because I have other things to do. I don't want to be at my computer that much. I am sitting at my computer less and I'm coding more. And these are the strategies that will get you there. That video has a lot of the good examples, but I do want to give a couple important points that I don't necessarily think were in that. And then also show you some dumb examples that weren't there that could be useful. The question I want you to start getting into your head is, how can I use more tokens to care less about this? When you have a problem
30:09that you're trying to solve, how early can you pull in the agent to start solving the problem? And how long can you have it go until you need to take another look? This is a twofold thing.
30:21You can even think of this in terms of a specific important button, merge. You want to derisk both sides of the merge button. You want it to be more likely that by the time you go to GitHub and you're looking at the merge button that it is safe to hit it because the agent has addressed as many of the potential problems as possible. It's kind of crazy to think of it this way, but I would honestly guess that of my token use, maybe 10 or 15% is actually coding and
30:49the other 80 to 85% and the other 85 to 90% is being spent verifying the code.
30:56So by the time I am actually looking, the work's done and not like it's kind of done but it has these bugs. If you can knock down the likelihood that the PR sucks or has some small issue from 5% to 1%, you'll be able to do way way more because that means you can focus less on each PR. So you want to derisk merge on that side and you also want to aftermerge. It needs to be easier to revert the things that are failing. You need systems that catch things before
31:25they hit your users or systems that will revert things after they hit their users and you've noticed the problem. If undoing a bad change takes more than 15 seconds, you probably shouldn't be vibe coding at all yet. You probably need to fix your systems. So, undoing something broken is way cheaper and faster. Then you can go a little harder. Holy [ __ ] I think I just saw the actual worst take of all time in my chat. The argument that I'm making is essentially you need to watch as much Netflix as you can because you need to maximize your
31:54subscription. I I'm sorry. I we might be using a different Netflix. I I just I personally don't know how I would use my unlimited movie watching to build real businesses. I don't see how I would use that to save 40x plus. I don't see how Netflix can help you escape the permanent underclass. I've actually never I hope you're rage baiting because this is legitimately the worst take I think I've ever seen and it's my job to
32:22read shitty takes which means congrats.
32:25You can probably provide a lot of value to the world through Netflix because you've successfully baited me in the middle of what will be a very good video by sending the actual stupidest possible [ __ ] message. So seriously, congrats.
32:39Salute. Hats off. Fantastic work. Back to real world. So, we've talked about d-risking before merge by making the code more likely to be good. And we talked about de-risking after as well.
32:51So, if the code is bad, it's easy to fix. Whether or not you are using agents, these are things worth doing.
32:57Make it easier to verify code is good before you look at the PR and make it easy to revert if the code is bad. But there is one other phase here, which is not really a button, which means this isn't the right UI, but whatever. You get the idea. Knowing what you want.
33:12There's a very good chance if you're watching this video and you're using agents for coding that by the time you have sent the prompt to the thread, you probably know what you want. I'm saying this because this is the case for me even just a few weeks ago. That is not the case anymore. I have finally rewired my brain where I don't bring in the agent when I'm done thinking. I bring in the agent when I start thinking so we can talk it out. Maybe we find a shortcut that I hadn't thought about before. I've had times where I thought
33:41about a problem, not like fully, but like pretty actively for 3 weeks, and then I went to an agent to build it, and I just asked, "Is there a stupid simple solution I'm not thinking of here?" And it gave me one. And I realized I had just wasted 3 weeks of my time. And I want to be really clear here because after the video about knowing your codebase, I realize you guys don't listen very well. So, I'm putting this in largely to have a quote that I can grab when people misquote me from this video. I am not saying you should stop thinking. I am saying you should figure
34:10out if something's worth thinking about before you think about it because you don't know until you put the time in thinking about it. But if the agent can solve it before you have to think about it, you both just saved a bunch of time and tokens too. Because if I ask the agent about a problem and the agent has a good solution, then I don't have to trick it with my bad solution and run in circles a whole bunch with it. You save your tokens and your brain if you give the agent the problem instead of the solution. And this was a hard habit for
34:38me to break. When people would DM me a bug in T3 Code, for example, I would think through the bug. I would think through where the problem was and I would go to my agent with a solution. I would tell it, I want you to change these things in this way and then it would change it and I would look at it and I would realize it doesn't actually quite solve the problem I wanted to. So now I just hand the screenshot to the agent and say, "Fix it." And it often does. And if it fails to, I now know this problem is too hard for the agent
35:06to do itself. At which point, I know it's time to think more about it. I think it's been really cool to learn that a lot of the problems I would have used a bunch of my mental energy on didn't need it. And also, and this is even cooler, some of the problems I thought were really simple weren't. So, the agent outright failed. You ready for the spicy take here? This is similar to how it felt to be a manager when I realized that my team doesn't need to be handed solutions. is they used to be handed interesting problems and they would usually come back with solutions
35:36and if the first few times they come back with a solution it's not quite right then I know this problem is novel and difficult in some weird way and I have to dive in and help steer it a bit.
35:44So instead of diving in once you know the solution dive in when you discover the problem see if it can solve it autonomously and if it can't whatever buy another cloud account or get your boss to. Seriously though if you're not employed as a dev don't stack subs until you're making money. Like, I know this shit's expensive, but devs make a lot of money. And the companies hiring devs also make a lot of money. Burn the money when you have it, burn somebody else's when you don't. So, the core thing I'm trying to say here is that at each of
36:11these stages from problem to knowing what you want to do to the code being written and merged, too many of y'all live not even this whole range, but like between a small set here where you are just letting the agent operate between knowing what you want and PR filed. I want you to do this. Go all the way from where you first hear about the problem to when you hit merge and maybe even let the agent merge itself once you build
36:39more confidence. And now your job is talking to users and of course S for when things do inevitably fail. The time thinking about this changes you in important ways. So again Dan, I respect you heavily. you were fantastic to work with. But I'll push back on this specifically because I again I even gave myself this quote earlier to make sure I have my get out of jail free card. I'm not saying think less and I personally
37:07believe it is a better use of our time to think about problems agents can't solve than the ones agents can. You will learn more and better yourself more thinking about the problems that your agents can't solve or reading through the solutions that they can solve than you would get thinking about a problem for days that an agent can solve in minutes. I'm not saying that you should replace your thinking with the agent.
37:30I'm saying that you should optimize for thinking about things that matter more.
37:33Don't outsource thinking, scale it. Yes, exactly. Pull in the model to figure out if you need to think about the thing.
37:39And if you do, that's where your brain energy goes. And we're going to get to this tip in a bit, but part of the skill here isn't to give it the task, let it go do the thing while you go and play a video game. Since the agent's doing the task, it's time for you to do the next task, and then the next one, and then the next one, and then you see the first one dung, so you go back and check it.
37:58And that's where this starts to get really cool. If you let the agent run for this whole path, this can take hours. It is possible that from when you show it the problem to when the PR is ready to go, the agent has to run for a couple hours. Maybe it runs for 30 minutes and it tests the thing quick. It gets up a poll request. It waits 15 minutes for getting, I don't know, like a review from any of our awesome AI code review sponsors, and then it spends 10 minutes fixing the changes, puts it up again, waits another 15 for a follow-up
38:26review, addresses those, and now it's good to go. That's an hour plus that you could be spending sending more prompts to other threads to do more things.
38:35Here's where that ends because I'm going to be so [ __ ] real. I have been very kind and polite to the alternatives to T3 code, but they [ __ ] suck when you have more than three threads going. And here is where we need to talk about a very good question we just got from Straw Man Twitch. How do you keep them all straight in your head? I'm going to be so real with you. If you're not using T3 code, I do not know how you do it. I I hate that this is the case because the
39:04competitors have been trying to copy us and they are not succeeding and it is really genuinely frustrating.
39:11I want the codeex activity thing to be good. I have offered to go there for free and write the code for them because I want to use these things well more than I want us to win with our open source project that makes no money. So what the hell am I talking about this arrogantly? I'm not actually arrogant about this. I'm pissed off about this because the solution was really simple.
39:31It was admittedly one that I thought about for far too long because I was trying to figure out how do I deal with the fact that I have, let's be real, far too many threads at any given time. I realized that I'm not treating threads properly. Threads are not histories that you actively go back to all the time.
39:48Threads are tasks. They're to-dos. And when they are not working, they should not matter. I have a bunch of stuff here cuz I've been streaming. So all of these things are done. Also, I was at demo day yesterday, but normally at any given time, I got 10 plus things running and five plus and the done state. And this is where things get really really cool with T3 code specifically. When you are done with a thread, you check settle and now it is gone. I know, I know Theo
40:18thinks he's so cool for giving a new word to archive. This is not that.
40:23Settle is different. The goal here is to turn your sidebar into inbox zero. You should try to end your day with all your threads gone or running. So when you go to bed and wake up the next day, you have some cool things to look at. But that doesn't mean we have cured the context switching costs. We haven't. We have just made it easy to see which context needs your attention. This is where you have to start rewiring a bit.
40:47This is where if you have ADHD, you have a solid advantage because once you get in the mindset of thread is up, not my problem now. You're not going to get out of it. That's something I want to add to T3 code. I'm going to be so real. T3 code is progressed so much and added so many of the things I wanted. I don't really think about it that much anymore.
41:06Like I I don't I'm running out of things to add, which is great. It means we're doing very well. Once orchestrator V2 is in, that will change. But let's say I want some things. And I do. I have a couple in mind. Here's the first one. I would really like a simple and minimal queuing system where I can send a message and it will show as pending and it will go up when the next tool call is completed or I can click steer and immediately send it. The attach screenshot shows an example of what I'm thinking of here from another app. Since I don't have a good way to get the
41:34screenshot cuz my codecs don't reset for another 3 hours or so, you should be able to find a good example of this in codecs. Look for screenshots online if necessary. I have two tips I'm about to give you that are important. They're going to be rapid fire, so pay attention right now. Tip one, look at the computer I have selected on the bottom left. This is one of the things I think T3 Code does exceptionally. Right now, it's Theo's MacBook Pro. That is the computer I'm currently using. This computer is
42:03going to have to be closed when I go downstairs later. I don't know if this thread will be done by then. So, I am not going to run this on this computer because I don't run anything on this computer other than my fleet management cuz I'm doing that in the loop. So, we're going to click here. This could be any repo and as long as the origin forget is the same on the different machines, they will all be bunched under here. So, I can pick any of my servers and pick where I want this to run.
42:27Usually, I pick one of these three cuz they're my Linux boxes. But, if I really want this to work on mobile or do computer use, which right now is better in Mac OS, I can pick my two Macs at the bottom here to do it. I don't care for this one, so I'm just going to throw it on a random cloud server. Cool. Now it's going. And I just realized I missed the second tip. So I will give that after showing another thing that I want to fix. If I hop in here to connections and scroll down, you'll see all the devices I currently have connected. I hate the UI here because you have update or
42:56disconnect and no remove for my T3 Connect machines. And disconnect doesn't seem like it is or isn't permanent. So it's not very clear. So here's what I'm going to do. Going to grab a screenshot of the whole UI here. Screenshot grabbed. We're going to go back. Command shift O. Enter. New thread. Cool. Paste.
43:13I really want to rethink the UI for the devices connected in T3 Connect in settings on the desktop app. There's a couple issues I have. First, I feel like disconnect doesn't seem as temporary as it is. It would be nice if that was a toggle instead. I also don't like that when update is an option, remove isn't.
43:32I should always be able to remove and it should be very clear that it's a permanent destructive action. If you think this is simple enough to do directly, go do it and send me a screenshot. If you feel like you aren't quite sure, make me a few mocks using my HTML skill and we can decide between them. Cool. There's a couple things I did here that might be useful. First, I gave it the exact issue I have with not a prescribed solution, but the problem and then some ideas of solutions. I then
44:00told it that it can go do it directly if it has a solution it's happy with. And I also told it that if it doesn't to give me mocks using a skill I built so that we can make a better decision together.
44:11And now if it does make one solution I don't like it, I can tell it again like go make the mocks and it knows what I mean. And here is where one of my favorite T3 code pro tips comes in. I'm not going to press enter. I'm going to press command enter which sends off the thread in the background and leaves me exactly where I am. So, I can now send off another prompt without having to do anything. It's very nice. And this helps you get into the mindset of, oh, that's a problem. I will go fire it off and I will look when I am done. I do not check my threads until they say done or input
44:40in the corner. And if you're too lazy to even pick which server to run, Maria has built an awesome feature for you in settings here, load balancing. You can now set up auto load balancing across all of the servers that you have connected in T3 Code. So it will fire off the threads across your different machines. And once you work this way, your job becomes different.
45:04You're no longer sitting there carefully babysitting the exact change. You are firing off different things you want to have done and then you go through them when they are ready for your attention.
45:14And here is the harshest reality. I need you guys to get through your thick [ __ ] skulls. And this was so hard for me. That's why I'm being mean. It took me forever to accept this. Agents can be multi-threaded. Humans are single threaded. You cannot focus on two things at once. You cannot do it. This means our jobs a bit different now. We are the bottleneck. So, how can you get the agents to unblock themselves as much as possible so you only have to come in when they need you? And how do you make it so they need you less and that you
45:43come in later? So, look at this one that I filed because people were mad about how I think about streaming. I wanted to move to a chunked by paragraph solution potentially. So I sent a prompt asking specifically how hard would it be to do this. If I really was committed and wanted this, I would have just told it to go do it. But I wanted to see if there's any difficulty here before doing it. Cuz if it's any friction, I just don't want this feature. So I asked and it said not hard. And then it hallucinated about half a day all in one. No, I don't [ __ ] care about what you think half a day of work is. It's
46:12kind of funny that these models still don't know how long work takes. Yeah, eventually it'll be fixed. Regardless, it said it would be pretty easy. It set a bunch of variable names that seemed fine. I then said, "Can you build it for me once you get it working? File a PR and spit up an environment with tail scale for me to try." Now, and this is very important. Imagine I didn't do the second sentence here. Imagine I just said, "Can you build it for me?" And then in 5 minutes it comes back, okay, I built it. And then I'm like, "Okay, can you spit it up on tail scale so I can
46:41try it quick?" I leave. And then I get the ding. I go back and then I click the link and I try. I'm like, "Okay, this is great. Can you file the PR?" And then it does. And I look, I'm like, "Oh, cool.
46:50You should babysit this, too. I probably should have included babysit here as well." So, I'll do that now. Can you babysit this PR and make sure everything is good? Cool. Now, it will continuously monitor that PR and it's not my problem again. So, your goal here is to make sure everything you need is there the next time you click. Ideally, you want to maximize the chance that the next time you check that thread, you're ready to merge. And really think about that.
47:15What can you add? What context can you give? What tools can you let your agent use to make it more likely that by the time you go back to that thread, you can merge the PR. And this is where one of my favorite T3 code features comes in.
47:28When you merge the PR, the thread disappears. Which means if you tell the agent you can merge the PR if it passes these conditions or meets these requirements, then you can send the prompt and never see the thread again. I would honestly guess that around half my threads are archived without me ever seeing their final message because it doesn't [ __ ] matter. I see people confused about the agents stepping on each other's toes while they're working.
47:52I thought we were developers, guys.
47:54We're not vibe coders. We made the right default in T3 code for this for a reason. By default, [clears throat] when you start a new thread, it starts in a new work tree. Not going to pretend work trees are perfect, but they are good enough. And now that the models are smart enough to deal with the weird [ __ ] that is git, they will fix your work tree issues for you. For example, one I run into a lot is I have a work tree with a branch that makes a pull request and then I want another agent, usually another model, to review it and
48:22give thoughts. And if those both are on the same machine and they're both on work trees, they can't both have the same branch and then git freaks out.
48:30Previously, models were dumb enough they would get stuck there. Now they can figure it out. They have workarounds.
48:34They'll make a new branch. They'll like pull it down in some other way. They'll make a clone. They'll do whatever they have to. It doesn't matter. I don't look. I don't care. Models are smart enough that if I give it a pull request with a branch that it already has in a work tree, it'll figure out how to read it. Don't spend your time thinking about those things anymore. They don't matter anymore. This is another habit I've gotten into. By default, T3 code auto settles threads that you have not touched for at least 3 days, which I think's the right call. I bumped it to seven because I usually have better discipline about clearing out my inbox,
49:03but I'm good enough at it here, which is why my sidebar is a little chaotic. I was exploring thread pop outs. This one I actually do want to play with later, but not now. So, I'm going to use the snooze feature to make this my problem later. You'll start to see my mindset as I go through this. I want to make it so I don't have as much stuff trying to take my focus. This is why I've done some strategic things with the UI and the sidebar. Notice that the ones that are working are semi-transparent.
49:31I wanted to hide working threads, but nobody would let me. So, I made them more transparent so you don't look as closely at them. I like this a lot. It has made it much easier for me to only prioritize things that need me, things that are done or things that are waiting for input. I also really, really, really want to emphasize there are few worse uses of your time than watching an agent as it works. If it works and it succeeds, awesome. Merge the code. If it works and it fails, maybe read the
49:59reasoning trace or even better, ask the agent why. And this is a good opportunity to pivot into the next section here, which is using your tokens poorly. I know it sounds silly, but I'm actually going to be framing these things in ways that could genuinely be helpful. God, I one last thing I want to crash out at a little bit. You can censor the chatter if you prefer, Jeff, in the edit. I see a lot of sentiment like this still.
50:24Like, if it works, that's fine, but what if it installs some malware on the side?
50:29I need to be so real with y'all. If you think your agent is more likely to install malware than you are, then you are dumber than your agent. If you are smart enough to realize that won't happen, then you're both smart enough to not install malware. If you think the agent is more likely than you are, you have installed malware recently and you probably have some on your machine. you should consider a new Windows install because I know you're also using Windows. Let's be real here. Anyways, we're going to go back to talking to real engineers cuz that's what we're here for. Let's talk more about using
50:59tokens poorly. One of the things I want you to think about is when you face a problem or have a question or uncertainty about something, how can you use your tokens instead of your brain?
51:10When I have a poll request that is hard to parse that a teammate filed, how can I use tokens to figure out what it is?
51:18If I'm curious what's changed in the Orchestrator V2 rewrite that Julius is working on, I could ask him, but he's busy. He's prompting. I could rather just ask my agent to get me the info on what he has been working on. If I've lost track of all my poll requests and I don't know what I should focus on today, I ask my agent to go through them and find something that I should be more focused on. If I can't find an email for one of my five email inboxes, I open chatg and tell it to go find it and it does. If I don't want to sit there and
51:46download all my medical records for my surgeries, I don't ask my assistant to do it like I used to, I ask Codeex to go through my inbox and go through all of my dashboards for my medical [ __ ] on my browser and get it all for me. And what's even more fun is I would have my computer going through and spending 30 plus minutes collecting all my medical records. And while it does that, I go back to T3 Code and I fire off four more prompts for things that I notice that are annoying me or for a feature I want to iterate on. And then every couple minutes I go back, I take a quick look
52:15and I see, oh, these things are done.
52:17These things are still working. Cool. I really need to finish the durable objects up. It's going to save us a shitload of money if I can get it right.
52:22This is how I think about it. The issue with the other tools that you can use for this is that the sidebar doesn't help you prioritize the work that you're doing and the work that is done. It doesn't make completed work go away. The search isn't trustworthy enough to get old work back. And you can't use the same sidebar across multiple machines.
52:42Like of the threads here, this first one is on one of my MacBooks. This one is on my server Alvin. This one's also on Alvin cuz I just spawned it. This one was too cuz I just spawned it. This one's on BB1. This one's on my MacBook.
52:55This one's on this MacBook. This one's on the other MacBook. This one's on my main cloud server. They're all in different boxes and it doesn't matter.
53:02They're all on the same UI and they're all on my phone app, too. I will say that right now, sadly, if you're not using tail scale, you'll only be able to connect three devices at once with T3 code for free by default. We're working on it. I want to bump the number a bunch, but since you're already going to be using CLI proxy, that means you're also going to be using Tailscale. That means you don't even need to use T3 Connect. You can just connect directly.
53:24This all works without using our servers at all. T3 Code is fully free and open source. It fully supports your subscriptions. It works incredibly with CLI proxy. So, if you have a couple boxes that have cloud code and codecs and you have them CLI proxied, you go to one of those boxes here. I'll show you just how hard it is to set up T3 code on a new box. I don't have a new box to set it up on. All my boxes are configured.
53:44Let's say they weren't. You go to the box, you shin, you're on npx t3 connect, and then you click the link. Oh, [ __ ] Email leak. Great. I should work on that. But yeah, you click the link that comes up. You sign in. And now, as long as you're signed in on the client, on the website, the mobile app. And then once you're signed in on the mobile app, the website, or the desktop app, you can now control that machine. You don't have to install anything. You just need quadcoder codecs, ideally with a sub or
54:12a proxy. You run the T3 connect command, and now you can connect through our layer. Or you can do npxt3 serve, and it will instead let you use tail scale. We even have a d-tail scale built in directly. or crazy thought, I don't know, kind of risky here, but we are in the use your tokens poorly section. You can just tell your agent to go set it up. You have a new server and you're already building a repo that manages your servers. You can tell cloud or codeex when you set up that server with cloud and codeex, you should also
54:42probably set it up with T3 code so I can connect remotely over tails scale and it will just do it and it will just work and it will be just great. I'm going to frame the stupid prompts a little silly.
54:51Have you ever found yourself Google searching where did I leave my keys because you got so in the habit of Google searching things that you just default to that? If you don't find yourself doing that with agents, you're not offloading to them enough. You need to get to the point where you ask a thing to agents that it obviously can't know and you feel silly for asking. If you don't find the things you're asking the agents a little bit silly, then you're not asking it silly enough [ __ ] and you're not pushing the limits here
55:19yet. Here's a fun example. I have a lot of side projects and I had a bit of inference to burn the day I sent this.
55:26So, I asked, I want you to go through all my GitHub projects as well as unfinished work on this machine using lots of sub aents to help me figure out what I should be putting more time into.
55:34What are some of my ideas and side projects that seem like they're more in demand now that I might have forgotten or that are worth bringing back and finishing? They might be halfbaked or projects that have been abandoned for a while and deserve another pass. go through all my work on this machine in GitHub and find things I should potentially revive. This is a shitty prompt for a shitty problem. None of this matters, but I feel bad when my limits reset and they weren't at zero.
55:57So, I was looking for things to burn them on so that I could have them all dead and not feel bad when they reset.
56:03And it found a handful of my things that it thought I should prioritize more. I think that is fun and cool. And I have already decided what I want to prioritize. So, now the next step, snooze. Well, settle. I don't need to snooze cuz I don't want it back. Context usage breakdown. I decided what I want to do on this one and what I want to do is kill it. So, we're going to go here and I'm going to close it and I'm going to archive it cuz I don't care anymore.
56:22Show work creation progress. I don't know how far we got with this. How far along is this work? Is the tail scale dev server still up? Spin it up if you can. I'd also love for you to file a PR and babysit it to make sure everything is good to go. And I'll come back later.
56:37Here's one that I've forgotten about cuz it was a while ago. It says 3 days, but I'm pretty sure I started this way before then. So, I'm just going to ask.
56:43I'll be so real. I haven't kept up with this workflow in a bit. I have no idea what the state of this is or what the value is. Can you give me a rough idea of where this PR is at and why we should merge it. This is going to fail cuz I'm still mostly out of Astra. It might route correctly. I have a weird routing issue right now. We'll figure it out.
57:00These two are still going. That means that they are legitimate prompts doing legitimate stuff. Awesome. I will come back to them later when they say done. I would not have ever clicked these two if I wasn't making content because I don't care when they are still working. Here's a real example of what my sidebar looks like when I'm actually working. Do you see how many threads I have here? There were even more below. 1 2 3 4 5 6 7 8 9 10 and a monitor. There's at least like two or three more underneath that. And
57:29across my five cloud code subs, I was able to do all of this and still have some usage left over. Not bad. And after this, I went to bed and I woke up the next day. I went through them one at a time. made a decision. Do I want to merge this? Do I want to do a follow-up?
57:45Do I want to get rid of it? What do I want to do? I went through and one at a time, did all that. And after I noticed some other things I wanted to fix. So, I spun up more threads to fix them. And I went back and looked to see if everything was done and went through and clicked all the ones that were done and did it. I really have been treating this like an email inbox now. And it has made me so much more productive. So, I just said they're confused because those threads have only been running or those threads only ran for 20 minutes.
58:09I had spun them all up within the last 20 minutes. Some of them took 30 minutes. Some of them took 18 hours.
58:16That was me trying to get everything going before bed. That wasn't work going. Those all said working. I had 12 plus threads that had been working for at least 20 minutes. Most of them kept working after. All of them actually did.
58:30I don't think any of them were done even with the next 10 minutes. So, as silly as the using your tokens poorly section may have been, I do want you to get into that mindset. Especially when you notice a limit's coming up and you haven't burned it. Find silly things to have your agents do. You'll learn things from that. You'll learn a lot more than you expect from that. I am still amazed at how good agents are at triaging large numbers of PRs, at finding things that I should bring in in merge that I might have missed or forgotten about otherwise. You'll be amazed at how much
58:58random [ __ ] you can find. And here is where we get into the last section. I've touched on a good bit of this here, but I want to really, really emphasize this part. using your tokens while sleeping.
59:08I mean this both literally and metaphorically. If you've ever had the feeling of I want to fix this right now, but I have a meeting coming up or I have to leave the office soon, so I'm not going to send this message to this thread because I have to close my laptop. I get it. I was like this not long ago. You need to move your dev work remote. You need to pick up a crappy little Linux box. You need to take some old computer that you haven't used in a while. Install Ubuntu on it, set it up
59:36with Tailscale, install CLI proxy on it so that it's always coming from your residential IP and then use it as a box that your agents can work with. Now connect it with something like T3 Code and you can fire [ __ ] off on that box and not have to think about it. I see people asking if they don't have a computer, what's the most budget option?
59:56The most budget option by far is to talk to your friends and family and find somebody with an old laptop or desktop that they'll give you for free that has at least 8 gigs of RAM and at least one or at least four cores and you'll be able to do some real work there. And I'm already seeing some very very stupid replies. If you think a VPS is the solution here, I might have to get into selling VPS's cuz those sound much more lucrative than bridges nowadays. need 32
1:00:24gigs of RAM and 16 threads. You can get it on HNER is only $275 a month. That's a great deal. Or you can waste all your money spending $700 on a box that has 32 gigs of RAM and a 1 TBTE drive. That would be terrible.
1:00:42That's such a waste of money. That's like two and a half whole months of renting a worse computer on an IP address that'll get you banned from Claude. Why would you ever buy hardware that you can run on an IP address that won't get you banned when you can rent something for two months that will get you banned? Hopefully, you understand sarcasm because there is pretty much no reason to do cloud servers unless you're grandfathered into a good deal. Theo, you're wrong here because insert something stupid. It's actually very funny someone said that in chat right after somebody complained about
1:01:11electricity costs for a computer that pulls conservatively 60 to 80 watts of power at worst. And remember, it's not doing inference. It's doing code. It's going to be using one or two threads most of the time. Let's go hop into my massively overspeced server to take a look. I have six threads or so running on Alvin right now. And let's take a look at BTOP quick. Huh. You can't see because my face is covering it. This is a box I
1:01:40have at least six threads running on right now. It's got 32 cores because I got it for a good deal. You might have noticed it's not using them very much.
1:01:48It's rounding to 0%. So once again, you don't need a lot, you just need enough.
1:01:55Sadly, that box I shared earlier with the GMK tech has gone up in price because my tweet caused them to sell out. You can still hunt, find something, but I would highly recommend you do not do a Mac for this because Mac OS is a [ __ ] show for parallel work. If you are looking at the numbers I'm sharing here and saying, "That's not right. I run three agents in codecs and my laptop's overheating." You're right. Your MacBook is overheating because your MacBook has a [ __ ] file system and an even shittier security policy that is hard
1:02:24bottlenecking how much [ __ ] you can run at once. Move to a real OS like Linux and you won't have these problems. You don't need a lot. And I bet my ass if you can find your mom or uncle's old PC that's got four to eight gigs of RAM in it and an okayish chip and you flash Buntu on it and you plug it into your router somewhere, you're going to have a better experience than you have doing dev work on a Mac like immediately. And if you live in an area that isn't San Francisco, well, let's be real. If you live in San Francisco, hopefully you can afford compute. And if you can't get
1:02:53out, the city's too expensive. So if you live anywhere else, Facebook Marketplace will be your friend. I bet you can find a surprisingly good deal. And any old laptop with any real RAM, you're going to be fine. And once you make that change, suddenly you're not bottlenecked the same way. There will obviously be edges like if your agents run CI on the machine locally and it's a Rust project that saturates all your cores when it compiles. You'll run into problems.
1:03:20You're an engineer though and you have tokens. Burn your brain and your tokens to solve the problem. Maybe you move CI to only run on GitHub or you use one of our awesome partners like Blacksmith or Depot. Maybe you have a different server that runs the CI. Maybe you set up a queuing system where the CI gets triggered one after another so you don't have five Rust compiles destroying your machine at once. Your engineers, you can solve those problems. I don't want to just tell you all the solutions to those things because your problems will be different from mine. Apparently, I succeeded with my goal of moving Maria
1:03:49to the Blacksmith CLI in order to let her run the CI via CLI instead of it running on her machines directly because she was a bit RAM constrained. And now she's way less RAM constrained. And that's how we end up with posts like this from Steven that Maria shared.
1:04:03Request, can you make T3 code less productive and experience so that he burns fewer tokens? This is where you want to be. And I have one last thing I want to lean into with this bit. This shit's so fun. I know a lot of engineers have felt like the thing they love died and that everything has changed too much and now engineering isn't fun anymore because you send off a prompt and then you sit there and watch it make a bunch of mistakes and then get annoyed that the code doesn't work and you could have wrote it faster yourself. I still have
1:04:32that experience. The only difference is I don't watch the thread. I go do something else and then another thing and then another thing. Maybe I spin up three threads and then I go to dinner with my friends. Maybe I spin up seven threads while I am waiting in the lobby to queue in a game. Maybe I go skate while I have some things running. And when I'm sitting and recharging, I pull out my phone and check on two of them quick and kick them off to go do a bit more work. And the fun isn't in those parts. I'll be clear. The fun isn't that I'm checking my phone and sending prompts. The fun is this new type of
1:05:01engineering problem. I can now justify spending more time micro optimizing how and when I trigger CI. I'm thinking more about my cores and how they're being used. I'm thinking more about how I can work on the same thing eight times on one box and not have to think as much after. And I love doing this. I find it so genuinely fun. And I hope this video helps inspire some more of this fun for y'all because that's my real goal here.
1:05:27All of this chaos has made engineering the most fun I've ever had with it. And if you're not having fun yet, I hope this can help you get there. And maybe, just maybe, you have a reason to go spin up a couple more accounts and burn a few more tokens. And if you do want to throw some of those tokens our way to make some real improvements to T3 code, we'll probably ignore them, but we might merge them. So, consider it. Oh, yeah. I did have this DM back and forth with Maria.
1:05:52I could probably scroll and find it, but I don't want to scroll through our DMs publicly. I set her up to do remote stuff, and she got a Linux box she could do at her place. She hits me up in all caps. Remote work is the best. This is so fun. And I said, "I'm sorry. You're about to burn so many tokens." And if you've been struggling to hit your limits, I really hope you don't struggle anymore because this [ __ ] is so so fun.
1:06:15I think I've said all I have to on this one. I'm going to go kick off some more threads and grab another drink. Enjoy the new world of agent maxing and get as many of these tokens out as you can before the subsidization ends, if it ever does. Let me know if you want a video about that, too, cuz I have a lot of thoughts about how this goes long term. But this wasn't about that. This is about actually using them. And I hope it was helpful. Let me know.