0:00A new Pi release. also could you please introduce yourselves?
0:05I mean Armin, I've never spoken to you yet.
0:08I've had Mario on twice. but yeah, could you do an intro?
0:13Hi, my name is Armin. I'm the MCPO at Earendil and resident intern.
0:18yeah, I don't know. We're we're we're just here working on Pi.
0:23but yeah. Well, When did you join? Like like working on Pi?
0:26I founded a company with calling a friend of mine, and we sort of subsumed Pi in in Nov in February.
0:39Why did you decide to do Yeah.
0:41That is an excellent question.
0:43no, it's like we we tried to get Mario on board for an extended period of time prior to that. it's just like in the in the wake of the OpenClaw craziness, the opportunity came that Mario was actually receptive to that idea. That's sort of my retelling of the story.
1:00but yeah, because like we wanted to work with Mario and we're like malleable software and and Pi was a is a great version of of an early experiment of that. so it just made sense.
1:16And now that you're so go ahead please Mario.
1:16and No, I just wanna say I'm Mario, I built Pi, I'm an idiot for building Pi and now I'm stuck with farming.
1:26Do do you feel like you're stuck with the project for real?
1:28Like is that a a concern of yours?
1:31Like it's so successful and what doesn't make sense to s to stop it?
1:37Well, I I mean I have like Armin, I have large open source history in a different field than Armin, but still. And once I do a open source project that gets users, as I said, I'm an idiot and keep maintaining shit for people, Yes.
1:50even though they don't pay me a single cent.
1:52And the negativity that comes in at times can grade you and but yeah, I think the goal is still to hand Pi not hand it over, but have more people Work on Pi, so the truck factor of me n needing to be around is less.
2:12And I think we have achieved this goal pretty pretty well so far.
2:16So I if I die tomorrow Pi will be in good hands.
2:19That's my That's that's a good place to be.
2:22Yeah. I started a meetup in Poland.
2:25when I moved here I didn't have any friends and the meetup just like exploded.
2:28It's I think it's the largest meetup in the country right now, on the meetup.com app. And now it just runs itself.
2:34I don't organize it, I just go there.
2:35I'm on the you know, my badge is there, so everybody's nice to me. But it's it's so beautiful when a project gets there.
2:41How how how Yeah, that's nice.
2:43has yeah, like finding contributors, how's that changed? Like more and more people are using AI.
2:49I know you were not a big fan of just like slopping things out.
2:52So how how's that experience been like?
3:00How Mario, how is that experience?
3:03So let me put it that way, Pi Durable would have happened a lot earlier if I didn't spend February to probably July spending most of my day maintaining Pi, which means going through issues, fixing issues. and the the the problem is agents can't really help you with fixing issues yet, at least in my opinion, at least for my project.
3:31I lack the trust to hand this over to them and judge which issue that's being reported is actually worth our while and needs fixing and which doesn't.
3:39So at the minimum I still need to go through all the submitted issues and triage them and select what's actually worth our time.
3:47And then the question is, can an agent fix issues?
3:51Yes, obviously. The question is just does it fix them well?
3:54And I would say for a lot of issues that works.
3:57I can trust the agent to fix it, especially if it's a tiny little thing.
4:01for anything that concerns the design or architecture you still kinda need a human because the agent will just locally shit into your code base and make things worse. Right Armin?
4:11I think the the the problem is also with the agent is like it will definitely solve the problem, but it's almost it's almost impossible that it will fix a problem without adding quote.
4:20Which seems weird because I think like like a lot of the issues actually shouldn't be net additive to the code base, but I know, I'm like I know, says the guy who permanently sends PRs with two hundred file changes it's ho it's it's horrifying. Like I'm I'm and twenty thousand lines of code.
4:36a I'm a victim to my own clanker use, but it like it is it is It is still not how things should be.
4:44and like particularly if you have like fully automated code base, I think it's just naturally going to grow with every single bug fix.
4:53So I had a conversation recently with DH who has given into the whole like slop you know slop operations. That that's w what what I call my job 'cause I basically use the models for everything.
5:05And on some level I feel like I'm definitely much dumber than I used to be.
5:09But on another level it's like very productive, in terms of if your job isn't like building something that is going to be replicated and used by millions of people, I don't see much of a downside of just like slopping out a script or Like create it so like what I've been doing recently is creating Docker containers for specific hardware combinations to get like a model running optimally on this like whatever random configuration of VRAM and RAM and and VME.
5:36And that's been working great, but it's just creating a Docker container.
5:39So like the thing, the design, the architecture doesn't matter that much as long as the thing, you know, like the the result works.
5:45Does it does what it's supposed to do, But you guys don't have that, you don't have that privilege.
5:46yeah. No, I agree. Sure we have that.
5:50we obviously have that. Lots of lots of things in Pi are actually pure slop.
5:55it it it just depends on what layer of the application y you you you pull the human in and on which layer you trust the agent to produce something that works and you don't care how it works.
6:07Like I give you an example, the TUI library of of of of of Pi.
6:13the initial design was At least love.
6:15yeah, but the initial design I I hand wrote, right? Back in early 2025 for another project.
6:22and that led to a pretty rob robust, minimal kind of thing that I knew would probably survive contact with the real world, even if it's not ideal, but good enough, right? And then after that, when I started building out Pi, I didn't go in and say, write me a markdown component, write me a text component, write me a editor component.
6:42I just well sorry, that's exactly what I did.
6:46I I just told the agent here's the API, use write against that fucking API and build me all of these components.
6:53I don't fucking care how they look like.
6:55They should just work and have a bunch of tests that ensure that the rendering is correct. Everything else I don't care.
7:01And I to this day I haven't looked at a single line of code of the markdown component or the the text component.
7:07So I I don't fucking care. It doesn't matter.
7:09the only time I do care is if performance goes down, and then usually the agent is great at identifying a bunch of shit, but it's very bad at coming up with a generalized solution.
7:20It will c go in and fix fix something locally that will improve performance for that one edge case, but it will not kinda think big picture and and f identify issues in the architecture itself that make things slow overall and led to this edge case of slowness, for example. So that's that's where I feel I still have value to some degree.
7:40I remember y you told me you like your processes creating like interfaces and like playing around with the APIs and trying to get a sense of what you're building.
7:49has that process changed at all?
7:53'Cause like for me, what I'm seeing right now is like the models have gotten much better in general at like doing things but the I have not seen much of an improvement in terms of like architecture.
8:03It's still very like messy and like have you have you noticed anything different there? Like these 'cause these new models are very capable but Yeah.
8:13I almost feel like they're worse.
8:14I don't know. I d like Nah, they're not for us. No.
8:18so I do think they're a little bit worse in in the architecture part.
8:22because now well so they are worse in the outcome because you give them way more control.
8:31I think this is basically what's happening.
8:33And I like with with almost every single model release, I'm like trying to build something.
8:42And I'm I feel like I'm progressively less happy with what they're doing on their own. And this is mostly because I'm like trusting them so much more.
8:50So like the the actual net outcome.
8:52So they're they're probably still better, but combined with my expectation to the model and how I'm using it.
8:58It's like like I I always build a C game, like a just a game in C. Because like it's just I don't know, it's just a thing. Like I I used to l love writing C, so it's like, okay, this is like a thing that I do.
9:11And I feel like they're getting worse.
9:12Like it's like maybe it's just pay less attention, but I can for sure say that it didn't get any better.
9:17I think the stuff that's coming out of them that it's getting complicated.
9:21Like the way they speak is getting more complicated.
9:23There's like they use just like way more technical terms.
9:26They go into like all these like weird niche technical topics in their responses.
9:32I'm not necessarily saying like the way that they built, but the way they communicate with us is getting weirder, I think.
9:39Although I I I I do think I Opus got better.
9:42yeah, like what was it, Fable five point whatever the fuck and the the previous opus release.
9:48Honestly I I s just stopped using them just because of the way they talked.
9:52That it was unbearable. I I I'm not a galaxy brain, Yeah.
9:55I every word was a new entry in a Thesaurus, so I don't I did I gave up and switched back to Sol completely.
10:04and then Astra, I think, also starting doing that bullshit.
10:09So I switched away from Astra.
10:10Estra is also terrible.
10:11but I'm pretty happy with with Sol 6.1 and Opus 5.5.
10:15They kinda toned it down again, and they're now normal normal people like machines.
10:20So that's good. But I think so one improvement that I've definitely felt is in Cyber on an unrelated project.
10:31last year I did a bunch of things to to secure some software that's desktop side with a clanker and that it just didn't like it couldn't comprehend what what was going on. And this year that software got cracked by some Chinese dude with two codex accounts. And I then used a reverse clankerer to replay the attack and and kind of figure out w how how they could have gotten through the defenses of my software.
11:01And that was extremely fucking impressive.
11:04Like really, really impressive.
11:06No reverser on Earth would have gone that far ever, in especially not in that amount of time.
11:11It was just fucking crazy. Like to give you an idea, that's a little low level, but so it's it's it's a JVM and then the JVM you have Java bytecode and that bytecode eventually gets compiled to to machine code through a cheat compiler.
11:23and I I secured all of that real well so people cannot just simply extract the bytecode or inject their own bytecode.
11:31But there was one little escape patch on the native side, which no human would have ever like been able to identify and exploit.
11:41And the clan just went in and reversed the native structures inside the JVM and started reading process memory directly, mapping that back to the correct build of the JVM, identifying which parts of that structure in native memory maps to which parts in in the Java representation, and it just It was like science fiction and it was amazing.
12:05Yeah. I mean I had the same experience.
12:05So I Like I I have this little I always show this is my little yeah.
12:09USB chip, the this Chinese CarPlay adapter.
12:13And like it f on its own figured out how to decrypt like this little thing on there that was supposed to protect the upload of like a new root image. And it found like a CGI file on the web interface where it can sort of like bypass stuff.
12:29And it's like I didn't do anything.
12:30I just plugged it in and like make it work.
12:33I don't know, like five hundred dollars later, like the whole thing was like mine.
12:36And it's it's impressive. It's really impressive.
12:40So I have them at home, like I I'm buying more and more hardware every or I'm given a lot of it now. And I'm just running these models and trying to see, like from the small smallest ones, the Qwen 27B all the way to the like DeepSeek and GLM.
12:53they've all gotten really good at doing things like I gave them a Nintendo DS ROM hack, like a ROM game, and I had it like build a Nintendo S DS emulator from scratch.
13:08Like it could look at Dolphin and other stuff, but it It has to write it like completely differently and it got it running.
13:14this was DeepSeek, it was able to get it running.
13:16yeah, I've I've found huge increases like w whenever the task is like very binary where it's either done or not done, they will eventually get it done.
13:28That's the kind of conclusion.
13:29Yeah, if if if if you have an Oracle that can tell the LLM if what it did is correct or not, then you basically won. and that that like for example Bun, right? Bun has a massive array of tests.
13:42I I'm sure it's not a hundred percent covering everything that needs to be tested, but it's enough for for Jared to to be able to tell a bunch of agents, okay, let's port this in this and this way, I'll steer you at at points. But mostly it's just try to make the test suit run.
13:59in your Rust re write off button.
14:00That works really great. I did similar the similar thing that you just said with the emulator, there's like a gazillion emulators for all kinds of systems out there and it's also in the training data, like crazy. Including the specs of those systems.
14:14So that is also a thing that is very easily testable with an Oracle.
14:18and and where the training data is just chock full of information, right? the same And it's also decently fuzzable.
14:26I think this is like like anything that can be fussed.
14:29Like I know I had this thing recently.
14:30I took like I took a bunch of like example files of any files, tomo files, CBOR files, message pack files, JSON files. And it's just like, here is my serialization library, just implement all the formats for it.
14:43So it was eventually done. I was like, okay, now fuzz the hell out of it.
14:47And it found a lot of issues with it.
14:49But because like I I know it was already not too hard anyways before to fuzz libraries, but now it's just a no brainer.
14:56It's like okay, make it work until It it it like it takes you hours to find another regression.
15:01Yeah, and I think that's an amazing use case.
15:04And then it it's like especially for fuzzing, like you you might have some automated fuzz generation, but with an LLM on top of that, right? That can also kinda reason about additional cases that your fuzzer that fuzzer generator doesn't cover.
15:18It's like rocket fuel for all for all of the kind of testing.
15:21So super impressive, super useful, really, really great. But then we come to things like you said earlier, like like architecture. And I don't wanna be the humans, not just I think DHH said something like and this you think you're special because Architecture safe.
15:36you can architecture and design and so on, think twice because the models are gonna snuff that as well.
15:41I think DHH ultimately doesn't understand what machine learning does.
15:41The thing is I I think that the models I think the models are actually like if you if you discuss with them what the right architecture for a problem is, they like I think they're pretty pretty good at it.
15:53But like they're like the the the the trajectory that's the most likely outcome for solving a problem goes down somewhere else entirely in the weights.
16:03so it's And it it can all Yes.
16:06it it can basically all be explained by the training data.
16:08I can I can easily set up training data for cyber stuff.
16:11I can easily set up training data for building something that already exists that has an Oracle.
16:17I can easily set up training data for building websites, for front-end stuff, for whatever the fuck you want.
16:23But what I don't have training data for is the process of building, of designing and architecting something.
16:29We that is the part, this is the only part in programming that is actually not usually encoded somewhere in textual form so an LLM can learn from it. That the process of building the architecture of something, and that's why we are stuck here, Hmm. Well I think.
16:41This is why we're all going to record now that our meetings and then every failed sh startup is going to have their meeting Yeah. We we're gonna sell our recordings and everything sent to the model apps.
16:52blackboard for two million dollars to open AI.
16:54S so j in I'm I met this guy in Germany who has a robotics company and they're the their main like business is actually selling training data.
17:04So they pay people to wear the hats and the arms and stuff and you know do the work and then they sell that over to the labs.
17:11And it's very good business. but I'm sure like it's it's all in the traces.
17:17So if if you are talking with the models about architecture, those traces should be weighted heavier.
17:24It's like how do you distribute how do you decide what gets higher weight?
17:28Yeah, it's it's very very interesting problems.
17:31But how It's a it's a I think the problem is also like it there's no signal in it because like as an middle sequence. Yeah.
17:34example, Pi Durable went through I don't know how many dollars.
17:38Yeah, it's like it went to multiple iterations.
17:40Like and and like how would how would the training know necessarily that one of those conversations actually ultimately ended up with the version that we shipped? And and obviously it's sort of possible and you can train it and you can and like so you can tag it, and like there are ways in which you could potentially do it, but it's a significantly harder problem, and there's just l fewer of those conversations that are really I I I would just like to think think that this is basically a classification problem. So you get a bunch of traces, i irrespective of your of your zero retention data policy.
18:10zero data data retention policy.
18:13so you have a bunch of traces as OpenAI or Anthropic or DeepSeek or whoever.
18:17Like a gasillion of those, right?
18:19And now you want to pick out the ones that are actually high value, high quality. How do you do that?
18:24You can't do it y with a human.
18:26I mean they do label stuff manually as well, but that doesn't scale. So then the next Best thing is a machine learning model that can classify traces into high quality, low quality.
18:38I think they have the opportunity, you know, like the thing that they do where if you're using Claude Code it just tags itself in the GitHub repo and the commits.
18:45So by building out like not only the conversational traces but also the the repository history itself, that is a probably a good way to do it because you just like map like what repos are still getting users and which of those PRs have gotten in and have like a percentage of the code still I think I think that makes it so yeah, I've been sharing my data with them.
19:07I d I just clicked the share.
19:08Right, but but the thing is this, like the amount of data they get, you you would think they would be able to do sophisticated shit, right?
19:15But from what I heard through the grapevine, they are not doing any kind of sophisticated shit yet.
19:22and I think once they Yeah.
19:24get a handle on that It's going to be with RSI. The models are going to do the sophisticated stuff because Yeah.
19:29all of a sudden it's going be easy.
19:32Okay, so I have to This is actually I think like why it could get better is that like right now we're sort of on the face of like it's really annoying for humans to do this, but like if the models think that it will be worth the effort and that will push them forward, like maybe because of self improvement will will get us there.
19:50With the latest Opus, like I have all this inference stuff in inference tasks that I give it and it is it is making breakthroughs like on a day-to-day basis.
19:58Just s it's gotten so good at like this art hardware to software layer, like navigating that. GPT is not good at that at all.
20:07Like I don't know that their models have they're better at math and s s some stuff like that, but yeah, I don't know. I I I think I've seen a breakthrough there, but that doesn't really I don't know, y you can't build apps with that.
20:21It's it's just a skill of like hill climbing which is a a little disappointing.
20:26I did wanna like ask you guys about like Pi durable.
20:29So y you you're using the word durable.
20:31What what's the logic yeah, what what's the story here?
20:36I I I try to do it in five sentences, okay? Agent starts up in a process, agent calls tools, process dies.
20:45What you want is restart the process, agent continues exactly from where it left off.
20:51That is durability. Without you having to type continue and without tool results or assistance streams getting lost, with your application state being also persisted durably and and and and and reliably resumed from and that is durability.
21:09And why do we need it? We need it because for long running agents, so anything that's a claw, you need exactly that.
21:16And the other thing that Pi Durable solves is that historically Pi has loaded the entire conversation basically, or the entire tree as it is in Pi, into memory, and that obviously also won't work if you have a Slack bot that lives in your Slack for years and years and there's like a gazillion messages in its full conversation trace.
21:33Not necessarily the thing that you s send to the LLM, but you still need to kind of be able to have all the history.
21:41And that is basically what Pi Durable does.
21:44plus it's easily deployable to a lot of places like Vercel, Cloudflare DO, E2B, your laptop, my phone, which makes it very interesting to work with.
21:57Maybe you can walk me through a little bit of like the architecture.
22:00So I I did read the repo today and I had a clanker also try to explain to me what's going on. the resumability I kind of understand is like you have to you have to commit an action if you this is this is what I understand at least. you have like an action and that action So give me needs to get committed into the file, f for it to be like yeah, for f for that to be canon. and if it drops like you have like this event queue.
22:24So you have events within a queue, at least That's what I think I understand.
22:28Yeah, that that that's the thing that we don't do. We don't do event sourcing.
22:32That is what everybody else is doing, Uh-huh.
22:34and we don't do that because there is a lot of problems.
22:36But instead of let me explain how we do it, okay? So the basic problem of durability is that you are executing something, it can break either before you start executing it, while you are executing it, or after you've executed it, right? And the thing that you execute is commonly referred to as an effect.
22:55So Most durability solutions then do the following.
22:58Before I execute the thing, I store something in some storage that says I'm about to execute this thing with this parameter so whatever inputs, right? So if I crash after storing that, I know that I've already evaluated all my inputs correctly.
23:15I verified them and I'm ready to execute the thing.
23:17If my process dies now before I executed the thing but already stored this intent, I can restart.
23:24Look into my database and see there was a thing that's about to happen with these inputs. Let me just do this because it hasn't executed yet.
23:31So I'm doing it now, right? That's the first part of the sandwich.
23:35The second part where you can crash is when the effect the execute is actually running. And then it depends on what is running actually.
23:44If it's an external system, like you send to GitHub Actions, built this commit in this action, right? And you send that.
23:53And then your process dies, and then you respawn, then you don't know necessarily if this has actually executed or not, because you haven't saved anything yet, because you didn't get the result back.
24:05You you were interrupted while you were talking with GitHub, right? so in that case, how do you continue?
24:12With GitHub is easy because you can say I'm about to schedule a run before it runs, and you get a key.
24:20and then you can save that before you actually execute that run, right? So an ID. so if I then execute the thing, I say execute with this ID, right?
24:31And if it already runs, GitHub tells me it's already running, fine, I'm just wait and I give you the result.
24:37If it's not running yet, GitHub will start that run with that ID, right? And that that is what's called idempotency Armin Your favorite word, idempotency.
24:51so cool. So if our external system has this feature that if I give it an ID and say do the thing with this ID and the the thing is already running and it tells me I'm already doing it, wait for for it to finish, then I can also survive crashes of my effect, right? And the final thing is once that effect has finished, I get a result back and I store that result and then my effect sandwich of intent, execution of the effect and storage of the result of the effect is done.
25:18And then I never have to think about it again.
25:19I can throw away all the information that happened, unless I need some trace for logging or audits or whatever.
25:24But the important part is I now have a result.
25:27That is all I wanted. And I wanted to get there irrespective of whether my process dies somewhere in the sandwich.
25:36Now there's complicated cases.
25:38For example, if I delete the file on disk as the effect, right?
25:45I usually cannot undo this unless I have Git or whatever, right? If I charge your credit card, I don't want to charge it twice.
25:54So all the effects that you're running also kind of need to be in a sense combinable with this durability system.
26:02And most external things are, but some are not.
26:05So then you have to decide how you handle with that.
26:07And so coming back to Pi's Pi durables architecture.
26:12Basically everything that happens in the harness is is a task that is just this effect sandwich. So for example, if I want to generate a LLM answer, I have a sandwich that is prepare the request, gather the context window and the parameters, thinking level, model, blah blah blah.
26:29Save that out, that we are now about to do this.
26:31Then actually call the LLM, this is the effect.
26:34And then you get the answer and you persist that answer.
26:38The same thing for a tool call.
26:40I'm getting the tool call from from the LLM.
26:42Then I have a task that executes one tool call from the LLM.
26:46And the the effect. Yes.
26:47And these are inside each other, I assume. Okay.
26:50One task can spawn another task and wait for it to stop.
26:54And that's the whole magic. There's nothing more to that.
26:57The other side is the application state I talked about, and there we we came up with something that's basically Do you know document databases like MongoDBs or document database?
27:08Where you basically store JSON objects, right? And you can freely modify them. And so for application state, we built a tiny little document store with the special feature that I have a document. And in the transcript at time zero, the document had that value, right?
27:27Say, for example, a to-do list, right?
27:29And then you work with the agent, user, assistant, user assistant, tool calls, changes the to do list.
27:35So now the document gets updated here in the transcript.
27:38Now, if I want to go back in time, which version of the document do I want at that point?
27:45This one or this one? Depends.
27:47Sometimes you want that version at that time, right? I want the old to do list when I was at this point in time.
27:54Other times, I just want the latest version of the document, irrespective of where I.
28:00And so our document store allows these two types of documents to be associated with any position inside of your transcript.
28:07And with that you can store your application state.
28:09For example, if I have a a canvas where I can draw on, that's a document that just stores a list of strokes that I can replay to draw the drawing, right? And as I talk with my agent and while I draw into the the thing and then say show this new drawing to the agent, that document changes and evolves together with the agent.
28:31If I go back in time, I wanna have the old version of that drawing, for example, for the agent. And I can do a lot of shit with just this basic primitive of version documents and tasks.
28:42And that's basically Pi Durable.
28:44So you have okay, you have like a version document and then as you're going through the session that document's being updated and each one of those is a version and what do you like store the difference between the versions or okay.
29:00Yeah, exactly. We we so there's something called JSON patch, which is an RFC, I think, I don't know.
29:06And we were looking at that first.
29:08So what that basically does is it defines operation operations on a JSON object, basically set this field to this value, create an array, set the element of that array to this value, and blah blah blah. And we looked at that or I looked at that and I found it to be lacking f for this use case.
29:26So we came up with something similar but optimized to the agentic.
29:31I would say like in general I'm a little bit surprised that I've been doing web development for I don't know like twenty years.
29:37And like the problem of partial like basically like React is all about state one, state two, how do you express that on immutable data structures?
29:47And somehow like it didn't actually ever really get significantly better, which surprises me.
29:53I mean there's things like Immer, right, which kinda mostly get it right.
29:58Yeah. I mean there's there's a little bit, but it's not like he's a well understood standard that everybody agreed on how to express like structural updates to nested data structures.
30:09and then you have like automerge, which is a project from the Inc and Switch guys that tries to give you some sort of like proxy object that records all the changes that you're doing to it.
30:20and then it's behind the scenes creates like whatever the CRDT is necessary to express that change.
30:28But it's like a huge hack in in how it's mapped onto JavaScript.
30:33It has so many holes in it. yeah.
30:35Just software development, right? I mean in in in Pi Durable there's obviously also some foot guns related to this kind of like we also use a proxy approach.
30:44What does that mean? that means basically I say to the storage, give me this document, you get a JavaScript object, which looks like a JavaScript object.
30:51You can operate on it like any other JavaScript object.
30:54I can see object.x equals ten.
30:56And that gets automatically recorded as an operation that is then applied to the real object.
31:04playback to kind of rest do this, right? and that also works with proxies, which is basically a hack, a slow hack, to be exact. But it works really well for the the things we need to do in an agentic harness.
31:15So There are tons of problems or like tons of like different opportunity spaces.
31:20I think like multiplayer is one of them.
31:22I don't really see much like good multiplayer harnesses.
31:24I'm trying to actually like onboard my wife to use Hermes more so that we have less issues coordinating things.
31:35but the yeah, th th there's just no good solutions.
31:39there's the whole fragmentation, I mean there's just like new models every day, new things, everything's changing constantly.
31:45Like why did you guys Decide that this is the problem to work on, as opposed to like some other agent harness issue.
31:55Well, I mean for me it was You mean multiplayer or durability?
31:57Durability, yeah, yeah, because you you've been speaking about a lot That there wasn't much out there actually.
32:03And and I mean there obviously things like Hermes, OpenClaw and all the other claws that exist, they or even AMP, right, with with their orbs where you can transfer sessions and they continue on. They all needed to solve the durability problem in some way.
32:16f for the agent harness specifically.
32:20But those are not generalized solutions that I can take and then build the next Earendil project on.
32:25Or or the next open source product.
32:28Or some internal Shopify agent thing that is a swarm or whatever, right? so it's very hard to take these things off the shelf and make them your own. So what I wanted to create was a very small primitive.
32:40I mean it's only like twelve thousand lines of code or so.
32:44And and and and make that generic enough, but also powerful enough that out of the box it can basically do everything you need from an agent.
32:51That includes compaction and all the other shit.
32:54And it's also, I think, like an important part here to to reason about is that you obviously don't need it.
33:01Like you definitely don't need Pi Durable.
33:03You you could totally build this yourself.
33:05You're so good at at marketing.
33:06But we had like multiple projects where we build agents and actually get them to work in a way where you trust them and they're like somewhat reliable and they keep like recovering from failure and but most importantly they keep doing that even if the human is not like babysitting on it and it's like, okay, this this crash now should we continue or something?
33:30Like actually making it like inherently reliable and like a thing that you feel like you want to build on, that actually in itself is is like is a lot of work.
33:40and not by the lines of code provided in the end, but just like the the c like the the core design that this actually works.
33:50And as agents are becoming more orchestrated and as they're becoming more running independent of like a human that observes them and and does something with them, you you kinda wanna like they are already like unreliable enough in a sense like there's a there's a LLM.
34:07It does whatever it wants, right?
34:08It's that it's like this this is organism of sorts.
34:11So the orchestration of that at least should be somewhat like understandable and like creating some some level of sanity.
34:21So I have a I'm trying to use my laptop or my computer less.
34:26Like one thing is like shutting the laptop down, even if you have caffeinate or whatever on, you're still like it's still it's still yeah, it's still annoying. So I'm trying to make it so that I can just send sessions between all my different machines.
34:39And I I asked GLM to do something like this.
34:43I called it parcels. It was made with Pi.
34:45So what it what it thought up of, I didn't decide on this, is just copy the JSON files that are the sessions from From one machine to another and then just like restart because it has all the environment variables already in there.
34:56and yeah, it's it should be it should be good enough.
34:59but it's still just like the the I would like it to just work by itself and just like spread over all of my devices at the same time.
35:07yeah, there's there's a lot, there's a lot of weird things happening.
35:12Yeah. So what what's next?
35:15Like what are you guys trying to do?
35:17What are you guys trying to do right now?
35:19Is it stabilizing this? Is it yeah, I I'm interested.
35:24Dog feeding mostly. Building lots of stuff internally and then deciding what to ship externally. I mean Armin, you wanna talk about that?
35:32Yeah, mean I I think like there there are two ways in which you can sort of like split Earendil the company. One of which is like the we've built some products and so we're working on products that ideally long term, even non-engineers find enjoyment in.
35:48not not coding agents. Yes.
35:48So not coding agents. We don't want to be a coding agent company next year.
35:53But at the same time we're also building like infrastructure pieces that people ideally find use for building their own agents on.
36:00so so like both of those are things that we're doing currently.
36:03and like originally like prior to working on on on Pi we worked on a product called Lefos, which was well now you call it like a personal agent, but it was like a an an agent sitting behind like a a pure text communication interface in our tr in our case it was email.
36:22and we really, really like that concept in general of like there being Like a a little entity that sits with you and helps you solve some problems but doesn't really impose itself onto like external parties.
36:36It's like it's like not like, hey, please book me, I have address an appointment, and then it goes on and annoys some human with like a i slop, but like a like a useful assistant that sits with you.
36:47and there's obviously there's a lot of stuff in that space, but simultaneously we don't really, really enjoy anything in that space.
36:56And they're like very quickly turning into some Massive crazy contraption of stuff.
37:01so in the same way as like a durable is I think it's also extremely risky. I mean like the the the proposition is like okay your ex you're already getting a lot of people trying to hack you and then you have millions of people who are using GrocBot and like logging in everything into the machine.
37:16I think that is like a catastrophe, like all all of these companies have like a catastrophe.
37:20I'm actually in a way constantly surprised not not more is getting like I'm entertained, man. It's it's super fun.
37:26Actually, like five minutes before we started talking, I saw a thread on Twitter by a guy who said, My grogbot posted all my personal finance information to my company Slack for everyone to see. And Elon said if there's anything like that happening, he will make me whole again with money, but he doesn't.
37:45Yeah. I I l I like the personal agents. I want I w I so I have Deepseek which I run at home.
37:52I used to use like Codecs for this, but the speed is atrocious. It's like thirty five I think I look thirty three tokens a second is the average via the Codex sub.
38:02It's just atrocious. And I I'm doing DeepSeek locally.
38:06They've gotten good enough, like these open weight models have gotten good enough.
38:08They can they can scan like fifteen different emails that I have and turn that into like linear and and take care of that. But it's still very messy.
38:18It's not very deterministic. you have you seen like the whole Jeff thing?
38:22I think that is that makes it to me that actually makes a lot of sense is having these Yeah.
38:27like semi what are they called?
38:31When I'm not head of MCP, I'm head of chef at Earendil He said Armin would marry Jeff if he could.
38:36He's very much in love with Jeff.
38:40Well merging the two I love chess.
38:41makes sense, like merging the two so you have like one that's like classifying for the other. I think that makes a lot of sense w when they're dealing with like high amounts of ball content.
38:51For example, emails.
38:52I mean it's ultimately the reason like finally MCP landed it in in in in Pi Yeah.
38:58had absolutely nothing to do with MCP was like I needed code mode for Jeff.
39:02So like from there to MCP it's like it's like five minutes.
39:06But for for Jeff to work or for any classifier model to work, it it needs another system to generate like a bespoke harness around it to kind of drive the thing.
39:19And and and there were it was just like obvious to build code mode.
39:22But yeah, the thing is like a lot of things are held back by the fact that they are like fundamentally very squishy tasks, so they need LLMs, but they're too slow.
39:33And so classifier models are like a a pretty neat hack now to I mean they are not new either, but they're like Jeff's really demonstrated that if you train it well enough to be a generic one-shot classifier.
39:47then like there is a product market fit for it.
39:49And then that started the whole thing.
39:51I just like in in a way it's surprising, but it's like I I love it for the fact that like a different a different category of model all of a sudden beyond like image generation L LMs all of a sudden is a thing that people are toying with.
40:04Yeah, and that is actually useful.
40:06Like I myself I have zero fucking use for image models, yeah. But I a chef model I can use for so many things and it's so much Do you not make memes, Mario?
40:15I am a meme, I don't need to make memes.
40:19So you also you guys also well at least you Mario you seem to like AMP and s specifically like the whore whole orc thing.
40:27W why is that? Like why does that stick out to you?
40:30I was gonna ask you about like how do you guys consider different harnesses?
40:33are you studying them? Is like if you look at like features, what's your approach to that?
40:42so if we deal with codex bullshit when the backend doesn't work, we clone the OpenAI slash codex repository and see what kind of fuckery they had to add to their shit to make their backend work.
40:52That's about the extent we look at other harnesses.
40:55with AMP it's different because AMP has been there way before Pi.
41:00Just like open code, also.
41:02been there since I think early 2025 or so.
41:04I I don't much look at open code anymore.
41:08I I did look at their B2, which is very interesting.
41:10They're also battling durability now, which is funny to to watch. but AMP I never used AMP.
41:18but I just like the team a lot.
41:20Because I think they they do kinda also have a vision of the future of software engineering and they're unsure of whether their vision is correct, but they're willing to bet their their house on on their vision, right? And they just execute and I think they execute well.
41:36I think I think Orbs works pretty well.
41:38I I met them in Munich earlier this year where they showed me the the the the Orbs thingy and I got to to to spend time with the rest of the team and it's just really nice people who who know their shit and who who don't bullshit a lot and I like that.
41:53I've liked their product. I I've been jumping around all the different harnesses.
41:56I think part of my job is sharing this stuff with other people.
41:58It's like telling them, you know, I've tried this, I've tried that.
42:00But I think Pi is really stuck out.
42:02Like and if I am ever building anything, I'm using Pi because well well one it's open source and two it's like actually comprehensible to me.
42:10But I I also think you you you did a lot with Omp, right? With with Shan's Pi thingy.
42:17So it's really good for local models if you have like a lot of VRM because they do this advisor thing. So you Mm-hmm.
42:24set up like Opus or whatever and it reads the entire conversation and it interjects whenever the model is like going in a loop or like veering off the the task, it'll like inject a message as if from the user telling it well like here's the scope and here's what you need to do differently.
42:42then you have like the sub agent things really like they I think It has a lot of like the clod code fanning out thing that Opus and Fable do.
42:51they have that really good, and then the the login, so just having the ability to do like slash connect and or whatever and and go in and like connect my stuff is very good.
43:02it's it's still like it's gotten better, but it it guzzles tokens that that's one issue with it, and its system prompt was huge.
43:11That was another issue. Deep C carnus I've been loving a lot.
43:14I don't know if you've taken time to look at it.
43:16It's it's phenomenal.
43:16I I I yeah, I I actually I read the source code.
43:20I I the plugin architecture is really interesting.
43:22They promise more than they can actually deliver, in just in terms of of actual plugin mechanism, but but the way they approach is really super smart.
43:31so I'm I'm really impressed.
43:33And it's defin that's actually one harness I actually looked at because obviously we want to evolve Pi and it's it's nice to see another harness kind of take the approach of the the thing itself must be self modifiable.
43:46And we now also Mm-hmm.
43:46see this in Claude Code, which is super fucking funny.
43:49when Boris says something like, my god, the mods they are amazing.
43:53How could we not like have Claude modify Claude Code itself?
43:58And fuck you motherfucker.
44:02I I think their harness has gotten a little bit better.
44:04I didn't I did Mm-hmm.
44:05not like it much last year, but it's gotten more stable.
44:09It still has so many bugs, like it just random errors or different sessions leak into each other. And that's that's one of the problems there.
44:16what about you, Armin? Is there any like harnesses that like stick out to you as something that you keep an eye on?
44:21I mean I I did I did look at DeepSeek harness.
44:24there was like when we started looking at the plugin architecture, there was a thing that like definitely w was worth looking at.
44:33I generally don't really look at that much what other people are doing because like it's a huge distraction. And you also don't necessarily know if they did it because it's actually solving a problem or if it's just like code is free, so it might as well have been like an experiment.
44:51When I did the MCP thing, I was basically looking at I was cross referencing every single other large harness that has something going on with MCP just to figure out if it's worth doing something.
45:03Because you get this list of stuff that people want.
45:05And it's like they really want to lend all of those and let me just check what the other ones are doing. I really hate elicitations.
45:13so there's like I think I used AMP quite a bit last year and I mostly stopped doing that, but I looked at the Orbs again.
45:23I do like the idea of orbs, but somehow I'm not very sticky to the web interfaces that much.
45:35So like I I do use the Cloud Mobile app and I use the Codex mobile app, but I never use the desktop apps.
45:42and and then like on the terminal, I actually don't think that they are particularly good.
45:49And I think it's not because I don't want to have a desktop app, actually I I really don't like the terminal all that much, but somehow I I haven't found this this perfect experience yet.
45:57Well th that's what I've liked about DeepSeek is it's it seems like the Pi of the desktop app.
46:02I've the the actual user experience is really minimal and very it's very clean and then because they spawn it as like a web process, you can just put it on tail net and then from my phone or from another machine and I I don't have to like do any weird like connect your device type things.
46:17That is th that's been a pain.
46:19do you guys think about doing like an app or or a web like yeah, web interface? I know I know you have kind of like yeah, I just slopped around the past couple of days on my phone building a thing that should primarily run on my phone.
46:37I call it PIM. Pi mobile. It's not the product I'm gonna put out as open source or anything. At least that's not the plan at the moment.
46:44but we're definitely exploring that for for the evolution of Pi itself.
46:49I think Pi is an initial proof of concept for self-modifiable agentic software.
46:56And now we need to get our asses up and evolve that into a direction where it becomes less I and the clanker and more many clankers on many machines talking to many humans.
47:09And it should be possible to to jump around and reconnect agents and humans in whatever configuration I want them to collaborate in.
47:19And I think that is where Pi should go.
47:20And with the PIM stuff I did, I just wanted to ensure that Pi Durable actually works.
47:25So talk fooding. I wanted to replace Claude for Android because I've also been a Claude for Android user and now it's completely replaced because I actually PIM works entirely on my phone except for the LLM.
47:37Everything else runs on the phone, which is I love it. it can also just connect to any of my other machines as the execution environment. So it goes to my Hetzner and then does sysadmin there for me while the brain runs on the phone.
47:49so I achieved already that part of my vision.
47:52And the next thing is to figure out how we turn this into the evolution of Pi with a plug-in system.
48:00And I I'm totally honest, Deep Sea Carness is definitely an inspiration there as well, not necessarily on a technical level, but on a conceptual level. and we o yeah.
48:10There's a surprising amount of complexity in a plugin system for web. And most of it actually for really dumb reasons.
48:18Most of it actually is just like purely like how do you make it secure?
48:21because you're sort of yeah, abstractly constrained by like it's almost like making it secure makes this the experience shit. So then the question is like, how do you find like a good balance between the whole thing?
48:36Yeah, the security stuff has been pretty rough for me.
48:39I've noticed like a few days ago one of the clankers was trying to share a file with another one of my devices and it just put my entire file system on the public internet for like a day.
48:50It was just like up there.
48:51I ha I have I have like wire like the the you know the guards and all that, but it's But but it's also like i the the goal should really be that multiple people can experience one session together in a in an operating environment that then our harness can extend. But the Mm-hmm.
49:07implication of this is that someone is more trusted than everybody else.
49:11And you don't want to end up in a situation where you have to fully trust the person to whom session you're connecting.
49:16Right? So imagine What about Docker? So I'm I I'm just thinking like it's this seems like a Docker type problem where you have something that is created and like it's an image.
49:26Yeah, but you're you're connecting i ignore for a second like the Docker part, right? Like that is almost trivial.
49:31Like how do you how do you protect the server?
49:33But like how do you protect the web interface?
49:35Because now you connect to my web interface, but the UI that comes up has like links in there that if I were to click it, then I would doing like like it's so easy that if we end up in a shared space where like I have an account, you have an account, the Clanker brings up some user interface that is literally XSS.
49:55to like exfiltrate my cookie or whatever.
49:58So like like pr it's and it's solvable, but it's it's annoying to solve.
50:04And it's actually really a lot of the solutions for solving this problems involve iframes, which suck for secure for for user experience.
50:10So like actually making an extensible web interface that's also secure is is way harder than than it looks and actually it almost seems like sandboxing Docker on the server is is easier.
50:22It's the easy part, yeah.
50:25But we're gonna find out. We throw some shit on the wall and find out.
50:27We're just not sure yet. But because the problem Yeah, Have you heard of IPFS? Sorry, IPFS.
50:32yeah, that comes out of the web three community where you have like a a ledger for your file system on the web that's distributed.
50:38yeah, I I don't think it it it solves the information exfiltration problem really.
50:44it might be an interesting way to store data.
50:49persistently, durably replicated, but I don't see it contributing to security much to be honest.
50:55It's mainly about like if if the clankers building something, it needs to put it somewhere that doesn't link back to your own.
51:02Like if you and me are working on something, we want a like a neutral space that we put it on that we could both trust isn't being yeah, like we could trust that it's gonna be there.
51:12It's like w whatever, but no, I s I I see this is this is too big brain for me to solve or even really think about.
51:19it's a big brain for us as well. We're just pretending we have a a clue of what we're talking about.
51:26Okay, so maybe w like what are you guys excited about?
51:30I have to I yeah. w Great amusing myself. I want to say like it's it's easy to ignore.
51:36You can basically say like, okay, let's just pretend we don't have to solve this issue.
51:42And then you can create a great experience, but it would be trust destroying.
51:48What was your question? What are we exciting about?
51:49I w yeah, what are you excited about like coming like in the next few months?
51:54building products that aren't coding agents.
51:57And evolving the one coding agent we have into a form that our existing community will actually like.
52:03That's a challenge. I already s we already saw just by releasing MCP and code mode support, which I don't know, that a lot of people can get upset by if you just add one feature that you previously said is not needed.
52:18But now things the models have been trained on this and the MCP server and the standards have become better.
52:23Now it is basically good enough to integrate as as as a built-in thing, which was our Mm-hmm.
52:28our thinking. And people get upset because how can you change your mind?
52:32Well, dude, I change my mind if new information is available.
52:35It's not like I'm running Not many.
52:37it's not like I'm on the hype train for MCP, my dude. It took us a year to add support for it, only after excruciating painful deliberation.
52:48So Hmm. What about you, Armin? Like what are you excited about?
52:55I mean I'm I think I'm primarily I don't know, excited is sort of like a thing.
52:59I I I feel like I'm really okay with how AI is right now.
53:03I just think the user experience around it can get better.
53:05So it's like what what really would excite me, I don't know, I feel like on a societal level, I think it would be great if the technology actually delivers something truly great.
53:14Valuable Yeah, like give me my room level superconductor, please. or like cancer medicine, whatever.
53:22On a purely technical level, I think like the main thing that I'm really hoping that we'll get to is that the hype dies down a little bit and we end up with really cool applications of this technology to just build better software instead of I don't know, like slop forks of of Photoshop.
53:40like like s re Yeah.
53:43like have actual true communities that build really cool software because like rather than stimulating their like In inner chemicals they're like actually working together to build really cool software. So that's I I give you an example.
53:58PS five emulation on PC is now solved thanks to clankers.
54:01Perfect. Fucking amazing man.
54:04That's even a contribution to society in the sense that it helps preserve PS five games for feet for the future.
54:09And I love that. That's awesome.
54:12Yeah, I've been saying to people like right now everything is just AI working on AI so and this is this also happened in crypto, like every bull bear market. The bull market, it's like it's not time to build anymore.
54:23It's just time to to market. It's time to sell and like you know, make the money. and then the the actual like serious builders, like they just kinda activate as soon as the bear market hits and they they start building the weird stuff that, you know, eventually pops off.
54:37But we w we're in that pro we're in that phase right now.
54:39It's it's extremely I think so, yeah.
54:41it's extremely insane.
54:42Yeah, I could see like it's it's not normal.
54:45But I don't think it'll last and this technology is very cool.
54:49Like I I I'm I'm really excited about like onboarding how can we take somebody that Yeah. Mm-hmm.
54:55has never used an agent and then have them use an agent?
54:57Because if if I could have my wife using an agent, my kid I so many things are gonna just stop being a problem.
55:07You just want your family l logistics to be handled by someone else.
55:12get other people hooked on the agent.
55:15No, but both me and Armin had experiences like that as well.
55:18Like I I had my wife introduced to that and and she did some linguistic research with Claude last summer. I I got my four-year-old then to write a game for my wife for her her birthday, stuff like that. It's that is cool, man. That's that's just that's just like how I like to use technology, right? And the onboarding stuff is exactly right.
55:36That is the biggest fucking problem.
55:38And nobody does anything. Everybody gives you a chat box and a sidebar and then Good luck motherfucker, figure it out.
55:45You are a secretary. Here, be more efficient with AI.
55:49Fuck you. That's just stupid.
55:53No, no effort has gone into like taking somebody and then making them like because the like the main thing is if you don't know, you can ask. And you will know.
56:03And the logical step from I don't know to like I can ask about it, that people do not have that skill set.
56:10this is like I think something you l I learned at least when learning to code is like you you shouldn't be asking and should like go figure it out and if you can't go ask somebody.
56:18Don't don't waste forever. but yeah, thank you guys for taking the time to do this.
56:29Even though some somebody who was it? yeah, it was it was Eric from OpenAI who who I think dunked on your your your S tier post, right? Unexpected.
56:44well I mean I think I think it's true.
56:46I think it I think it's absolutely true.
56:48It's the only one that's like well thought out.
56:49Like well like it's clearly well engineered.
56:52And I like droid, so I use droid a lot and I I love the team.
56:57But I think what what happens with like any AI project is that you can it's code is cheap, code is free, or you know like generating the slop is free and you're just sitting with like so much slop that is thirty percent finished.
57:10And I think that over Over time that really destroys the user experience because you're constantly running into like bugs and errors and I don't know, like the agent keeps using the ask question tool and I just wanted to do the goal or loop and this is yeah, that this this is what I mean, is like this stuff by going the minimal route you overcome a lot of the issues that the models like are naturally prone to.
57:33Yeah. Okay guy. Thank you for having us and have a nice have a good night.
57:37Thank you, I'll see you in the next one.