lukebuehler 1 day ago

Very cool to see Pi build a durable agent harness too. I've been building in this space for quite some time myself [0][1] and it is a super interesting place of innovation. Less hype-y that on-your-machine coding agents, but all major players are building products in this space: LangChain Deep Agents, Vercel Eve, OpenAI Agents API, Anthropic Managed Agents, etc.

The main reasons are:

1) they are "durable", i.e. easier to make long-running in an unattended way, and easier to implement recovery, monitoring, etc

2) separating the harness from the compute brings safety and scaling benefits

3) easier to make multi-player.

[0] https://github.com/smartcomputer-ai/lightspeed

[1] https://github.com/smartcomputer-ai/agent-os/

  • the_mitsuhiko 1 day ago

    It's quite an interesting place to hack on, because it's both a well understood problem but with so many wrinkles to it. I can't count how many earlier designs we chewed through before we ended up with the final one and I would not be surprised if we learn even more about it.

    • lukebuehler 1 day ago

      Yes, from your release post I can tell that you put a lot of thought into it!

      I like your structured concurrency approach with tasks, which is similar to how I do it in Lightspeed too.

      Also, the durable state implementation as documents is elegant! Question, though: why directly write/read to the store, why not abstract it and do more of a reducer/redux pattern and hide the persistence of the documents?

      • the_mitsuhiko 1 day ago

        > why directly write/read to the store, why not abstract it and do more of a reducer/redux pattern and hide the persistence of the documents?

        We tried so many things. At one point it pulls in so much more complexity. At one point we had half of automerge's proxy system in there. In the end we felt like this is a reasonable line to draw, but we will see!

      • badlogic 1 day ago

        Things are still not settled, and while the API looks like store i/o it's actually more similar to Immer's drafts, just with a different encoding, as JSON patch can't deal with the kinds of data we encounter in our workloads, at least not in a way that keeps memory and perf within some bounds.

        The good thing is that this more low level API can be easily papered over with a nice sugary thing.

    • rcarmo 1 day ago

      I'm loving it. Already converted a few of my smaller tools to it, and am ripping out the guts of piclaw to replace them (in time)

  • anilgulecha 18 hours ago

    Pi is anyway a better option than all the others, because of it's focus on agnosticism. Vercel's SDK will work slightly better with it's AI gateway, OpenAi's Agent will work better with codex models, and so on.

    Pi-durable makes pi a good acquisition target for Cloudflare - nothing like durable objects (with containers no less) really exists in other clouds. Wonder what the_mitsuhiko thinks about this.

    • cryptonym 15 hours ago

      > focus on agnosticism

      > acquisition target for Cloudflare

      Once it gets acquired, it won't be any better than all the others.

      • aquariusDue 14 hours ago

        From what I know (someone correct me if I'm wrong) Pi/Earendil is VC-backed so it's just a matter of time really. I don't see any other end game for them as a company other than being acqui-hired. In this regard I have more faith in something like Zed (the code editor, VC-backed too) to retain their "independence".

        Pi is alright but it's only virtue so far is being the Neovim of harnesses, minimal yet (incredibly) extensible. That being said it's still early days and unclear how the whole landscape regarding agentic stuff will play out at various different levels. For example I prefer stuff like Claude Code and Pi but I've seen a friend use Kiro at work with some crazy workflows all basically structured around Markdown files, spec writing, ingesting tickets from Jira and then validating/testing the code written automagically.

    • lukebuehler 14 hours ago

      What is interesting here is the _library_ approach to durable agents. All the other options I listed, including mine, take more of an SDK/batteries included approach. So this is very much in line with the Pi philosophy in general.

      So, it'll be interesting to see if these durable agent setups need more of a complete product approach, or if people want to compose them as libraries.

vito 8 hours ago

Will be following this closely! I've been working on something similar called `dagger agent`. The idea is to build on Dagger's sandboxing and (ab)use its existing infrastructure of "reproducible function recipes flowing through OpenTelemetry" - by replaying those recipes from their trace in your local engine or from Dagger Cloud.

The durability has been a lifesaver when I'm dogfooding and the TUI crashes or the session spawns 5 ambitious subagents and OOMs my laptop. Quite a few times I've migrated a session to Cloud's engines and kept going from there. Effectively the entire state of multiple sandboxes is put through storage as humble OTLP data and revitalized on whatever hardware you bring it back on, like thawing Walt Disney (and giving him a stimpack I guess).

My end goal is to have long-running sessions that hold curated context via tool state so it's durable to compaction, and to be able to keep a multi-agent month-long workstream going from my phone by messaging its leader through Cloud. It's been a passion project for over a year and it's taken a lot of world-building along the way (new Dagger core APIs, and Dagger 1.0 being the top priority independent from all this agent stuff). I'm finally at the point where I'm using it productively but I don't have a quick-start yet (edit: here's a quick and dirty one - https://gist.github.com/vito/bed465ed43a09a433b85a360311c7d3...)

Find me in Dagger.io's Discord if you're interested :)

  • gavmor 7 hours ago

    First I'm hearing of Dagger.

    I'm still using Concourse every day... for fun (and GPU mutex) — I guess I'll switch to Dagger? If it's light enough to run agents, I'm very curious.

    • vito 7 hours ago

      Happy to hear from a Concourser! The transition to Dagger isn't fully defined yet, but we're getting there. It's taken a lot of time to find the right primitives and build a solid engine for maximal caching, but we're excited about what we're launching soon. It covers the middle parts of a Concourse pipeline really well (build + test, CI checks) but we still need to solve the beginning/end (monitoring inputs, shipping/releasing).

  • thinkxl 7 hours ago

    I'm a happy Dagger user. I use it heavily for GitOps, but I never thought to use it like this.

    Definitely joining the Discord server.

lemming 21 hours ago

One decision here which seems like a large break from the original pi is that Durable doesn't support branching conversation trees, it only supports conversation forks with ancestry information. Can anyone speculate (or confirm, if you happen to be Armin or Mario) why this is, and if that is necessary for the durable guarantees? The branching conversations are still an immutable data structure, so I can't see why this would be necessary, but perhaps I'm missing something.

  • unified101 18 hours ago

    A fork is branch right? This is how pi's branches are built.

  • CGamesPlay 18 hours ago

    I think it’s just for consistency. A fork and a tree navigation are the same operation conceptually. Now, unlike older Pi, a fork is not a copy of the session, it just has a pointer to the older session. The only losses that I can see are: now /resume shows every conversation rewind; and /tree is harder to implement. Neither of those is provided by Durable, so the gap is left to the implementor.

  • badlogic 12 hours ago

    Oh, the tree is still there. It's just flatter :)

    In Pinthe coding agent, each transcript entry is parented to another entry. That was actually exceptionally dumb.

    If you do /tree in pi, pi needs to flatten that tree into linear, nested conversations.

    In Pi Durable, we corrected this mistake. A conversation is a chronological, immutable list of entries. A conversation can be parented to an entry in another conversation, and thus inherits that parent's older conversation entries starting from that entry.

    So, exactly the same functionality, just less dumb.

pulkitsh1234 11 hours ago

Can someone explain how these durable agents handle state present on the VMs/Sandboxes ? I get it that the agent state can be recreated from checkpoints/logs, but what about the state present on the runners (i.e. container, VMs, Sandboxes, etc). How are both states kept in sync ?

Like if I have a web-app running on the runner and the agent is navigating the web UI and then the runner (or the agent) crashes. When the agent is recreated back from the checkpoints (or a new runner is launched), it will think it has already navigated to page N, but in reality the browser on the runner might be on page 0.

  • arnorhs 9 hours ago

    I'm glad you pointed this out. This is my main question as well.

  • phoghed 9 hours ago

    Think about how you'd do it as a human, that's usually the answer for these things in my experience.

    If you've ever worked with any workflow engine, doing it with agents is largely the same. If you were writing some automation that used a browser, how would you handle recovery for any given step of your workflow? It depends on what you're doing, the specifics of the web app you're interfacing with, etc.

    Contrived example, but let's say you're sending an email. Load the page, click the button, enter text in the various fields, etc. Since there's no side effect of consequence until you hit send, you could just make sure your failures clean up drafts, and replaying the whole thing is safe.

    I've admittedly done very little browser automation like this though, mainly I've just called APIs, created and uploaded files, done db operations, normal dev stuff.

    I'd expect RPA platforms to be way ahead on agent automation like this, I haven't kept with any of them though. If they aren't, real missed opportunity for them.

  • everforward 7 hours ago

    The good version of it is to only allow interacting with the environment declaratively (preferably idempotently as well), and then checkpointing those.

    So rather than saying “open chrome, go to this page, click next page 5 times”, it would be something like “chrome is running; url is X; url is X/page/1; url is X/page/2” etc. Ansible, basically.

    Other than that, you could just replay bash tool calls. That’s full of holes, though. Anything that relies on “date” will return different stuff, and if you try checkpointing the system time then TLS breaks due to timestamp differences.

    If you wanted to go absolutely wild, some hypervisors can checkpoint the memory of a running VM and revert back to a prior version memory and all. I can’t imagine a way to make money off that (you’d be writing gigs of data per checkpoint), but I suppose it’s technically possible.

    • svieira 5 hours ago

      > If you wanted to go absolutely wild, some hypervisors can checkpoint the memory of a running VM and revert back to a prior version memory and all. I can’t imagine a way to make money off that

      You may want to check out what Antithesis is doing in this space - https://antithesis.com/ Specifically, https://antithesis.com/docs/resources/deterministic_simulati...

      • everforward 26 minutes ago

        It’s neat but I don’t think fixes these issues because they have to integrate with live systems. You can lie about system time in tests by making everything else have the same time.

        That doesn’t apply if I need to hit gmail.com and can’t login because my system time is a week behind and the JWT says it isn’t valid for another week.

        Even if you make everything else accept the time, timestamps will be screwed. Like in a fake gmail service that accepts mocked times, what do you use for timestamps on emails?

  • jmtulloss 7 hours ago

    With the one I’ve made, it’s best effort. We record whether or not tool calls have succeeded and if the agent dies without knowing whether or not the call succeeded, then when we resurrect it we tell it the state is unknown and it can either inspect the resource or try again depending on the side effects.

  • the_mitsuhiko 6 hours ago

    > Can someone explain how these durable agents handle state present on the VMs/Sandboxes ?

    That greatly depends on your agent design. If you give a user a full sandbox then you're going to be in a position where you probably need to snapshot it. But there are plenty of agent designs that are not using full VMs and for those the state story is way easier.

phainopepla2 23 hours ago

What are people using these infinitely-running agents for?

  • rubslopes 22 hours ago

    one example: I have several cron jobs that monitor different projects and notify me via my claw agent in Telegram. Whenever it sends me a message, I can ask it to address the issue within the same conversation.

  • shepherdjerred 22 hours ago

    I don’t really understand the infinitely running case, but I do have schedule agents to do things like open PRs when some events happen, or triage alerts every day

    • miki123211 21 hours ago

      Something I've had Chat GPT do for a while was writing "Daily Presidential Briefings" for me. Basically non-clickbait, well-summarized, priority-ordered news.

  • plaguuuuuu 19 hours ago

    The extremely basic use case is that I'm doing something at work and it's not finished yet when I leave the office.

    Or my laptop crashes, ugh.

    Yes, if I could ssh into a random server it'd be fine. But I can't.

  • lukebuehler 14 hours ago

    Not infinitely, but 6-12 hour long individual runs: complex analysis in enterprise across many data sources. Basically, long running investigations that touch databases, many files, apps via computer use, and so on.

arbuge 7 hours ago

> A harness is storage plus the machinery needed to run one or more conversations with large language models in parallel. It provides the tools those models call, and the execution environments the tools run in.

I don't want to quibble on definitions and semantics here but by this definition, there wouldn't be a single harness out there that I can see except those that include paid cloud storage like Dots.

ernsheong 22 hours ago

This stuff is really complicated. Just trying to build a harness coordinating multiple instances of vanilla pi has been a bit of a nightmare. I'm not sure if the huge added complexity is worth it but kudos for trying and labelling as experimental.

ice3 9 hours ago

Very similar to my work flow with pi.

I do use a git as pi memory with commit history and a worklog. Restarting pi just continues from the worklog + last commit/uncommitted changes.

Durable ... is different but I definitely will try it out.

vmg12 1 day ago

Brilliant. I wish the the durable application state wasn't restricted to just the json documents though. There should be some sort of integrated way of implementing the outbox pattern so external stores can be synchronized with the conversation state.

  • badlogic 1 day ago

    You can already (sort of, kind of) do that via a task (please excuse the agent slop, it's midnight and it's been a long day):

    ``` const SyncToPostgres = defineTask<{ entryId: string }, { phase: "send" }, void>({ kind: "app.sync-postgres", version: 1, initial: () => ({ phase: "send" }), phases: { send: async (task, runtime, context) => { const entry = await runtime.read(/* the entry */); await postgres.upsert("messages", { id: task.input.entryId, ...entry }); // idempotent by id await runtime.commit(() => ({ status: "terminal", outcome: { status: "completed" } }), context); }, }, });

       await root.commit(async (tx) => {
          const id = await tx.entry(AssistantEntry, answer);
          await tx.createTask(SyncToPostgres, { entryId: id }, { ownership: { kind: "conversation" },
     background: true });
       }, context);

    ```

    The commit on the root conversation picks out the last agent answer id from the transcript, and durably schedules a task that then syncs it to postgres. inside the task, you fetch the answer by id and send it over to postgres indempotently.

    What's missing here is sugar, basically a hook that runs inside each commit so the outbox write is atomic with the state change, with ordered delivery, and possibly a durable change feed with cursors.

    Thanks for the input!

snarfy 9 hours ago

Dots, durable, digital worker, it seems the whole agent space is converging on always-on packaged agents.

rsalus 1 day ago

super interesting. I have so many half-considered questions... like sandboxing (it seems like it is BYO). would love to have some kind of policy engine.. perhaps an integration with https://github.com/NVIDIA/openshell in the form of an extension?

also, I see most of the durability promise comes from persisting JSON documents locally and minimizing the amount of context/data kept in-memory, even during SQLite mode. while this makes sense, my own experiments with a process that relied on a JSONL-based event store have led me to prefer keeping things in-memory to avoid all the friction with I/O.. am I crazy for preferring just a straight .db file being persisted?

saagarjha 1 day ago

Nice! I made my own version of this for Pi but I’m excited to see if I can just replace it lol

skeledrew 23 hours ago

I like the multi-user bit the most. Should make it easier to build my remote control tool, as something I've had to hack around is not being able to use ACP while the TUI is active in an instance.

wolfcola 9 hours ago

Is Earendil backed by Peter Thiel or did they just also choose a LOTR name?

imtringued 12 hours ago

Please do not make the stupid mistake of baking in default tools again. Just don't do that. It's dumb as hell. Make the default tools very easy to opt-in by giving it a default profile. Then have two different commands. One command just uses the default profile, the other command e.g. call it pi-durable-agent or whatever, just to distinguish it, should run with a completely empty configuration.

If you add e.g. bash as a forced default tool, then someone can't come up with an extension called "sandboxed-bash", which internally runs the sandboxing logic and then delegates back to the bash tool.

By baking in your specific personal use cases you have made your software tool useless to the vast majority of people on the planet. Some of those people might decide to go ahead and use your software anyway and then run into massive headaches along the way and pretend those headaches aren't real, but that doesn't change the fact that the software design is incredibly poorly though out.

A coding agent doesn't necessarily need to write files. A review agent can just read the code, maybe it doesn't even read files on disk, maybe it just looks at a code diff on github and then posts a line by line comment. It does not need bash or node or whatever default tool you think is cute. It needs the tools I give to it and if it uses only the tools I give it, then I don't have to babysit it. If you let it run bash or node just to be cute, I have to babysit your agent harness. Is that so hard to understand?

If I need 100 different agent types, and they all have bash or node and there is a risk of them using bash or node when I only want it to use exactly the tools I want it to use, then why the hell would I choose your software? I wouldn't. I don't want to use it. It is completely illogical. Some people want to run agents as if they are microservices. Yes, that's me. I don't want to babysit every single microservice. You guys want to build the ultimate agent monolith and then call it minimal.

I got burned so I'm going to write my own harness anyway. Have a nice day.

lostmsu 18 hours ago

> A requestId makes a submission exactly-once, so a client that retries after a crash gets the original submission back instead of asking twice.

Sounds like a bug under a false assumption. Just having an ID cannot alone guarantee exactly once semantics AFAIK.

  • badlogic 12 hours ago

    The requestId is an imdempotency key, which is exactly how you get exactly once semantics.

  • VGHN7XDuOXPAzol 11 hours ago

    This phrasing from the article feels so Claudish to me. Is it just me?

azuanrb 23 hours ago

Cross-post from the 1.0 thread. I’m currently building a harness for Slack to support our on-call and support channels. It’s been working great so far.

The harness is built on top of the Pi SDK. I initially used Codex, but Pi seems more hackable, and I like that it’s vendor-agnostic by default.

Running it on Kubernetes works, but dealing with the JSONL session files and making sure sessions survive pod interruptions adds some complexity. I’m using DBOS for that right now, which works well, although it still feels like overkill.

This came at just the right time. I’m looking forward to removing the pieces I no longer need and simplifying the architecture. Thanks Pi team!

  • ghola2k5 20 hours ago

    I’m shoving this into Agent Substrate on Kubernetes with a different storage interface

croemer 23 hours ago

That code font is painful to read, no syntax highlighting and extremely pixelated. It's retro but an eyesore.

  • anentropic 11 hours ago

    agreed, it is bad - pixellated in a blurry way that turns letters into blobs

zmmmmm 20 hours ago

It's an interesting concept. This is half way to replicating pieces of Gastown. I like the idea, but I'm disappointed these tools still fail to address sandboxing as a first class citizen. I want to be able to declaratively set rules for what sandboxes agents execute in and mark context as tainted when untrusted etc. So far I still don't see any of these harnesses properly addressing this space. I'd be interested in knowing if it can be done through the extensibility of Pi, but since it operates directly on the trust layer, it feels like the type of thing that really needs native support.

  • jlkuester7 20 hours ago

    Not familiar with the details of Pi Durable, but I have tinkered a bit with different sandboxing strategies for Pi. IMHO it would be hard to trust a sandboxing layer built into a harness that is so focused on being fully pluggable/moddable/self-improvable.

    When I am using Pi to write extensions for Pi, I feel better running Pi wrapped in a separate os-level sandbox. I guess Pi could do it all, but I am content with how it is.

    • LeBit 19 hours ago

      After looking at so many options, that is also my take.

      These should be decoupled.

      Maybe I need nono in one context and smolvm in another or both.

      I would not want to trust the harness to self policy.

    • dbmikus 19 hours ago

      I agree in using a separate OS-level sandbox or a VM. Better to have the option for modularity.

      However, for ease of use, it is nice for harnesses to by default run with sane and safe sandboxing setup. Then give the option to disable them.

    • zmmmmm 18 hours ago

      if you only have one level of trust then running the harness itself in a sandbox and leaving it at that is fine. This works for coding. For more complex enterprise style scenarios it stops working. Say you have an agent reading emails for you to action high priority ones. You have to assume it is going to get prompt injected constantly. But you want to have an escalation pathway for a high priority email, so somewhere you need a tool that can modify state in a database. You can't give that trust to the email reading one. So you need a higher level agent that can spin up a low trust sub-agent, get an output from it, and then feed the sanitised output into a different agent that has rights to update the database. This is obviously simplified / toy scenario, but it just illustrates that there are different trust levels, and different agents need to be authorised to do different things.

      • elesiuta 16 hours ago

        I'm currently working on this [1] and can almost support this exact workflow. However the multiple agent orchestration is a sequential state machine, and other than network which can be set for tool states, filesystem access is still set only for the entire state machine.

        [1] https://github.com/agent6-dev/agent6

  • antonok 17 hours ago

    Earendil's own Gondolin tool is the best sandboxing model I've found so far. It just executes the toolcalls in a minimal ephemeral VM, unlike most others which run the whole harness inside the sandbox. It's a bit rough around the edges (doesn't play well with other plugins and doesn't work under Bun), but it's great if you're willing to put in some effort to tweak your setup. Much more comforting to fire off long-running parallel tasks when you know the blast radius is fully contained lol.

    • patates 14 hours ago

      I'm not trying to be defeatist but with these models, is there even a real way to contain the blast radius? I also run things sandboxed but it feels like it taking over the whole computer is at the distance of just one probability calculation going awry.

  • NitpickLawyer 16 hours ago

    > fail to address sandboxing as a first class citizen.

    Isn't it better if the tool is sandbox agnostic and you as the developer / integrator choose what's best for your use case? There are several levels of sandbxing, with many degrees of "freedom", so it would be really hard/confusing/overly-complex to build something ootb that suits everyone, no?

  • whazor 15 hours ago

    The benefit of Pi in my eyes is that the TypeScript interfaces make it easy to build your own sandbox.

  • badlogic 12 hours ago

    I do not see any resemblence to Gastown at all? Pi Durable is a library for writing durable agents. You can plug in any sandbox solution you like (aka execution environment in Pi Durable speak).

    Within a session, you can give each conversation its own sandbox, based on your application's needs and policies.