# Pi Durable: An Agent Harness That Survives Its Own Crashes

> Earendil shipped Pi Durable, an experimental MIT-licensed TypeScript package, alongside Pi 1.0 on October 1, 2026. Its agents survive process crashes: every model request, tool call and compaction is a task that checkpoints to SQLite or JSONL storage, and each tool declares whether a rerun after a crash is safe. In a kill -9 test, a replay-safe search ran again while an unsafe booking was reported to the model as interrupted instead of repeated. The package is about 15,000 lines, and one process owns a storage at a time.

URL: https://blog.bokvi.com/blog/pi-durable-agent-harness/
Published: 2026-10-03
Tags: ai-agents, llm-integration, javascript, architecture

**On October 1, Earendil, the company behind the Pi coding agent, shipped [Pi 1.0](https://earendil.com/posts/pi-1-0/) and, next to it, an experimental package called [Pi Durable](https://earendil.com/posts/pi-durable/): a TypeScript harness for agents that keep going when the process under them dies.** Every model request, tool call and compaction is a task that writes a checkpoint to storage before it moves on, so a fresh process opens the same file and carries on from there, even if the last one died in the middle of a tool call. That is the piece missing between a coding agent you babysit in a terminal and the Slack bot, triage bot or research agent you actually want to leave running for weeks.

I read the announcement and the [package README](https://github.com/earendil-works/pi/blob/main/packages/durable/README.md), replayed Earendil's demo recording frame by frame, and then killed a Pi Durable process with `kill -9` myself to see what came back. The short version: it's small, it's honest about side effects, and it isn't distributed. I think that's the right trade.

![A terminal UI with a task tree: vacation.research #12 running in the background and owning conversation 13, pi.generation #21 waiting on 26, 27 and 28, and three pi.tool tasks all running execute](https://blog.bokvi.com/_astro/vacation-parallel-searches.DzTkfZbp_lT8cc.webp)

*Earendil's vacation planner hands research to a subagent, which runs three searches at once. Every row in the task tree is a durable task with its own checkpoint. Frame from Earendil's demo recording, re-rendered.*

## Pi stays in the terminal, Durable goes everywhere else

Pi, the coding agent, is a one-person terminal tool, and Pi 1.0 keeps it that way: if it dies, you look at what happened and tell it to continue. Pi Durable is a separate package for agents you leave running: reachable from any client, steerable by several people, able to survive its own crashes. It shares Pi's model layer, `pi-ai`, but it is a framework for building agent applications, coding agents included, not a new Pi.

The size is part of the pitch. The whole package is about 15,000 lines without tests, which Earendil puts at about 150,000 tokens for GPT and 250,000 for Claude. The intended reader of that code is your agent: the announcement's try-it section tells you to point your agent at `packages/durable` and have it read the README and the examples. I like that. It is a framework sized to fit in a context window, and the 3,000 lines of storage backends are the part your agent can usually skip.

## Everything is a task, and every task keeps a checkpoint

Earendil has [a whole post](https://earendil.com/posts/what-is-a-harness/) on what "harness" means. For Pi Durable it is storage, the machinery that runs many model conversations in parallel, the tools those models call, and the execution environments the tools run in. A conversation is a transcript, an agent is a model with its settings and tools, and everything the harness runs, from calling the model to executing a tool, is a task. If you have read [how agent architectures are usually layered](https://blog.bokvi.com/blog/ai-agents-architecture/), that last part is the new one: the agent loop itself is made of durable tasks.

Storage is pluggable. Pi Durable ships SQLite and JSONL backends that persist, an in-memory one that doesn't, and a conformance suite for writing your own on a key-value store or Postgres. The SQLite and JSONL code uses no Node APIs, so with a small adapter it runs on Bun or inside a Cloudflare Durable Object. On SQLite only the working set stays in memory, and because compaction keeps active transcripts inside the context window, memory stays flat even at tens of thousands of messages.

Tools reach files and a shell through an execution environment. The shipped one is local Node; write a remote one and the harness can run on one machine while the files and shell its tools work on live on another.

## Kill the process, lose almost nothing

The demo's most important moment is the least dramatic one. The planner's subagent is researching weather, museums and trains. Weather and museums finish. Then the process dies in the middle of the trains search, and someone types `vacation --continue` at the shell.

![The last frame of the dead TUI, with one pi.tool task still running, followed by a shell prompt where vacation --continue has been typed](https://blog.bokvi.com/_astro/vacation-crash.C_KbcCWs_Z2nsOvh.webp)

*The process is gone mid-search, with only the trains search, pi.tool #28, still running. Frame from Earendil's demo recording, re-rendered.*

A new process opens the same storage, finds the unfinished tasks and continues each from its last checkpoint. A model request that was cut off is sent again, and the partial answer stays in the transcript, marked as aborted. Queued messages are still queued. In the demo, the finished weather and museum searches stay finished, the conversation and its task tree come back as they were, and only the trains search runs again.

That last part isn't automatic. Whether a cut-off tool call runs again is up to the tool.

## Replay is a promise each tool makes

A harness can't know whether rerunning your code is harmless, so each tool says so itself, and the default is no. I wanted to see that with a real `kill -9` rather than a recording, so I wrote a short script against the package in a checkout of the Pi repository. It uses Pi's built-in `faux` provider, a scripted stand-in for a model, so no API key is involved. The "model" asks for two tool calls at once: a search that is safe to rerun, and a booking that charges a card.

```ts
const search = defineTool({
    name: "search",
    description: "Search train times",
    parameters: Type.Object({ route: Type.String() }),
    replay: "safe", // only reads, so a rerun after a crash is fine
    execute: async (args) => { /* two seconds of searching */ },
});
const book = defineTool({
    name: "book",
    description: "Book and pay for a seat",
    parameters: Type.Object({ train: Type.String() }),
    // No replay: an interrupted booking is reported to the model, never repeated.
    execute: async (args) => { /* two seconds of charging the card */ },
});

const job = { type: "input", content: "Find a train to Graz and book it", requestId: "job-42" } as const;
await root.submit(job, context);
```

I started it, killed it 1.5 seconds in while both tools were running, and started it again on the same SQLite file:

![A terminal: the first run starts search and book and is killed with kill -9; the second run resubmits job-42 and gets submission 7 back, runs the search again to completion, and records the book result as an error saying the tool was interrupted and may have partially run](https://blog.bokvi.com/_astro/kill-9-resume._yBLxzfD_yordR.webp)

*The second process gets submission 7 back for job-42, reruns the search, and hands the model an error for the booking instead of charging the card again.*

Three things happened, all as documented. The `requestId` made the submission exactly-once: resubmitting `job-42` returned the original submission instead of starting a second job. The search ran again from scratch. The booking didn't; the model got `Tool book was interrupted and may have partially run` as its result and has to decide what to do next. The final answer in my run is scripted; with a real model, that decision is the model's.

That is the honest answer to the oldest problem in durable execution. `replay: "safe"` makes a tool at-least-once and the default makes it at-most-once; an effect outside the harness happens exactly once only if the far side is idempotent. Earendil's longer example, a checkout that charges several cards in parallel and refunds the rest when one is declined, passes an idempotency key to the bank for the same reason.

Hooks follow the same rule. They can step into tasks, including the built-in ones for model requests, tool calls and compaction, and because a hook can run again after a crash, a hook that makes a decision stores it in a memo, a small value saved with the task where the first write wins. This one, from the announcement, goes in an extension's `hooks` list and gates deploys behind a Slack approval:

```ts
hook(ToolTask, {
    beforeTool: async (call, api, context) => {
        if (call.name !== "deploy") return undefined;
        // After a restart, the hook finds the stored answer instead of asking again.
        let approved = await api.memo<boolean>("approval:deploy", context);
        approved ??= await api.memo("approval:deploy", await askInSlack(call), context);
        return approved ? undefined : { block: "Nobody approved the deploy." };
    },
})
```

Without the memo, a crash after the answer came back but before the deploy started would ping your team a second time. With it, the approval survives the process.

## Subagents, forks and steering are just conversations

Pi Durable has no built-in subagents; Earendil says they "take a few lines of code to build". A tool gets the harness API for its call, so it can create a conversation it owns, give it a smaller model (`gpt-6-luna` in the example) and its own instructions, submit the work and wait. The subagent counts its own cost and a UI can show it under the call that started it. After a crash, a replay-safe subagent tool finds its subagent again and waits for the answer, while a crash in a call that isn't replay-safe aborts the subagent with it.

Conversations and tasks form one ownership tree. Aborting a task aborts everything it owns, bottom-up, so each piece cleans up its own effects first. Tasks are foreground by default, so Esc stops them; a background task, like the research job in the demo, keeps going when the conversation goes idle.

Because the subagent is a real conversation, you can switch to it and talk to it while it works:

![The subagent's view in the planner TUI, with the trains search printing source 8 of 15, a steer message queued that reads skip the trains, we're driving, and the status line labelled subagent 13](https://blog.bokvi.com/_astro/vacation-steer-subagent.D8P1bbx8_Z1ybOg9.webp)

*Inside subagent 13, mid-search, with a steer message queued behind the running tool call. Frame from Earendil's demo recording, re-rendered.*

Any client can do that. Everything a UI needs is committed state, so a second client can attach to a running conversation late, get the current view and then only the changes, and send a message with `whenBusy: "steer"`. Forks are just as cheap: a conversation can fork another at any entry and see the parent's history without copying it. Earendil's example is a Slack channel as one conversation and each thread as a fork at the message it replies to, both running at once, with the thread allowed to search but not deploy:

```ts
const thread = await channel.fork(answered.answer!, { ownership: { kind: "ownerless" } }, context);
await thread.configure({ tools: { remove: [deploy] } }, context);
```

Each conversation stores its own agent, from the model and thinking level to its tools and working directory, so a reviewer next to the main agent can run a cheaper model with read-only tools in its own checkout. Application state, like a todo list or a plan, lives in typed JSON documents committed together with the transcript, so the two can't disagree, and each document says what a fork inherits.

Even the code can change underneath a running agent. Load a new version of an extension with your own loader and `registry.install()` it under the same name, and it replaces the old one in one step; a call already running finishes on the old code, and the next call uses the new one. Conversations store extension names, never code, so after a restart they bind to whatever the new process installs.

## Compaction runs while the agent keeps talking

In many harnesses a long conversation hits a wall: the agent stops, summarizes, and you wait. In Pi Durable compaction is a task like any other. With the settings in Earendil's example, a summary of older messages starts in the background once the conversation is within 49,152 tokens of the context window (`reserveTokens` of 16,384 plus `backgroundTokens` of 32,768) and lands at the next turn boundary. The conversation only waits for it inside the last 16,384 tokens, where the next request wouldn't fit otherwise. If a provider still rejects a request as too long, the harness compacts and retries once, and you can compact by hand, with your own instructions, while the agent works.

![The main conversation answering a question about Sachertorte while a user message reads two sentences, please, and the task tree shows pi.compaction #39 summarizing alongside two running pi.generation tasks under the label Compacting (manual)](https://blog.bokvi.com/_astro/vacation-compaction.Cl6cU01v_4QW0m.webp)

*A manual compaction, a new answer and the subagent's model request run at the same time. Frame from Earendil's demo recording, re-rendered.*

Nothing is deleted. Older messages stay in storage, so `reset()` can start a fresh context from a handoff note while a `search_history` tool you write can still search the messages from before. That's all it takes to build an agent that hands off to itself.

![The planner's final answer, a Vienna weekend plan with a day 1 route from St. Stephen's Cathedral to the State Opera and Albertina and a day 2 route through Upper Belvedere and Naschmarkt, with no live tasks left and a note that the conversation was compacted](https://blog.bokvi.com/_astro/vacation-plan.BSTRXRl__ZhweaR.webp)

*The subagent's report arrives as a message and the compacted main conversation turns it into a plan, which still says to take transit rather than drive. The harness guarantees the steer arrived, not what the model makes of it. Frame from Earendil's demo recording, re-rendered.*

## What it unlocks

Here is my read of what changes when the agent loop is durable by default.

**Agents become services, not sessions.** Earendil says it will show the small tools it builds with Pi Durable for its own work in the coming weeks, like a Slack bot or a GitHub triage bot. That is the shape: a bot that labels issues with a cheap subagent and doesn't lose its queue on a redeploy, a reminder that fires tomorrow because its deadline is stored, a research run that outlives the laptop going to sleep.

**Agents that live at the edge.** Storage without Node APIs means a harness inside a Cloudflare Durable Object, with tools reaching out to a remote execution environment. Timers only fire in a process that has the harness open, so an evicted object needs a Durable Object alarm to wake it and call `resume()`. I still expect "one Durable Object per conversation tree" to be the first serious hosted pattern people build on it.

**Agents you share.** Several people watching and steering the same conversation, joining late from a different client, is the default rather than a feature you bolt on. An on-call agent the whole team can nudge is mostly UI work now.

## Where it stops

It is experimental, and the README says the API "changes without notice between releases." Build on it knowing you will chase changes.

Durable here means surviving a crash, not running distributed. One process owns a storage at a time and other clients attach to it, but nothing enforces that: the README says there is no cross-process locking, so a lock or lease is your job (the vacation demo takes a lockfile on its session directory). Mind it in rolling deploys, where old and new processes overlap. Something still has to restart that process when it dies, and nothing in the package spreads one harness across machines.

Crash means process crash. The SQLite backend runs with `synchronous = NORMAL`, so commits survive the process dying, but the newest one can be lost if the host or its power goes; JSONL needs `{ fsync: true }` to flush before each commit. Lose the commit that recorded a tool call's start and that call runs again, replay-safe or not.

Replay safety is a promise you make, not one the harness checks. Mark a non-idempotent tool `replay: "safe"` and a crash mid-call can run it twice. The design makes the decision explicit; it can't make it correct.

It is TypeScript only for now; Earendil's FAQ says a Rust port isn't ruled out, but isn't the focus. And the API is low-level. Earendil calls a subagent "a few lines of code", and the triage example in the announcement is about 50 with its comments. That's the price of no magic, and I think it's the right price, but this isn't a five-minute framework.

## Try it

From a checkout of the [Pi repository](https://github.com/earendil-works/pi):

```sh
npm install && npm run build
node packages/coding-agent/src/experimental/durable/main.ts
node packages/coding-agent/src/experimental/vacation/main.ts
```

The first is a [small coding agent on Pi Durable](https://github.com/earendil-works/pi/tree/main/packages/coding-agent/src/experimental/durable); the second is the [vacation planner](https://github.com/earendil-works/pi/tree/main/packages/coding-agent/src/experimental/vacation) from the recording, about 1,300 lines of TypeScript, most of it TUI code built from the coding agent's components. Kill the planner mid-search and start it again with `--continue`, or it opens a new session. After the build, most of the package's [examples](https://github.com/earendil-works/pi/tree/main/packages/durable/test/examples) run from `packages/durable` on the same `faux` provider I used, so you can watch recovery without an API key. To build on it in your own project:

```sh
npm install @earendil-works/pi-durable @earendil-works/pi-ai @earendil-works/chord
```

Both Pi 1.0 and Pi Durable are MIT-licensed. If you have been holding agents together with retry loops and a JSON file of "where was I", this is the first framework I'd point my own agent at and say: read this, then rebuild it on top.