Skip to content
On this page

    Beyond the Ephemeral Loop: The Architecture of Durable AI Agents

    AI agents built on while loops lose state when processes crash. Durable agents require an embedded database and an operating system task scheduler.

    12 min read

    Developers build AI agents as a while loop around an LLM client.

    You send a prompt, execute the requested tool, append the result to an in-memory array, and repeat. When a container restarts, a process crashes, or a network drops during a 45-second tool call, you lose the session. The user receives a severed socket. The database retains orphaned records. Recovery requires starting over.

    Earendil paired the release of Pi 1.0 with an experimental runtime: @earendil-works/pi-durable. At v1.1.0 (released October 7, 2026), the package holds about 22,500 lines of TypeScript in src/, 72 test files, and a 263 KB design spec.

    Pi-Durable replaces the memory loop with an embedded storage engine and an operating system task scheduler.

    It is not alone. In 2025 and 2026, nearly every agent platform shipped some form of durable execution: Temporal, Restate, DBOS, Inngest, Cloudflare Workflows, and LangGraph checkpointers. The defaults still bite. LangGraph’s InMemorySaver keeps every checkpoint in RAM, so a redeploy wipes them all. Most teams find out in production.

    What makes Pi-Durable worth studying is where it puts the durability: inside the process, next to the agent, in one SQLite file.


    1. From Terminal Loop to Durable Engine: The Evolution of Pi

    The architecture of Pi-Durable reflects a concrete journey documented across two releases.

    Pi originated as an interactive terminal CLI. Created by Mario Zechner at Earendil, @earendil-works/pi-coding-agent version 1.0 started as an antidote to bloated developer environments. As detailed in Pi Agent: The Minimal Harness That Became My Multi-Tool Glue, the design embraced radical subtraction: no native subagents, no plan mode, no permission prompts, and a prompt under 1,000 tokens.

    That original runtime ran inside a terminal interface powered by @earendil-works/pi-tui with four built-in tools (read, write, edit, bash), appending messages to a JSONL session tree. If the process died, a human operator restarted the command-line interface and re-prompted the model. The operator acted as the crash recovery system.

    @earendil-works/pi-durable addresses autonomous server workloads. When an agent runs background triage, executes multi-step deployments, or supports team channels, you cannot rely on a human watching the terminal to restart broken processes.

    The Evolution of Pi: From Terminal CLI to Durable Harness Architecture comparison contrasting the classic Pi terminal coding agent loop with the multi-client Pi-Durable harness engine. PHASE 1 · PI CODING AGENT 1.0 The Interactive CLI Loop Single Human Operator Terminal TUI · Single keyboard driver In-Memory AgentSession RAM transcript · 4 built-in tools Appended JSONL File ~/.pi/agent/sessions/ tree log FAILURE PROFILE: EPHEMERAL Process crash loses active turns. Human operator restarts and re-prompts. PHASE 2 · PI-DURABLE 1.0 The Embedded Durable Harness Multi-Client Ingress Web UI / Slack bot / RPC workers Single Mutation Line Atomic commit · Tasks + Entries + Docs SQLite WAL / JSONL Storage Working set in RAM · Rest on disk FAILURE PROFILE: CRASH-RESILIENT Survives SIGKILL mid-turn. Resumes checkpoints; marks unsafe calls.

    This transition did not discard Pi’s minimalist foundation. As shown in Harness Engineering: Stop Blaming the Model, Fix the Environment, reliable agent behavior requires engineering the operating context rather than expanding prompts. Pi-Durable moved that boundary into an embedded runtime.


    2. The Single Mutation Line: Storage Precedes Visibility

    Standard agents publish before they persist.

    An agent streams text over a WebSocket: “I cancelled your subscription and credited your balance.” The process runs out of memory before executing the backing Stripe call. The customer reads the promise, but the transaction never occurred, and the in-memory transcript disappeared.

    Pi-Durable prevents this race condition through a single mutation line.

    The Single Mutation Line: Storage Precedes Visibility Three-tier isometric stack. Input drops to the storage foundation first, then rises to the client observation plane. The session kernel in the middle holds the single write lock. TIER 1 · STORAGE FOUNDATION SQLite WAL · JSONL · Durable Objects WAL mode · synchronous = NORMAL TIER 2 · SINGLE MUTATION LINE Scheduler · Commit Queue · Chord Docs One writer at a time · working set in RAM TIER 3 · CLIENTS (READ-ONLY) Web UI · TUI · Slack Bot · RPC Subscribe to committed viewState() INGRESS Submit input 01 · commit 02 · publish Nothing reaches Tier 3 until Tier 1 commits.

    In Pi-Durable, state changes pass through atomic storage commits:

    • Transcript entries (pi.user, pi.assistant, pi.tool-result)
    • Task state transitions (pending, running, waiting, terminal)
    • Document mutations tracked as Chord JSON deltas
    • Queued inputs in the conversation inbox

    The harness commits state to disk before streaming tokens to users. If any storage call throws, the Session fails and closes itself: every later call and every pending wait rejects with SessionFailed. Nothing retries. A half-written session is worse than a stopped one.

    Cloudflare reached the same rule from the other side. A Durable Object’s output gate holds every outgoing message and response until the storage writes before it are confirmed. If the write fails, the message never leaves. Pi-Durable applies that rule to tokens, tool output, and UI state.

    import { openNodeSqliteStorage } from "@earendil-works/pi-durable/storage/sqlite/node";
    import { Harness } from "@earendil-works/pi-durable";
    
    // Open the session over SQLite WAL
    const harness = await Harness.open(
      await openNodeSqliteStorage("./agent.sqlite"),
      { models, registry },
      context
    );
    
    const root = await harness.root(context);
    
    // Submitting work returns a durable submission handle
    const submission = await root.submit(
      {
        type: "input",
        content: "Refactor auth middleware to use JWT",
        requestId: "job-1042" // Exactly-once deduplication key
      },
      context
    );
    
    // If the worker crashes here, the next process reopens storage and resumes
    const settled = await submission.wait(context);

    Anchoring user output in a storage commit removes network race conditions. A client reconnecting after a socket drop resumes from the committed database state.

    What “Committed” Means on Each Backend

    “Committed” is only as strong as the storage under it. Pi-Durable ships four backends, and they make different promises:

    BackendImportSurvives process crashSurvives power loss
    MemoryMemoryStorageNoNo
    SQLitestorage/sqlite/nodeYesNewest commit may be lost
    JSONLstorage/jsonl/nodeYesYes, with { fsync: true }
    Durable Objectstorage/sqlite/cloudflareYesYes (Cloudflare replicates writes)

    The SQLite backend runs in WAL mode with synchronous = NORMAL. That is a deliberate trade. In WAL mode, a commit appends pages to a log file instead of rewriting the database, so readers never block the writer. With NORMAL, SQLite skips the fsync on each commit and syncs only at WAL checkpoints (by default, every 1,000 pages). A kill -9 loses nothing. A pulled power cord can lose the last commit.

    For a coding agent on a laptop, that is the right call. For a payment agent, pick JSONL with fsync: true or a Durable Object. Cloudflare introduced SQLite-backed Durable Objects in September 2024 and made them generally available in April 2025, with up to 10 GB per object and 30 days of point-in-time recovery. One Durable Object per conversation gives every agent its own database and its own single writer.

    One more constraint: one process owns a storage at a time. There is no cross-process locking. Point two workers at the same SQLite file and nothing stops both schedulers from running the same tasks.


    3. Document State: Bases, AST Splice Deltas, and Fork Pruning

    Agent state cannot live in freeform transcript prose alone. Complex workflows demand structured data: task lists, review diffs, active files, and billing ledgers.

    Pi-Durable uses Chord to store typed JSON documents next to transcripts. Writing full JSON snapshots on every turn exhausts disk space, while storing raw diffs slows down read queries over time.

    The storage engine balances this trade-off by alternating between bases and operational deltas.

    Document Storage: Bases, Deltas, and Fork Pruning Sequence diagram showing how documents store occasional full bases and compact AST splice deltas with latest versus rewindable retention. HISTORY: REWINDABLE · APP.TODOS Preserves Full History Seq 2: Base v1 Initial: {"items": ["write docs"]} Seq 3: Delta v1 Splice: ["p", ["items"], 1, 0, ["fix build"]] Seq 5: Base v1 (Checkpoint) checkpointWhen: deltasSinceBase >= 2 snapshotAsOf(entryId) Reconstructs value at any historical turn HISTORY: LATEST · PI.LIVE / PI.INBOX Automatic Storage Pruning Seq 2..4: Older Records Deltas deleted upon new base commit Seq 5: New Full Base Storage purges seq 2..4 upon new base Fork Policy Options: · asOf: Copies state at fork entry · current: Copies current head state · initial: Runs initial() on child access

    Storage writes changes as two record types:

    • Bases: Full JSON copies written at creation, during migrations, or when checkpointWhen() returns true.
    • Deltas: Compact Chord operations (Op[]), such as array splices (["p", ["items"], 1, 0, ["fix build"]]) and value replacements (`[“s”, [“text”], “Done.”]).

    The engine provides two history modes:

    1. history: "latest": Storage deletes older deltas when a new base commits. Used by pi.live and pi.inbox to bound disk usage.
    2. history: "rewindable": Storage preserves historical bases and deltas. snapshotAsOf(entryId) calculates document values as of past commit points.
    import { defineDoc } from "@earendil-works/pi-durable";
    
    interface TodoState {
      items: string[];
    }
    
    const TodosDoc = defineDoc<TodoState>({
      kind: "app.todos",
      version: 1,
      scope: "conversation",
      history: "rewindable",
      fork: "asOf", // Child conversation starts with parent state at fork entry
      initial: () => ({ items: [] }),
      // Commit a full base whenever two or more deltas accumulate
      checkpointWhen: (_value, _ops, info) => info.deltasSinceBase >= 2,
    });

    4. The Effect Sandwich: Executing External Effects Without Locks

    External effects cannot run inside a storage commit.

    A commit holds the database write queue. Executing an HTTP request, a Stripe charge, or a remote deployment inside a transaction stalls concurrent writers. Pi-Durable isolates side effects into a three-phase state machine: the Effect Sandwich.

    The Effect Sandwich: Intent, External Effect, Outcome Flowchart showing the three-phase pattern of committing intent, executing the external effect outside the commit, and committing the outcome. PHASE 1 · TRANSACTION LINE Commit Intent & Idempotency Key Checkpoints { phase: "charge", key: "pay-42" } to storage before acting Crash here: rerun prepare, nothing sent PHASE 2 · OUTSIDE THE TRANSACTION (UNLOCKED) Perform External Effect Stripe API / Cloud Deploy / Remote Bash via network (key: "pay-42") Crash here: rerun with same key "pay-42" PHASE 3 · TRANSACTION LINE Commit Outcome & Receipt Entry Marks task completed and appends result entry in one commit
    1. Commit Intent: The prepare phase generates an idempotency key derived from the task ID (pay-42) and commits it to disk.
    2. Perform External Effect: The charge phase executes the network call using the committed key. This runs outside the database transaction, keeping storage unlocked.
    3. Commit Outcome: When the external call succeeds, the phase writes the terminal outcome and appends a receipt entry in one commit.
    import { defineTask } from "@earendil-works/pi-durable";
    
    export const PaymentTask = defineTask<{ card: string }, any, any>({
      name: "shop.payment",
      version: 1,
      initial: () => ({ phase: "prepare" }),
      phases: {
        // 1. Commit Intent
        prepare: async (task, runtime, context) => {
          const key = `pay-${task.id}`;
          await runtime.commit(() => ({
            status: "running",
            checkpoint: { phase: "charge", key },
          }), context);
        },
    
        // 2. Perform External Effect (outside commit line)
        charge: async (task, runtime, context) => {
          const key = task.state.checkpoint.key;
          const receipt = await payments.charge(task.input.card, key);
    
          // 3. Commit Outcome with receipt entry
          await runtime.commit(async (tx) => {
            const entry = await tx.appendEntry(runtime.conversationId, {
              kind: "app.receipt",
              data: { receiptId: receipt.id },
            });
            return {
              status: "terminal",
              outcome: { status: "completed", result: { entryId: entry.id } },
            };
          }, context);
        },
      },
      abort: async (task, runtime, context) => {
        await payments.refund(task.state.checkpoint.key);
        await runtime.commit(() => ({
          status: "terminal",
          outcome: { status: "aborted", reason: "user" },
        }), context);
      },
    });

    Crash Recovery Taxonomy

    Process crashes divide into three recovery states:

    Crash PointState on Reopening StorageRecovery Action
    Before Step 1Checkpoint is prepareReruns prepare. No network call occurred.
    Between Step 1 and Step 3Checkpoint is charge with key: "pay-42"Reruns charge using the recorded key. The Stripe Idempotency API recognizes the key and rejects duplicate charges.
    After Step 3State is terminal with receiptSkips execution. Work is complete.

    Creating the idempotency key during the intent phase ties remote deduplication to the local database checkpoint. Generating the key inside the effect phase produces a new key on restart, breaking idempotency.

    The key has a shelf life. Stripe keeps idempotency keys for at least 24 hours, then prunes them. A task that crashes in charge and resumes three days later sends a key Stripe no longer remembers, and the charge runs again. Durable state on your side does not extend the provider’s memory. For effects that can sit for days, query the provider for an existing charge by your own reference before you retry.


    5. Tool Replay Policies: Safe vs. Unsafe Recovery

    Process crashes during tool calls break standard agents.

    Rerunning an active tool on restart triggers duplicate side effects, like double charges or repeat git commits. Dropping the active tool leaves the conversation broken. Pi-Durable requires every tool to declare a replay policy:

    Tool Replay Under Failure: Safe Rerun vs. Interrupted Yield Flowchart showing how process interruption splits based on replay policy between safe idempotent reruns and unsafe error yields with partial output. CRASH DURING TOOL EXECUTION Task in state: running · checkpoint: phase execute Reopen: Evaluate Replay Flag Harness resets task status to pending replay: safe replay: unsafe IDEMPOTENT READ OPERATION Reruns Tool From Scratch File search, knowledge retrieval, pure calculation Completes run and appends pi.tool-result State: terminal · outcome: completed MUTATION WITH SIDE EFFECTS Synthesizes Interrupted Error Card charge, container deploy, remote bash Packages buffered stdout from pi.live Yields control to model for next action
    import { Type } from "@earendil-works/pi-ai";
    import { defineTool } from "@earendil-works/pi-durable";
    
    const searchDocs = defineTool({
      name: "search_docs",
      description: "Search internal knowledge base",
      parameters: Type.Object({ query: Type.String() }),
      replay: "safe", // Safe to rerun on recovery
      execute: async (args) => {
        return { content: [{ type: "text", text: await index.query(args.query) }] };
      },
    });
    
    const executePayout = defineTool({
      name: "execute_payout",
      description: "Release funds to customer",
      parameters: Type.Object({ amount: Type.Number(), accountId: Type.String() }),
      replay: "unsafe", // Do not rerun after a crash
      execute: async (args) => {
        const receipt = await stripe.transfers.create({
          amount: args.amount,
          destination: args.accountId,
        });
        return { content: [{ type: "text", text: receipt.id }] };
      },
    });

    Before execute() runs, the harness commits an execution intent checkpoint to disk:

    {
      "phase": "execute",
      "arguments": { "amount": 250, "accountId": "acct_8821" },
      "replay": "unsafe"
    }

    When a crash occurs during tool execution:

    • replay: "safe": The scheduler reruns the tool function on restart.
    • replay: "unsafe": The scheduler refuses to rerun the tool. It synthesizes an interrupted result, packages whatever partial output was committed to pi.live before the crash, and hands control back to the model.

    Testing Crash Recovery Under SIGKILL

    Consider a test scenario where a kill -9 signal halts a process running two concurrent tool calls in one SQLite database:

    // The SQLite row left by the killed process
    {
      "id": 24,
      "kind": "pi.tool",
      "owner": 18,
      "state": {
        "status": "running",
        "checkpoint": {
          "phase": "execute",
          "arguments": { "target": "weekly" },
          "replay": "safe"
        }
      }
    }

    When a second process opens the database:

    1. Harness.open() runs a reconciliation commit: all tasks marked status: "running" return to status: "pending".
    2. The engine skips the unsafe deploy tool and appends an error entry: "[error] Tool deploy was interrupted and may have partially run".
    3. The engine reruns the safe fetch_report tool from scratch.

    The model receives the partial output, reads the interruption message, and decides how to proceed.

    This is the key design choice: the model, not the runtime, decides how to recover from an unsafe interruption. The runtime knows a deploy may have half-run. Only the model, with the transcript in front of it, can decide whether to check status, roll back, or ask the user.

    Unsafe tools also shape what happens below them. A tool can call other tools as nested calls, each its own pi.tool task. When the calling tool is not replay-safe, its unfinished nested calls and child tasks get abandonOnRestart: true by default. On the next start, the scheduler aborts them, with everything they own, before any of them runs again. You never get an orphaned child finishing a job its parent will never collect.


    6. The Replay Paradox: Why Workflow Engines Clash With Agents

    Engineers evaluating durable execution often consider Temporal.

    Temporal manages distributed microservices well. Running interactive, multi-turn agents on Temporal introduces architectural friction.

    The two systems achieve durability through contrasting models:

    Event Sourcing History Tape vs. Checkpoint Slabs Isometric architectural comparison showing Temporal replaying historical event blocks from line 1 versus Pi-Durable landing on the active checkpoint slab. TEMPORAL PARADIGM · EVENT SOURCING Linear Event Log (Replay From Line 1) Event 01: Line 1 Events 02..14 Event 15: Crash Replays all Constraint: Strict Code Determinism Worker restarts and re-runs workflow from line 1. Editing a prompt or changing an if branch diverges from history, throwing non-deterministic fault. PI-DURABLE PARADIGM · STATE MACHINE Checkpoint State Machine 1. prepare 2. execute 3. checkpoint 15 Direct resume no re-execution Advantage: Zero Replay Ceremony Worker loads active checkpoint from SQLite disk. Runs next phase with current registry code. Hot-reloading prompts and tools works natively.
    Architectural DimensionTemporal Workflow EnginePi-Durable Harness
    Durability TechniqueEvent Sourcing: Re-runs workflow code from line 1, mocking completed activities.Checkpoint State Machine: Never replays past code; loads the latest committed task state.
    Code DeterminismMandatory. Calling Date.now(), random IDs, or changing an if block breaks history replay.Not required. Code can change between turns. Only inputs, checkpoints, and outputs persist.
    Code Hot-ReloadingComplex. Requires patch markers (patched() in TypeScript, getVersion() in Java and Go) or Worker Versioning.Native. Swap the extension definition in the registry; the next phase runs the new code.
    Token StreamingInefficient. Pushing 500 token chunks through Temporal history hits payload and event limits.Throttled WAL commits. Flushes token and tool stdout chunks every 100ms into a live document.
    Runtime InfrastructureExternal cluster (Temporal Server + Postgres/Cassandra + gRPC workers).Embedded engine. Runs inside Node, Bun, or Cloudflare Durable Objects over SQLite WAL.

    Temporal enforces deterministic execution. If an agent executes tool calls for 40 minutes and the worker restarts, Temporal replays the workflow function from line 1. Every code branch must match the recorded event history.

    Updating a prompt template, editing a system instruction, or modifying a tool definition during execution causes the replayed code to diverge from history. Temporal aborts with a non-deterministic workflow error.

    Pi-Durable stores explicit checkpoints instead of replaying code. On restart, the engine loads the active checkpoint from SQLite, resolves the tool registry, and runs the next phase. Past code does not run again.

    The Limits That Hit Agents First

    Temporal’s hard limits were set for business workflows, not chatty LLM loops:

    LimitValueWhat it means for an agent
    Payload size2 MB per input or resultA long tool output or a file read must be offloaded
    Event history51,200 events or 50 MBWarnings start at 10,240 events or 10 MB
    Transaction / gRPC message4 MBCaps a batch of tool results in one step

    Temporal’s answer for big payloads is the claim-check pattern: write the blob to S3 and keep only a reference in history.

    Run the numbers on streaming. A 4,000-token answer streamed over 40 seconds, recorded every 100 ms, is 400 updates. Store each as an event and you hit the 51,200-event ceiling after 128 answers. A busy support conversation gets there in a day. So in practice teams stream tokens through a side channel (Redis, a WebSocket hub) and store only the final message in Temporal, which brings back the exact “user saw it, storage did not” gap from section 2.

    Pi-Durable avoids the problem by not keeping history for live state. Partial output goes into pi.live, a document with history: "latest". The harness flushes it at most every 100 ms with one commit in flight, and older deltas are deleted when a new base lands. The transcript gets one pi.assistant entry per response.

    Where the Other Engines Sit

    Durable execution engines split into three camps by what they store:

    ModelEnginesWhat a restart runs againCode-change cost
    Event sourcingTemporal, RestateWorkflow code from the top, with recorded results fed backMust stay deterministic; patch every change
    Step journalingDBOS, Inngest, Cloudflare WorkflowsFunction from the top; finished steps return saved outputStep order must stay stable, or version the workflow
    State checkpointsLangGraph, Pi-DurableNothing that finished; resume at the saved checkpointCode can change between steps

    DBOS is the closest relative in spirit: a library, not a cluster, that writes each step’s output to Postgres and skips finished steps on recovery. But it still re-enters the workflow function from the top, so step order must match. LangGraph checkpoints graph state after every super-step, which is close to Pi-Durable’s model. The difference is scope: Pi-Durable also commits each tool call’s intent before it runs, owns the scheduler, and ties the transcript, documents, and tasks to one commit line.

    Pi-Durable sits at the far end. It stores the state machine, not the function call stack. That is why live code swaps work, and also why you must write every phase so it can start cold from its checkpoint.


    7. The Step Check Engine and Live Code Handover

    Reloading extension code during live runs requires isolated execution phases.

    In Pi-Durable, each in-memory task run is an invocation. Invocations execute sequential phases. Between phases, the harness evaluates a check called a step.

    The step executes as a commit on the Session write queue, verifying all changes committed by the prior phase before scheduling the next.

    The Step Check Engine and Live Code Handover Decision flowchart showing the six prioritized rules evaluated before each task phase execution, including live code handover and fault detection. Phase Returns Committed runtime data THE STEP EVALUATION ENGINE (PRIORITY ORDER) 1. Outcome committed or waiting? → Stop invocation 2. Harness is closing? → Keep checkpoint 3. Abort mark requested? → Schedule abort 4. Phase handler threw error? → Write faulted 5. Checkpoint changed? New registry code? → Handover to new definition → Next phase 6. Checkpoint unchanged? → Write faulted Rule: Unchanged checkpoint halts loop

    Preventing Infinite Loops With Step Faults

    If a phase returns with the same checkpoint it started with, it loops on the same inputs.

    The step compares old and new checkpoints as JSON, ignoring key order. When checkpoints match, the harness marks the task faulted: "Task phase returned without durable progress".

    Committing entries or documents does not satisfy this check. Only changing the checkpoint advances task state.

    Live Code Handover

    At each phase boundary, the step inspects the in-memory registry:

    • Updating an extension file on disk triggers registry.install(NewExtension).
    • The step detects the updated definition for that task kind.
    • The task hands over: the engine returns it to pending with its current checkpoint, and the next scheduler pass executes the phase using the new definition.
    • Predecessor and successor execution handlers never overlap.

    Handling Missing Extensions via the Blocked State

    Uninstalling an extension or deploying a task version without a migration script pauses execution:

    • The harness does not terminate tasks when code is missing.
    • The task enters a blocked runtime state (missing_task, task_too_old, or migration_failed).
    • Storage retains the task in a pending state.
    • When you register the matching extension, the scheduler detects the definition and resumes execution.

    8. Hierarchical Concurrency: Bottom-Up Abort and Fail-Fast

    Most multi-agent implementations dispatch subagents as detached background async promises. If a user presses Esc or a payment fails, the parent halts, but child subagents continue burning API credits in the background.

    Pi-Durable enforces hierarchical task ownership:

    Hierarchical Task Trees: Fail-Fast Branching and Bottom-Up Teardown Architecture diagram showing parent Checkout task waiting on child payments with failFast cancellation and reverse bottom-up abort execution. PARENT TASK · SHOP.CHECKOUT Status: waiting on [8, 9, 10] Policy: failFast · Checkpoint: decide CHILD 8 · PAYMENT A Card A ($50) Abort mark injected CHILD 9 · PAYMENT B Card B ($100) FAILED: DECLINED Triggers failFast CHILD 10 · PAYMENT C Card C ($25) Abort mark injected TEARDOWN SEQUENCE · BOTTOM-UP CLEANUP 1. Children abort handlers run in parallel: refund(pay-8) and refund(pay-10) 2. Children commit terminal outcomes and retire child documents 3. Parent abort() handler runs ONLY after all child tasks settle terminal
    import { defineTask, type TaskId } from "@earendil-works/pi-durable";
    
    type CheckoutState =
      | { phase: "pay" }
      | { phase: "decide"; payments: TaskId<string>[] };
    
    const Checkout = defineTask<{ cards: string[] }, CheckoutState, string>({
      name: "shop.checkout",
      version: 1,
      initial: () => ({ phase: "pay" }),
      phases: {
        pay: async (task, runtime, context) => {
          await runtime.commit(async (tx) => {
            const payments: TaskId<string>[] = [];
            for (const card of task.input.cards) {
              payments.push(
                await tx.createTask(Payment, { card }, {
                  ownership: { kind: "task", taskId: task.id }, // Child task
                })
              );
            }
            // Stop running code until all payments finish.
            // The first failure aborts sibling payments.
            return {
              status: "waiting",
              checkpoint: { phase: "decide", payments },
              on: payments,
              policy: "failFast",
            };
          }, context);
        },
        decide: async (task, runtime, context) => {
          const outcomes = await runtime.outcomes(task.state.checkpoint.payments, context);
          const paid = outcomes.every((outcome) => outcome.status === "completed");
          await runtime.commit(() => ({
            status: "terminal",
            outcome: paid
              ? { status: "completed", result: "Order placed." }
              : { status: "failed", error: { message: "Payment failed." } },
          }), context);
        },
      },
      abort: async (task, runtime, context) => {
        // Runs after child payments finish aborting and refunding
        await runtime.commit(() => ({
          status: "terminal",
          outcome: { status: "aborted" },
        }), context);
      },
    });

    Two rules govern the task tree:

    1. failFast vs. allSettled: In policy: "failFast", a failure in child task #2 marks siblings #1, #3, and #4 for cancellation.
    2. Bottom-Up Abort Order: When you abort a task or conversation, the cancel signal cascades to leaf tasks first. A parent task’s abort() handler runs only after child tasks reach a terminal state.

    Cascading cancellations down to leaf tasks releases child holds, such as room bookings or card authorizations, before the parent marks the workflow failed.

    Deep trees used to be slow. The first scheduler visited every live task on each pass. The current one keeps in-memory ownership indexes and starts abort cascades only from tasks with cancellation intent. The changelog reports the gains:

    WorkloadBeforeAfter
    Chain of 500 owned tasks settling12 s50 ms
    Chain of 20,000 owned tasks settlingn/a1.5 s
    One tool making 8,000 parallel nested calls18.8 s2.7 s
    Tasks kept in memory after 10,000 ended subagent calls10,0000

    The last row matters for long sessions. A parent that spawns subagents all day no longer grows its memory with every ended child.


    9. The Semantic Inbox: Steers, Follow-Ups, and Passive Writes

    Users do not wait for an agent to complete a 60-second tool turn before sending input. They interrupt.

    Standard architectures drop incoming messages or cancel running turns when users type. Pi-Durable routes incoming messages through a structured inbox (pi.inbox) supporting three operations:

    The Semantic Inbox: Three Ingress Routing Paths Flowchart tracing how user messages route between mid-turn steers, post-run follow-ups, and non-conversational passive writes. INGRESS QUEUE Client Input Submitted while run is active whenBusy: steer Course correction whenBusy: followUp Subsequent task type: write Audit / data sync PATH 1 · STEER INJECTION Injected at Active Tool Round Boundary Inserts into transcript between tool rounds Model reads steer alongside latest tool result PATH 2 · QUEUED FOLLOW-UP Held in Inbox Until Run Completes Preserved in pi.inbox during active turn Triggers successor turn once conversation goes idle PATH 3 · PASSIVE TRANSCRIPT WRITE Direct Append Bypassing LLM Appends structural entry (telemetry, audit, sync) Placed with zero token expenditure
    // 1. STEER: Injected after the active tool round finishes.
    // Joins the running work to redirect the current turn.
    await root.submit(
      {
        type: "input",
        content: "Don't use Postgres; use SQLite instead",
        whenBusy: "steer",
      },
      context
    );
    
    // 2. FOLLOW-UP: Queued until the agent finishes its current response.
    // Starts the next turn once idle.
    await root.submit(
      {
        type: "input",
        content: "Now write the unit tests",
        whenBusy: "followUp",
      },
      context
    );
    
    // 3. WRITE: Appends an entry to the transcript without invoking the LLM.
    // Used for audit markers, file drop events, or CRM status sync.
    await root.submit(
      {
        type: "write",
        entry: { kind: "app.audit", data: { userTier: "enterprise" } },
      },
      context
    );

    There is a fourth mode: whenBusy: "reject" throws ConversationBusy instead of queueing, for callers like cron jobs that should never stack work. By default the inbox places one steer or follow-up per turn; setting steeringMode: "all" or followUpMode: "all" places every queued item at once. If a run fails, queued items stay in the inbox until the next submission places them, oldest first.

    When a tool finishes, the harness inspects the inbox before scheduling the next generation task. If a steer is pending, the engine injects the input into the ongoing run. The model receives the correction alongside the latest tool output, updating its trajectory without discarding prior execution.


    10. Two-Tier Compaction: Active Head vs. Permanent Audit

    Enterprise agents manage competing requirements: models have finite context windows, but compliance regulations (SOC 2, HIPAA, FINRA) require complete audit retention.

    Pi-Durable decouples the model context window from the underlying storage transcript.

    Two-Tier Compaction: Active Context Projection vs. Immutable Storage Storage architecture diagram showing raw append-only SQLite log entries preserved permanently while the active context window reads only from the compaction head. PHYSICAL STORAGE · APPEND-ONLY LOG (SQLITE WAL) Entry 1 User Entry 2 Assistant Entry 3 ToolResult Entry 4: Compact head pointer: 3 Entry 5 User Entry 6 Assistant DUAL THRESHOLD GATES · backgroundTokens (32,768): Summarizes in background while turn continues · reserveTokens (16,384): Pauses generation near limit until summary commits ACTIVE MODEL CONTEXT WINDOW Model sees: summary + Entry 3 onward Entries 1..2 stay on disk but leave the context window Result: Zero token bloat · Raw audit trail permanently intact
    1. Dual Compaction Thresholds:
      • backgroundTokens (32,768 tokens below limit): A background compaction task summarizes older turns while user interactions proceed.
      • reserveTokens (16,384 tokens below limit): If background summarization has not completed and the context window nears exhaustion, the next turn pauses until the summary settles.
    2. Head Pointers (head: EntryId): Compaction writes a pi.compaction entry whose head references the oldest retained entry.
    3. Audit Preservation: Context reduction supplies entries from head forward to the model, warm with Anthropic Prompt Caching. The underlying database retains every raw entry from entry #1. Historical turns remain accessible through cursor queries (tx.scanEntries()).
    4. Recent Context Stays Verbatim: keepRecentTokens (default 20,000) sets roughly how much of the newest context the summary leaves untouched.
    5. Overflow Recovery: When a provider rejects a request because the context is too long, generation compacts and retries once.
    6. Stale Summaries: A background summary that would cut before the start of the current context settles as stale. When several compactions are in flight, the furthest cut wins.

    The cache detail is easy to miss. Each conversation stores a UUIDv7 in pi.provider and sends it as the provider sessionId. It survives reopen, retries, compaction, and model changes, so a restarted worker lands on the same prompt cache. With Anthropic, a cache read costs 10% of the base input price. A worker that crashes and resumes with a fresh session ID pays full price to rebuild a 150,000-token prefix; one that keeps the ID pays a tenth.

    For a hard break, reset() starts a new context. It can take a handoff note ("We were fixing the flaky login test. Continue."), and a tool can request one with control: { handoff: "..." }. The model stops seeing older entries; storage keeps them.


    11. Decision Framework

    Choose Pi-Durable when building:

    • Customer agents where process restarts must preserve conversational state
    • Terminal tools or edge services requiring an embedded footprint
    • Interactive systems needing live token streaming and mid-turn steering
    • Workloads targeting single-tenant SQLite databases or Cloudflare Durable Objects

    Choose DBOS or Inngest when building:

    • Agents that already live next to a Postgres database or a serverless function host
    • Step pipelines where step order rarely changes and you want each step’s output queryable as rows

    Choose Temporal when building:

    • Multi-service orchestrations spanning distributed databases and external teams
    • Asynchronous workflows that pause for weeks awaiting human approval
    • Batch billing, data synchronization, and server provisioning pipelines

    The Bottom Line

    Production AI agents demand systems engineering, not prompt tweaks.

    In-memory loops fail under production traffic. Reliable agent runtimes require:

    • An atomic commit line that persists state before publishing output
    • An effect sandwich pattern isolating network calls from storage locks
    • Intent checkpoints paired with explicit tool replay rules
    • Hierarchical task trees with bottom-up abort cascades
    • A step verification engine that detects stalled loops and handles live code swaps
    • A semantic inbox routing steers, follow-ups, and background writes
    • Compaction that trims the active context window while retaining audit history

    Pi-Durable demonstrates that reliable execution does not require a distributed server cluster. You can host crash-resilient, multi-conversation agents within an embedded library of about 22,500 lines on SQLite.


    Building durable agent architectures or wrestling with state management in production LLM apps? I would love to hear what patterns you are using. Reach out on LinkedIn.