Definition
Durable execution persists enough execution progress for unfinished work to resume after a process, machine, or network interruption. The runtime records completed steps, pending work, and recovery state, then uses an explicit replay policy when execution restarts.
For an agent, this can include model requests, tool calls, queued messages, and delegated tasks. Saving the conversation is useful, but recovery also needs to know which actions completed and which effects remain uncertain.
Origin and agent usage
Durable execution was established in workflow systems before its recent use in agent harnesses. Tom Wheeler's May 2025 Temporal explanation describes persisted progress and recovery after failure. It documents the mechanism and terminology; the first coinage is not established by that article.
Earendil's October 1, 2026 Pi Durable announcement applies these recovery ideas to agent harnesses. It describes model requests and tool calls as checkpointed tasks. An interrupted model request can be repeated; an interrupted tool call is repeated only when its replay declaration permits it. Otherwise the agent receives an interruption report. The package is experimental.
The same announcement uses a request identifier to deduplicate submissions. That guarantee applies to the submission boundary it describes. It does not imply that every external action occurs exactly once.
Scope of the guarantee
Products differ in what they persist and replay. A workflow runtime's guarantee about its recorded execution does not automatically cover a remote payment or deployment. Read the recovery policy at each boundary before relying on a claim of exactly-once execution.
Operational significance
A read-only lookup can often be repeated safely. A payment, deployment, or message may already have taken effect before the runtime lost its acknowledgement. Repeating it without a deduplication key or an authoritative status check can cause a second effect.
Define recovery per operation: what is persisted before execution, how completion is recorded, which actions may be replayed, and how uncertain completion is reconciled. A durable checkpoint cannot undo an irreversible external action or prove that the action was authorized.
Durability also depends on storage. State held only in the process's memory cannot survive that process being lost. Recovery claims should name the failures, storage guarantees, and tool contracts they cover.
Distinguish it from nearby terms
- Durable memory preserves knowledge between runs. Durable execution preserves the progress and recovery obligations of work.
- A checkpoint is saved state. Durability requires the runtime to use that state correctly after failure.
- Retrying repeats an operation. Durable execution decides what should be resumed or repeated using persisted progress.
- Rollback restores an earlier state. Resumption continues unfinished work and may require compensation for effects already produced.
Check your understanding
An agent loses its connection immediately after requesting a deployment. What records or external checks would let it distinguish a completed deployment from a request that never arrived, and when would a retry be safe?