An agent can stop, fail, or be replaced without losing the work if state lives in evidence-bearing artifacts.
The sentence “the worker says it finished” carries almost no weight in my systems. A summary is a lead. It points toward output worth inspecting. It doesn't establish that a file contains the change, an external system accepted the update, or a command produced the claimed result. I've received plenty of cheerful summaries from work that wasn't done.
This becomes critical when work spans more than one agent or run. Conversation is temporary. Processes end. Context windows compress. None of that cares that we're almost finished. The durable artifact has to preserve enough state for another worker to inspect what happened and continue without trusting a story about it.
Checkpoint shape
A checkpoint should answer practical questions: what unit of work is this, what state is it in, what inputs were used, what outputs exist, and what remains? I keep those fields structured rather than burying them in a paragraph. The next worker shouldn't have to interpret my prose mood.
The state vocabulary is small and explicit. planned, in_progress, awaiting_verification, accepted, and failed are more useful than a percent-complete guess. The checkpoint also records the current attempt and the location of artifacts. A replacement worker can read it before touching the task.
The checkpoint isn't a transcript. Saving every thought produces a large artifact with very little authority. It should contain decisions, identities, references, and conditions needed to resume. Logs can remain attached for diagnosis, but they don't substitute for the current state.
Writing checkpoints at meaningful boundaries matters too. If the worker updates state before the output is durable, a crash can leave “awaiting verification” pointing at nothing. That's a lovely state name with no state behind it. The artifact should be written first, then the state should reference it. Ordering is part of the mechanism.
Provenance
Every meaningful result needs a path back to its inputs and method. For a code change, that includes the candidate revision, files changed, and commands used for verification. For collected evidence, it includes the source anchors and observation time. For a generated document, it includes the source material it was based on.
This is how read-back becomes possible. Another worker doesn't accept “updated configuration” from the summary. It reads the stored configuration and compares it with the requested change. If the claim concerns execution, it reruns the relevant check or inspects the external result.
Provenance also limits accidental reuse. A test result attached to one candidate can't silently validate a later candidate. An evidence snapshot with an old observation time can't present itself as current. Context is useful only when its scope survives along with its content. Otherwise it's just a plausible old fact.
Side-effect state
External effects need their own record because retries can be dangerous. A worker preparing a change is different from a worker attempting it, and an attempted change is different from a confirmed one. Collapsing those states invites duplicate execution, and external systems don't grade retries on intent.
Before an external action, I record the intent and a stable operation identity. After the attempt, the checkpoint records the response or the fact that the result is uncertain. If the worker loses its connection after sending a request, the next worker should see outcome_unknown, not assume failure and send again.
Confirmation has to come from the external system or a reliable read-back. Agent memory that it called a tool isn't enough. For a file, read the file. For a submitted record, retrieve the record by its stable identity. For a command, preserve the output and exit status.
Uncertainty is a valid state. It prevents a pleasant fiction from becoming a second side effect.
Resumability
Resuming work means selecting the next safe transition from durable state. It doesn't mean asking a new worker to infer the whole plan from the last summary.
Suppose a worker produced an artifact and stopped before verification. The checkpoint identifies the artifact, inputs, and required acceptance check. A replacement can start at verification. It doesn't regenerate the artifact, because regeneration would erase useful evidence and potentially introduce a different result.
If verification fails, the worker records the failure against that artifact and creates a new attempt. The old output remains identifiable. This gives the task a history without making the history authoritative. Only the currently accepted artifact can satisfy downstream work.
Resumability also depends on environmental assumptions. If a step requires a particular candidate revision or source snapshot, the next worker confirms that condition before continuing. When it can't recreate the context, it stops with a precise mismatch rather than improvising from nearby state.
Idempotency
Idempotency is often described as “safe to run twice,” but the useful implementation begins with identity. The same logical operation needs the same key across retries. The system can then return the existing result, continue an incomplete attempt, or refuse an ambiguous duplicate.
Not every action is naturally idempotent. “Just run it again” isn't a strategy for sending things. In those cases, the checkpoint and read-back path must prevent blind repetition. An outcome_unknown action should be reconciled before any retry. A completed action should expose the external proof that closes it. A failed action should say whether another attempt is permitted.
This is also why worker summaries stay outside the acceptance boundary. A summary may say “done” after a tool call returned. The durable state may show that the response was never read back. Acceptance happens only when the artifact or external effect is verified and the checkpoint moves to accepted with that evidence attached.
The mechanism is intentionally indifferent to which worker returns. A fresh process can read the checkpoint, resolve the referenced artifacts, inspect side-effect state, and choose the next permitted transition. If it cannot do that, the previous worker preserved a narrative of the work, not the context required to continue it.