Useful automation needs a represented stop state, a retry budget, and a deliberate way back into motion.
Recurring automation earns trust in an unglamorous way: it runs into something outside its authority and stops cleanly. Mine can inspect infrastructure or classify support work. It can gather evidence and prepare the next step. It can't restart a service or send a customer reply just because those actions look likely to help.
That boundary can't live only in prompt prose or a comment beside the scheduler. “Ask before acting” is a wish until the workflow has somewhere to wait, something to record, and a specific transition that requires approval.
A stop is a state
I model stopping as part of the workflow rather than as failure around the workflow. A run might move from observing to classified, then to ready_for_review. It doesn't drift from classification into mutation because a generated explanation sounds confident. Models are rarely short on confidence.
The state needs durable data. What was observed? When was it observed? Which action is proposed? Why does that action require authority? An operator should be able to inspect the record without reconstructing the run from a stream of logs. I've done that reconstruction. It's not a feature. The proposed customer response or restart action belongs with the evidence, but it's still a proposal.
This matters most during ordinary operation. If approval is represented only by an emergency dialog, the happy path will eventually route around it. A first-class awaiting_approval state makes the boundary visible in queues, reports, and tests. It also prevents a later run from mistaking unfinished work for a fresh task.
The same design handles exceptions better. An item that can't be classified should enter an exception state with the evidence collected so far. It shouldn't be shoved into the nearest category just to keep the automation moving. “Needs eyes” is useful information when it's explicit and owned.
Retries need a budget
Stopping rules aren't limited to high-consequence actions. They also govern persistence. A retry makes sense when the failure is plausibly temporary and the operation is safe to repeat. It becomes noise when the same condition keeps returning and nothing about the next attempt will be different. The sixth identical retry isn't more determined than the fifth.
I give recurring work a bounded retry policy: a defined class of retryable failures, a delay, and a maximum attempt count. The count is stored with the item, not held only in the memory of one worker. When the budget is exhausted, the state changes. The scheduler can't reset the story by discovering the same item again.
The distinction between retryable and terminal failures matters. A brief network failure may justify another read. Missing authority doesn't. Neither does an ambiguous classification that requires judgment. Retrying those conditions turns a decision problem into a resource-consumption problem.
The budget should follow the item across scheduler runs. Counting attempts only inside one process creates a loophole: every fresh invocation believes it is making a first retry. I store the next eligible time as well, so two workers cannot both decide the delay has elapsed and perform the same inspection. The mechanism remains modest, but the boundary survives concurrency and restarts.
There is a subtle trap in “safe” read-only work too. Inspection can be repeated without changing production, but repeated alerts, duplicate review items, and endlessly regenerated drafts still create side effects for people. Bounded retries keep the automation from becoming the loudest broken thing in the room.
When a run exhausts its budget, it should leave a compact account: attempts made, last evidence, and the condition required to resume. That is much more useful than a generic failure mark. It tells the next operator whether to approve, correct input, wait for recovery, or close the item.
Approval is a transition
Approval has to authorize a particular move, not bless the automation in general. I'm approving this action against this evidence. I'm not handing out a hall pass. For a proposed restart, the approved object should identify the target action and the evidence on which it was based. For a prepared customer reply, approval should apply to the exact draft or clearly invalidate when the draft changes.
That makes the transition testable. An item in awaiting_approval may move to approved only when an authorized actor approves the current proposal. Execution can then move it to completed or to an evidence-rich failure state. Rejection is also a real transition. It may return the item for revision or close it without action, but it should not look like a missing approval.
Freshness belongs in this path. Infrastructure evidence can age while a request waits. If the observed condition isn't current anymore, approval shouldn't release an old action against a changed system. The workflow can require another inspection and generate a new proposal. That inconvenience is cheaper than executing a stale remedy.
Approval records also need an actor and a time. Those fields are operational inputs, not audit decoration. They let execution confirm that the approving identity had the required authority and that the approval applies to the current proposal. If either check fails, the item stays where it is and explains why.
I trust this arrangement because the automation cannot narrate its way past the boundary. The scheduler may wake again, a different worker may pick up the record, and a model may produce a more persuasive recommendation. None of those events supplies authority. Only the recorded approval transition does.
The final test is easy to state and worth running: place an item at every stop-line, restart the worker, and let the recurring schedule fire again. The item should remain stopped, with its retry history and proposal intact, until the required approval or corrective input arrives. If another run can silently carry it forward, the stop was documentation, not mechanism.