# Preserve / retry / wait / validate matrix

Apply this to every step of a broken production run during the diagnosis phase, before
retrying anything. It is the decision rule behind the real recovery described in the
[recovery use case](https://kittyclaw.dev/recover-a-stalled-video-render-without-rebuilding-it).

| Step state | What it means | Action | Real example |
|---|---|---|---|
| **Preserve** | The step's output is done and correct. Nothing about the failure touches it. | Leave it exactly as is. Do not re-run it, even if it is upstream of a failure. | 89 clips and 107 images already correct were left untouched during the recovery. |
| **Retry** | The step itself failed, or its direct input was wrong (wrong format, bad reference, corrupted output). | Re-run only that step, or the smallest group of steps that share the same broken input. | 18 clips regenerated because their reference frame was wrong; 11 images regenerated because they had failed outright. |
| **Wait** | The step cannot proceed because of something outside the pipeline's control: a provider quota, a rate limit, an external dependency, a scheduled resource. | Do not retry on a fixed interval and do not treat it as failed. Record the wait condition on the ticket and schedule a resume for when it should be checked again. | A daily generation quota blocked several clips; the ticket only resumed once the owner raised the quota from 20 to 40 generations per day. |
| **Validate** | The step technically succeeded, but the result needs a human or an independent check before it can be trusted (a claim, a legal detail, a creative decision). | Route it to a review step. Do not let it advance past a gate on its own. | The narration text needed an independent fact-check pass before generation, not after. |

## How to use this during a diagnosis

1. List every step in the broken run with its current status.
2. For each failed or blocked step, ask: is the step itself broken, or is it waiting on
   something outside the pipeline? That single question sorts "retry" from "wait."
3. For everything else, the default is "preserve." Only demote a step out of "preserve" if you
   have a specific reason tied to the actual failure, not a general feeling that a full rebuild
   would be safer.
4. Anything that reaches a publish-facing gate goes through "validate" once, regardless of how
   it got there.

## Anti-pattern to avoid

Relaunching an entire production because a handful of steps failed risks being slower and more
costly than diagnosing the actual scope of the failure. In the real recovery, retrying 18 clips
and 11 images out of a graph of nearly 300 steps resolved the incident; relaunching the whole
production would have discarded weeks of already-correct, already-completed generation work.
