# Video pipeline recovery kit

Recover a stalled multi-step generative production (video, or any pipeline whose steps have
individual status) without rebuilding the parts that already work. This kit is the reusable
version of the workflow described in
[118 Clips Were Ready. The Documentary Still Would Not Publish.](https://kittyclaw.dev/recover-a-stalled-video-render-without-rebuilding-it),
a recovery coordinated through KittyClaw for a stalled long-form AI-generated documentary
(originally scoped at about 20 minutes; the finished film runs about 12 and a half minutes).

## What this kit assumes

- A KittyClaw project holding the ticket for the production itself (the "source ticket").
- A pipeline whose individual steps expose their own status (done, waiting, failed) instead of
  only one overall pass or fail for the whole run.
- A way to re-run an individual step or a small group of steps without re-running the rest.
- A defined set of checks the finished output has to pass before a human is asked to approve it.

## Files

| File | Purpose |
|---|---|
| `recovery-ticket.template.md` | Open this as a new ticket when a production breaks partway through. It only holds the diagnosis: what is valid, what failed, what depends on what, and the exact retry scope. |
| `preserve-retry-wait-validate-matrix.md` | The decision rule applied to every step during the real recovery. Use it to classify each step instead of relaunching everything. |
| `master-qc-checklist.md` | The checks the finished video had to pass before the human approval step. Adapt the specific numbers to your own format; keep the principle: check before you ask. |
| `schedule-resume-after-quota.example.json` | An illustrative call to KittyClaw's `PATCH /tickets/{id}/schedule` endpoint (`fireAt` / `author` / `targetStatus`) showing how to make a ticket resume automatically once an external quota or dependency clears. The dates in this file are fictional; they do not reproduce the real recovery's actual schedule calls. |
| `build-script.example.mjs` | A cleaned, trimmed adaptation of the real production script's recovery-relevant logic: aggregating step status, retrying only failed branches while preserving valid ones, and rebuilding a merge cascade. All local paths, internal service URLs and the original 23-scene creative content have been removed; two short placeholder scenes stand in for illustration. This is not the full original file. |
| `node-manifest.example.json` | A trimmed excerpt of the real node-tracking manifest's structure: how a production keeps track of which node produced which output, so a retry can target exactly the right nodes. |
| `delivery-comment.template.md` | The final comment template used to confirm, in the source ticket, that the gate cleared, the publish step ran, the callback fired, and the public URL is live. |

## How the real recovery used these ideas

1. A new ticket held the diagnosis for the whole broken run: which of roughly 300 production
   steps were valid, waiting, or failed, and why.
2. Only the steps named as failed were retried. Everything already valid was left untouched in
   the same production graph, no full rebuild.
3. A provider quota was tracked as an explicit wait on the ticket, not as a failure, and the
   ticket only moved forward once the quota was actually lifted.
4. The finished file was checked against the quality checklist before anyone was asked to
   approve it for publishing.
5. Once approved, the publish step reported back to the original production ticket with the
   result and the public link.

## What is deliberately not included

No local file paths, no internal service hostnames or ports, no account identifiers, no OAuth
or authorization details, and no unreduced proprietary creative content (the original narration
script and full 23-scene shot list) are included. `build-script.example.mjs` keeps the
structure of the real recovery logic with placeholder scenes and placeholder `.example.invalid`
URLs in place of the real local addresses; it is meant to be read and adapted, not run against a
real Kinoboard instance without changes.

## Limits

- The decision matrix and checklist reflect one real recovery. Extend them for whatever your
  own pipeline's steps and output format actually require.
- Not every failure has a code-level fix behind it. See the use-case page's "What actually got
  fixed, and what only got worked around" section before assuming a recovery playbook alone is
  a permanent solution.
