← Back to kittyclaw.dev
I use KittyClaw to… recover a stalled video render

118 Clips Were Ready. The Documentary Still Would Not Publish.

A small studio called Bloomii was building a long-form documentary about a French village that became self-sufficient in electricity, generated almost entirely through an automated video pipeline: still images turned into short motion clips, clips merged into longer segments, then one final assembly pass with narration and subtitles. The brief had originally targeted about 20 minutes; the finished film runs about 12 and a half. By late July the shot list was down to a clean 118 clips. Then the production missed its release date, and a status check showed real breakage: eight merge steps had failed outright, and eleven clips were stuck behind eleven failed images further upstream.

Rebuilding a hundreds-of-step production from zero would have thrown away weeks of completed generation work to fix a handful of broken joints. KittyClaw kept the production ticket as the one place to track what was actually broken, retried only the failing steps, tracked a generation-quota wait as a wait rather than a failure, and held a quality gate before the owner's final go-ahead. The video went live and verified on YouTube about five weeks after the original build began, with a confirmation sent straight back to the ticket that started it.

196already-valid clips and images left untouched during the recovery (89 clips + 107 images)
3×final-assembly attempts that failed at almost exactly the same six-minute mark
0failed steps left in the finished 298-step production graph

Why 118 finished clips were not a finished film

Bloomii generates video through Kinoboard, a node-based production engine: one workflow request holds every step of a video from the first reference image to the finished file, and each step tracks its own status (queued, done, failed). For the Ungersheim documentary, that meant a single request eventually holding 298 steps chained together: reference images, six-second motion clips, a cascade of merge steps that combine six clips at a time into longer stretches, one final assembly pass with narration and subtitles, a quality gate, and a publish step.

On 22 July, removing duplicate shots brought the plan down to 118 distinct clips across 23 narrated scenes, the number that later gave this story its ticket title. Over the following weeks the release date moved twice, from 31 July to 2 August and then to 16 August, mostly because the outside service generating images and clips enforces a daily generation quota, not because anything was broken; the owner raised that quota from 20 to 40 generations per day on 11 August. Then, on 16 August, the release date passed with nothing published. A diagnosis the next day found the real problem: eight merge steps had failed because a batch of clips generated in portrait orientation had ended up mixed in with the landscape clips the rest of the film used, and eleven clips were still stuck behind eleven failed images further upstream.

A second, smaller retry pass later that same day pushed the shot list to what the ticket records as 118 finished images and 118 finished clips. That is a real, dated checkpoint from 17 August, not the final count in the film that eventually published, since more clips were generated after that point. Either way, the formatting problem was now fixed. Once the merge cascade finished and the picture format was confirmed consistent, the final assembly step, the single pass that stitches every merged segment into one file with narration and subtitles, still failed three separate times on 18 August, each one stopping at almost exactly six minutes into rendering, with an orphaned encoding process still visibly running in the background after the failure was declared. Every clip going in was now correct. The failure had moved from the media to something else entirely: a fixed time ceiling on the machine doing the rendering, silently capping a step that a video this long needed far more time to finish.

What KittyClaw coordinated

KittyClaw did not generate a single frame of the film. Its role was to keep one ticket as the source of truth for what was actually happening across a production spread over five weeks and hundreds of steps, and to make sure nothing shipped in public until it had actually been checked.

01Diagnose
02Retry only failures
03Wait out the quota
04Check the master
05Gate + publish + confirm

1. Diagnose the graph before touching it

On 17 August the team opened a separate recovery ticket (bloomii #1442) with one purpose: classify every one of the roughly 300 production steps as valid, waiting, or failed, and why. Keeping that classification in one ticket instead of scattered across logs is what turned "the video is broken" into a scoped fix: 18 clips and 11 images to redo, not a full rebuild.

2. Retry only what failed

Only the identified failing steps were relaunched: 18 clips and 11 images. 89 clips and 107 images that were already correct were left completely untouched, in place, in the same production graph. That distinction between "broken" and "already good" is the difference between a targeted fix and throwing away weeks of completed generation work.

Decision rule: a step only gets retried if the diagnosis names it as failed or blocked by a specific upstream failure. Everything else is preserved as is.

3. Treat a quota wait as a wait, not a failure

Earlier in the same production, the release date had already slipped from 31 July to 2 August to 16 August because the outside image and clip generation provider caps how much can be produced per day. The workflow tracked that limit as an explicit waiting state tied to the ticket, so a paused production was never mistaken for a broken one, and the resumption happened once the owner raised the daily quota from 20 to 40 generations on 11 August rather than through repeated manual re-checking.

4. Check the finished file before asking for a decision

Once the merge cascade completed and the final assembly succeeded, the team did not treat "the render finished" as "the film is ready." They checked the output directly: a 1920×1080 frame, 755.167 seconds of video against 755.160 seconds of narration audio (no meaningful drift), loudness at -19.8 LUFS, and no silence longer than two seconds anywhere in the narration track. The narration itself had already passed an independent fact-check review earlier in production, which required two corrections before the build even started. Only once every one of those checks passed did the workflow ask the owner for the two decisions only a person could make: renew the expired YouTube publishing authorization, and watch the finished cut before releasing the quality gate that unlocked publishing.

5. Publish, verify, and report back automatically

The documentary went live on YouTube on 23 August at 18:15, and a confirmation was sent automatically back to the original production ticket (bloomii #1125). Both the recovery ticket and the original ticket were closed a week later, on 30 August, with the owner confirming the result directly in the ticket.

What actually got fixed, and what only got worked around

It would be convenient, but not accurate, to describe this as one bug with one fix. There were three separate things going on, and they deserve to be kept apart. First, the clips generated in the wrong orientation: that was a data problem, not a code problem, and it was resolved by the targeted retry described above, a production fix with no code change behind it. Second, a genuine, unrelated code fix already existed for a different failure mode on the same merge steps: on 20 July, the same day this build started, a Kinoboard ticket (kinoboard #44) shipped a real code change, verified directly in the repository as commit 609e0bb, replacing the merge steps' fixed five-minute processing-time limit with one that scales with each clip's own duration. That fix addresses how long a merge step is allowed to run, not what orientation a clip was generated in; the two failures happened to hit the same class of node without being the same problem.

Third, the final assembly step's six-minute wall, hit three times on 18 August, has no matching fix in the available record. It was worked around: the team re-ran the assembly once conditions changed, and it went through. That is worth stating plainly. One of these three issues has a documented, verified engineering fix behind it (and it isn't the one that looks most related to the recovery's headline failure); one was resolved by data, not code; and the last is a recovery that worked, not a guarantee that the same ceiling will not resurface on the next long production.

How to do the same with KittyClaw

You need a KittyClaw project that already holds the ticket for whatever you are producing, a video, a data pipeline, a multi-step build; a pipeline whose individual steps expose their own status instead of only one overall pass or fail; and an agreed set of checks the output has to pass before a human is even asked to decide.

1. Open a recovery ticket with one job

When a production breaks partway through, do not edit the original ticket in place. Open a new ticket whose only content is a diagnosis: which steps are done, which are waiting, which failed, and why. Link it back to the original ticket. This keeps the record of what actually happened during the incident separate from the plan, and gives you one place to scope the fix.

2. Classify every step before touching any of it

For every step in the broken run, decide whether to preserve it untouched, retry it, wait on it, or send it for human validation. The kit below includes a plain-language decision matrix built from the exact choices made in this recovery; use it to resist the instinct to relaunch everything "to be safe," which can be the slower and more costly option.

3. Turn external limits into a scheduled resume, not a guess

KittyClaw's API exposes PATCH /api/projects/{slug}/tickets/{id}/schedule, which takes a fireAt time, an author, and an optional targetStatus. Call it when you are blocked on something outside your control, a provider quota, a dependency, a scheduled resource, so the ticket wakes up and moves itself once the wait is actually over instead of relying on someone remembering to check back. The kit includes an illustrative example of that call for a quota wait (its dates are fictional, not a copy of the real recovery's actual schedule calls).

4. Define the checks that must pass before a human is asked to decide

Write down, once, what "the output is actually ready" means for your kind of production: format, duration, sync, loudness, whatever applies to what you build. Run that check before creating the approval step, so the person approving is confirming a result that already passed, not doing the first check themselves.

5. Gate, publish, and report back to the source ticket

Keep the human decision (approve for release) separate from the technical one (does the output pass the checks). Once approved, the publish step should report its result, what published, when, and the public link, directly back to the ticket that originally requested it, the same way this recovery's callback closed the loop to bloomii #1125.

6. Start from the kit

The bootstrap kit below has the recovery ticket template, the decision matrix, the master-quality checklist, a worked scheduled-resume example calling the schedule endpoint, and a cleaned, trimmed adaptation of the real production script and node manifest referenced as evidence for this story, with local paths, internal URLs, and the proprietary scene content reduced to a short illustrative excerpt.

What the evidence supports

Diagnosis

A dated, per-step classification (bloomii #1442, 17 August) turned a full production into an 18-clip, 11-image fix.

Preserved work

89 clips and 107 images already correct were left untouched through the whole recovery.

Verified master

1920×1080, 755.167 s video against 755.160 s audio, -19.8 LUFS, no audio silence over two seconds, checked before any publish decision.

No performance claim

This case shows a recovery reaching a verified publication. It says nothing about the documentary's views, watch time, or audience reception.

Watch the published documentary →

Limits and recovery

  • A number pulled mid-incident is a checkpoint, not a final count. The 118 in this story's own title is a real, dated milestone from 17 August, not the clip count of the film that finally published. State what a number actually measures before repeating it.
  • Not every failure in a recovery gets a matching code fix. The merge-step timeout does, verified in the repository. The final assembly step's worker-level cap does not, in the evidence available. Track that gap instead of assuming a workaround is a permanent solution.
  • External provider quotas are outside your workflow's control. Track them as an explicit wait tied to the ticket instead of retrying blindly, and expect the release date attached to any generative pipeline to move.
  • A quality gate is only worth what its checks cover. The checklist here covers frame size, duration sync, loudness and narration silence length; it does not measure visual cut or transition length, which the real recovery did not check either. Add whatever your own format actually requires before trusting a gate that omits it.
  • Never publish local paths, internal service URLs, account identifiers or authorization details from a production project. The kit below has had the real ones replaced with clearly fictional placeholders.

Kit files and detailed evidence

Use these files to run the same recovery workflow on your own multi-step production. The examples intentionally contain no local path, internal service URL, account identifier, or unreduced proprietary content.

Recover your next stalled pipeline instead of rebuilding it

Open a recovery ticket, classify every step with the matrix, schedule the wait instead of guessing, and gate the result before it goes public.