---
title: "Novel Recap — a book or chapter → a fast cinematic recap reel"
tier: hero
format: short-form-video
theme: book-recap | booktok | cinematic-narration | series
persona: book reviewer, BookTok creator, study-aid maker, publisher social
duration: "1 topic → 45–90s vertical short in ~3–6 min"
budget_usd: "$0.40–$3.00 per short (6–8 scene images/clips + TTS + music + finishing)"
caps: ["gpt-image", "seedream-5-lite", "flux-dev", "kontext-edit", "ltx-q-i2v", "seedance-i2v", "inworld-tts", "music", "ffmpeg-burn-subtitles", "hyperframes-lower-third", "ffmpeg-concat", "ffmpeg-mux", "ffmpeg-export"]
skills: ["short-form-video"]
showcases: []
status: "playbook (2026-06-08) — ported from the Pixelle-Video format study; runs on existing Storyboard caps. Showcase render pending."
reliability: 3.9 # 5 − cast anchor ref .3 − scene gen .3 − i2v .6 − tts .2 − music .2 + .5 proven E2E (Frankenstein recap reel re-animated via ltx-q-i2v, 2026-06); subtitles/export deterministic; ceiling is anchor-grade cast hold (~85–90%, LoRA for 95%)
---

# Novel Recap — recap a book in 90 seconds, cinematically

Point this at a book, a chapter, or a plot summary and it builds a **fast cinematic recap reel**: a cold-open hook, 5–8 illustrated plot beats narrated dramatically, and a cliffhanger that drives the comments. It is the BookTok / "recap this in 60 seconds" format, done entirely on Storyboard's existing Livepeer caps — no new models.

## What you'll get

> **A finished reel = vertical 9:16 + voiceover + looped music bed + 45–90s, every second covered by audio.** A silent tail, or one whose music cuts out partway, is a draft — not a deliverable. The numbered enforcement below is the *definition of done*; also see the `short-form-video` skill.

A 45–90s **vertical (9:16, 1080×1920)** short: 6–8 cinematic illustrated scenes, a dramatic narrator voice-over, building tense strings that run the full length, burned-in captions, and a stitched + aspect-correct MP4 ready for TikTok / Reels / Shorts.

## The cap chain (the enforced recipe)

```
topic (book/chapter)
  → agent writes the recap script + 6–8 scene beats (hook → beats → cliffhanger)
  → generate_project  aspect_ratio:"9:16"        # one cinematic image per beat
       image: gpt-image | seedream-5-lite | flux-dev
       consistency: MANDATORY cast lock — do Step 0 first (anchor URL + restate wardrobe; LoRA for a 95% hero)
  → motion: animate the 2–3 highest-impact beats, duration 8–12s   # ltx-q-i2v (resolution:"auto") or seedance-i2v
                                                                    #   — fire in waves of ≤4–5 (see Watch-outs)
  → create_media action:"tts" model_override:"inworld-tts"   # dramatic narrator (chatterbox = clone-from-ref only)
  → create_media action:"music"                  # "tense cinematic strings, building, ominous, no vocals"
       lyrics_prompt:"[Instrumental]"             #   MANDATORY — else the bed is sung aloud
  → ffmpeg-burn-subtitles                         # captions in post (do NOT trust AI text-in-image)
  → ffmpeg-concat                                 # stitch the per-scene clips → base reel
  → create_media model_override:"ffmpeg-mux"  audio_fill:"loop"   # VO + looped ducked score over the FULL reel
  → create_media model_override:"ffmpeg-export"   # finalize 9:16, 1080×1920
  # audio_fill:"loop" repeats a short (15s) bed across the whole reel — without it the score cuts out
  #   and the last third runs silent (the silent-tail bug, fixed 2026-06-09). NON-NEGOTIABLE.
```

## How the agent should run this (interaction contract)

1. **CONFIRM (one message, ≤1 question):** restate in 1 line ("90s 9:16 recap of <book>, 7 beats, hero cast-locked, inworld VO + tense strings, ~$1–3, ~5–10 min") + real choices: 6 vs 8 beats · animate only the 2–3 key beats (default, cheaper) vs full i2v pass · spoiler-marked vs stop-before-the-twist. If the book is named, state assumptions and start.
2. **PREVIEW CHECKPOINT:** show the SCRIPT (hook → beats → cliffhanger, $0) + the Step-0 hero reference frame (~$0.05), then all keyframes BEFORE any i2v fires — a recap with the wrong-looking hero or a flat hook dies at the script gate, for pennies.
3. **NARRATE:** per wave: "animating beats 2, 5, 7 (ltx-q-i2v, ~2–4 min each), job …"; poll `get_create_media` every 10–15s, ≤4–5 i2v in flight, post a 1-liner at least every 2 min.
4. **FAIL GRACEFULLY:** gpt-image 422 on dark/violent plot beats → soften ("the aftermath of the duel", not gore) → seedream-5-lite / flux-dev; `seedance-i2v` hang → `ltx-q-i2v` (`num_frames:241`, `resolution:"auto"` — the Frankenstein reel's actual fix) → Ken-Burns push on the still; cast drift on a scene → re-fire once with the anchor URL + wardrobe restated. ≤2 retries per scene; a reel with one still beat beats no reel.
5. **DELIVER:** final MP4 + one honest line ("hero held on N/7 scenes — anchor recipe is ~85–90%, not LoRA-grade; recap stops before the ending, marked spoiler-free") + ONE next step ("Part 2, or train a hero LoRA if this becomes a series").

## ⚠ Step 0 — Lock your cast (mandatory — this was tested)

A recap with 6–8 shots of a recognizable hero is the hardest consistency case there is, and the auto character-anchor is **not enough on its own** (in testing it failed to extract characters from terse prompts and let them drift). Do this first, in tiers:

1. **Render ONE clean hero reference** — a clear portrait/establishing frame of your protagonist (and antagonist, if recurring).
2. **Pass it as an explicit anchor on every scene** — `generate_project` with `character_anchor: "<that image URL>"` (a real URL, **not** `"auto"`). This locks **face, age, and signature props/costume** across shots.
3. **Restate name + wardrobe in EVERY scene prompt** — "the SAME young dark-haired Victor in the white shirt", "the SAME tall stitched Creature". The anchor holds the face; the prompt holds the wardrobe (without it, costume drifts ~0.8).
4. **For a 95% hero series** — train a per-character `/lora` and apply it to every scene. The anchor recipe gets you **~85–90% fast**; a LoRA gets you **~95%** and is worth it for a recurring series (e.g. a multi-part book).

One hero → reliable. Hero + antagonist both on screen → lock the hero, expect the antagonist to drift more; a LoRA per character is the polished-series fix.

## Script structure (beats)

1. **Cold-open hook (0–4s)** — the most shocking line of the book, no setup. "She married the man who killed her father. She just didn't know it yet."
2. **Setup (1 beat)** — who, where, the stakes, in one sentence.
3. **Plot beats (4–6 beats)** — the rising action, one beat per scene, each ending on a turn. Don't summarize evenly — accelerate.
4. **Cliffhanger / CTA (last beat)** — stop before the ending. "And then she opened the letter. Part 2?"

Keep each beat's narration to ~1.5–2.5s of speech so 6–8 beats land in 60–90s.

## Pacing & aspect

- **Fast cuts** — 6–10s per scene, faster as tension rises. 9:16 vertical, 1080×1920.
- Music swells under the cliffhanger and runs the FULL reel, ducked under narration — that is what `audio_fill:"loop"` in the mux step guarantees.
- Animate only the 2–3 highest-impact beats (`ltx-q-i2v` for quality, `seedance-i2v` for cheaper motion, both at `resolution:"auto"`); leave the rest as stills with a slow Ken-Burns push (ffmpeg `zoompan`) if your finishing supports it.

## Watch-outs

- **Render in WAVES — never fire every clip at once (tested the hard way, 2026-06-08).** Animating a whole reel by firing all its scene i2v jobs simultaneously — ×N reels — overwhelmed the SDK `/inference` worker (BYOC video caps have capacity ~2 each behind a single dispatch worker); a ~30-clip burst crashed it and most renders failed. Cap in-flight **video** renders at **≤4–5**, animate **one reel at a time**, and poll each batch to *done* before firing the next. Images are cheaper (cap ~4) but still wave large sets.

- **Character consistency is the #1 failure — see Step 0 (mandatory).** Text-to-image has none by default and the auto-anchor alone won't save you: your protagonist will look like a different person each scene. The anchor-URL + restate-wardrobe recipe holds face/age/props (~85–90%); a per-character `/lora` is the gold standard (~95%) and the right call for a recurring series. Budget for the cast-lock step.
- **AI text-in-image is unreliable.** Never bake the title, chapter number, or a quote into the generated image — it will be garbled. Put all text in post via `ffmpeg-burn-subtitles` / `hyperframes-lower-third`.
- **The multi-link quality ceiling.** A weak link drags the whole reel: a flat script, an off-tone narrator, or muddy art. Review each stage before stitching — read the script aloud, spot-check 2 scene images, and listen to the TTS once.
- **Copyright/spoilers.** Recap, don't reproduce. Paraphrase; don't read long verbatim passages. Mark spoiler reels as such.

## Cross-surface (MCP / CLI / Webapp)

- **MCP:** `generate_project` (`aspect_ratio:"9:16"`, image per beat, `character_anchor:"<hero URL>"`) → `create_media` action:"animate" (waves of ≤4–5) → `create_media` action:"tts" (inworld-tts) + action:"music" (`lyrics_prompt:"[Instrumental]"`) → `ffmpeg-concat` → `create_media model_override:"ffmpeg-mux"` with `audio_fill:"loop"` → `ffmpeg-export`.
- **CLI:** `livepeer story "<book> recap, 7 cinematic beats"` → `livepeer media create --model inworld-tts` → `livepeer media create --model music` → `livepeer media export --aspect 9:16 --subtitles captions.srt` (the export loops the bed over the full reel).
- **Webapp chat:** `recap <book title> as a 90s cinematic reel` — the agent writes the beats, generates scenes, and offers TTS + music + finishing as one-click follow-ups.

## Execution mode (autopilot / director / auto)

This is a multi-scene project, so pass a `mode` to `generate_project` to control the **pre-finish gate** (after every scene renders, before delivery): `autopilot` (certain brief → ship in one call, no pause), `director` (always pause for a human pick/steer before delivery), or `auto` (recommended — ship UNLESS the auto-critic is unsure: score < `confidence_threshold` (default 0.8), verdict `iterate`, or a flagged shot). Setting `mode` runs the project async (returns a `job_id`); resume a director/auto pause with `review_checkpoint` + `resume_from_checkpoint` (decision `continue` to deliver / `abort` to discard). Omit `mode` for legacy straight-through behavior.
