---
title: "History Explainer — a historical event or figure → a cinematic documentary short"
tier: hero
format: short-form-video
theme: history | documentary | education | edutainment
persona: history educator, documentary creator, museum social, study-channel
duration: "1 topic → 45–90s vertical short in ~3–6 min"
budget_usd: "$0.40–$3.00 per short (5–8 scene images/clips + TTS + music + finishing)"
caps: ["gpt-image", "mai-image-2.5", "recraft-v4", "flux-dev", "kontext-edit", "seedance-i2v", "ltx-q-i2v", "inworld-tts", "music", "hyperframes-lower-third", "ffmpeg-burn-subtitles", "ffmpeg-concat", "ffmpeg-mux", "ffmpeg-export"]
skills: ["short-form-video"]
showcases: []
status: "playbook (2026-06-08) — ported from the Pixelle-Video format study; runs on existing Storyboard caps. Showcase render pending."
reliability: 3.7 # 5 − image gen .3 − i2v .6 − tts .2 − music .2; finishing chain deterministic; chyrons/subtitles optional; no proven E2E showcase yet
---

# History Explainer — a cinematic documentary in 90 seconds

Point this at a historical event, figure, or turning point and it builds a **cinematic documentary explainer**: a striking date/fact hook, the context, 3–4 key moments, why it mattered, and a takeaway — narrated by an authoritative voice over epic scoring, with date/place chyrons. Runs entirely on Storyboard's existing Livepeer caps.

## What you'll get

> **A finished reel = vertical 9:16 + voiceover + music + ~45–90s.** Pass `aspect_ratio: "9:16"` to `generate_project` (it now threads to every scene), set the i2v `duration` to 8–15s **on EVERY scene** and animate **all** of them (default ~5s + skipped scenes = the "38s instead of 90s" bug), add `inworld-tts` narration (not the chatterbox default voice) + `music`, and **mux the bed with `audio_fill:"loop"`** so the score covers the whole reel (not just the first part). A silent landscape montage — or one whose music cuts out two-thirds in — is a draft, not a deliverable. See the `short-form-video` skill's *definition of done*.

A 45–90s **vertical (9:16)** documentary short: 5–8 period-accurate cinematic scenes (plus maps/portraits), an authoritative narrator track, epic documentary music, lower-third chyrons for dates and places, burned-in captions, and a finished MP4.

## The cap chain

```
topic (event / figure)
  → agent writes the documentary script + 5–8 scene beats (hook → context → moments → significance → takeaway)
  → generate_project           # one period-accurate cinematic image per beat
       image: gpt-image | recraft-v4 | mai-image-2.5 | flux-dev
       maps/portraits: same image caps with explicit "antique map" / "oil portrait" prompts
  → motion                     # seedance-i2v (resolution:"auto") or ltx-q-i2v — slow cinematic push, 8–15s per scene, animate EVERY scene
  → inworld-tts                # authoritative documentary narrator (NOT the chatterbox default voice)
  → music                      # "[Instrumental] epic documentary orchestral, solemn, building" — a 15s bed is fine; you loop it
  → hyperframes-lower-third    # OPTIONAL polish — date / place chyrons ("Constantinople · 1453"); skip on a first cut
  → ffmpeg-burn-subtitles      # OPTIONAL polish — spoken-word captions in post; skip on a first cut
  → ffmpeg-concat                                      # stitch the per-scene clips
  → ffmpeg-mux  audio_fill:"loop"                      # lay the looped, ducked score under the full reel
  → ffmpeg-export                                      # 9:16, 1080×1920
  # NOTE: the music MUST be muxed with audio_fill:"loop" so a short bed repeats to cover the WHOLE reel —
  # otherwise the score cuts out and the last third is silent (the bug fixed 2026-06-09).
```

## How the agent should run this (interaction contract)

1. **CONFIRM (one message, ≤1 question):** restate in 1 line ("90s 9:16 doc short on <topic>, 6 beats, inworld VO + looped score, ~$1.50–3, ~10–15 min") + real choices: 5 vs 8 beats · flux-dev (fast) vs gpt-image (cleanest, ~150s/scene) · stills-draft vs full animated reel. If the topic is clear, state assumptions and start.
2. **PREVIEW CHECKPOINT:** show the SCRIPT (beats + hook fact + every date, $0) and then the still keyframes (~$0.30) BEFORE any i2v fires; invite fact corrections — a confident wrong date is the worst failure this playbook has.
3. **NARRATE:** per wave: "animating beats 1–4 (seedance, ~2–4 min each), job …"; poll `get_create_media` every 10–15s, ≤4–5 i2v in flight, post a 1-liner at least every 2 min.
4. **FAIL GRACEFULLY:** gpt-image 422/timeout → soften or switch to flux-dev; seedance hang/"fetch failed" → retry once → `ltx-q-i2v` (`num_frames: 241`, resolution "auto") → Ken-Burns still as last resort. 1-line WHAT/WHY/fallback each time; ≤2 retries per scene; a reel with one Ken-Burns beat beats no reel.
5. **DELIVER:** final MP4 + one honest line ("beat 3's crowd has minor anachronisms — period-accuracy ceiling; all dates verified") + ONE next step ("add chyrons + captions, or run the next topic").

## Script structure (beats)

1. **Hook (0–4s)** — a striking date or fact, stated cold. "In 1815, a single volcano caused a year with no summer."
2. **Context (1 beat)** — the world right before. Who, where, why it was about to change.
3. **Key moments (3–4 beats)** — the turning points in sequence, one per scene. Each gets a date/place chyron.
4. **Significance (1 beat)** — why it still matters. The ripple forward.
5. **Takeaway (last beat)** — one resonant closing line.

## Pacing & aspect

- **Measured, authoritative** — 8–12s per scene, slower than a recap reel; let the narration breathe. 9:16, 1080×1920.
- Chyrons appear as each key moment lands; keep them on-screen ~2.5s.
- Music stays solemn and low under narration: layer narration + looped score via `ffmpeg-mux` (`audio_fill: "loop"`); duck the bed so the VO sits on top.

## Watch-outs

- **Render in WAVES — never fire every clip at once (tested the hard way, 2026-06-08).** Animating a whole reel by firing all its scene i2v jobs simultaneously — ×N reels — overwhelmed the SDK `/inference` worker (BYOC video caps have capacity ~2 each behind a single dispatch worker); a ~30-clip burst crashed it and most renders failed. Cap in-flight **video** renders at **≤4–5**, animate **one reel at a time**, and poll each batch to *done* before firing the next. Images are cheaper (cap ~4) but still wave large sets.

- **Period accuracy drifts.** Image models cheerfully add anachronisms (zippers in the 1600s, modern fonts on "old" signage). Prompt the era explicitly, and review every scene for clothing/architecture/tech that doesn't fit. Maps and portraits are safer than crowds.
- **Don't put dates or place names INTO the generated image.** AI text rendering is unreliable — a garbled "Constantinopl" kills credibility. Use `hyperframes-lower-third` chyrons in post for every label.
- **Facts must be right.** This is education — a confident wrong date is worse than no video. Verify the hook fact, the dates, and the significance claim before narrating. The pipeline won't fact-check for you.
- **The multi-link quality ceiling.** A flat script or a robotic narrator sinks an otherwise gorgeous reel. Review the script for accuracy AND drama, and listen to the TTS once before stitching.

## Cross-surface (MCP / CLI / Webapp)

- **MCP:** `generate_project` (period scenes) → `create_media` (TTS, music) → lower-thirds + `director_export`.
- **CLI:** `livepeer story "<event> documentary explainer, 6 beats"` → TTS + music via `livepeer media create` → `livepeer media export --aspect 9:16 --lower-third dates.json --subtitles captions.srt`.
- **Webapp chat:** `make a 90s history explainer about <event>` — the agent drafts beats with date chyrons and offers finishing as follow-ups.

## Execution mode (autopilot / director / auto)

This is a multi-scene project, so pass a `mode` to `generate_project` to control the **pre-finish gate** (after every scene renders, before delivery): `autopilot` (certain brief → ship in one call, no pause), `director` (always pause for a human pick/steer before delivery), or `auto` (recommended — ship UNLESS the auto-critic is unsure: score < `confidence_threshold` (default 0.8), verdict `iterate`, or a flagged shot). Setting `mode` runs the project async (returns a `job_id`); resume a director/auto pause with `review_checkpoint` + `resume_from_checkpoint` (decision `continue` to deliver / `abort` to discard). Omit `mode` for legacy straight-through behavior.
