---
title: "Emotional Micro-Story — a short relatable narrative with an emotional payoff"
tier: hero
format: short-form-video
theme: micro-story | emotional | storytelling | relatable
persona: storytelling creator, brand-narrative social, greeting-card brand, lifestyle channel
duration: "1 topic → 45–90s vertical short in ~3–6 min"
budget_usd: "$0.40–$3.00 per short (4–6 scene images/clips + TTS + music + finishing)"
caps: ["gpt-image", "flux-dev", "seedream-5-lite", "kontext-edit", "ltx-q-i2v", "seedance-i2v", "inworld-tts", "music", "ffmpeg-burn-subtitles", "ffmpeg-concat", "ffmpeg-mux", "ffmpeg-export"]
skills: ["short-form-video"]
showcases: []
status: "playbook (2026-06-08) — ported from the Pixelle-Video format study; runs on existing Storyboard caps. Showcase render pending."
reliability: 3.0 # ref-frame gen −0.3, scene gen −0.3, full-reel i2v −0.6, tts −0.2, music −0.2, 7-stage happy path −0.4 (subtitles/export now OPTIONAL); no E2E showcase yet
---

# Emotional Micro-Story — a tiny story that lands in 90 seconds

Point this at a small human moment — "the last voicemail she never deleted", "the dad who learned to braid hair", "the stray cat that waited" — and it builds a **short relatable emotional narrative**: a warm setup, a turn, and an emotional payoff. Intimate, slow, lingering. Runs entirely on Storyboard's existing Livepeer caps.

## What you'll get

> **A finished reel = vertical 9:16 + voiceover + music + ~45–90s.** Pass `aspect_ratio: "9:16"` to `generate_project` (it threads to every scene), set i2v `resolution:"auto"` + `duration:8–15` **on EVERY scene** and animate **all** of them, add `inworld-tts` narration + a `music` bed, then lay the bed under the whole reel with `create_media model_override:"ffmpeg-mux"` + `audio_fill:"loop"` so the score never cuts out. A silent landscape montage — or one whose music dies in the last third — is a draft, not a deliverable. See the `short-form-video` skill's *definition of done*.

A 45–90s **vertical (9:16)** micro-story: 4–6 warm, intimate cinematic scenes, a gentle narrator (often first-person), emotional piano/strings, soft captions, and a finished MP4.

## How the agent should run this (interaction contract)

1. **CONFIRM** — one line: the moment + beat count (4–6) + narrator choice (first-person vs third-person is the one question worth asking) + "~$0.40–$3.00, ~3–6 min stills / longer with i2v". Then start.
2. **PREVIEW CHECKPOINT** — the **script (setup → turn → payoff) + the Step-0 cast reference frame** (~$0.05). The emotion lives in the script — redirect here, not after $2 of video. Second cheap gate: show all keyframes before ANY animation fires.
3. **NARRATE** — one line per wave with job_id + ETA; animate in waves of ≤4–5 and poll `get_create_media` every 10–15s — a 5-scene i2v pass runs ~5–10 min; never silent through it.
4. **FAIL GRACEFULLY** — `gpt-image` 422 on dark/grief prompts → soften the wording ("an empty chair", not "death") → `flux-dev`; `seedance-i2v` hang → `ltx-q-i2v resolution:"auto"` (`num_frames:241` for ~10s) → Ken-Burns slow-zoom on the still (gentle zoom suits this format anyway). One line per swap: what / why / fallback. ≤2 retries per clip; a partial reel with a note beats nothing.
5. **DELIVER** — final MP4 + one honest note ("character held on N/N scenes; wardrobe drifted slightly in scene 4 — anchor recipe is ~85–90%, not LoRA-grade") + ONE next step: train a per-character `/lora` if this becomes a series.

## The cap chain

```
topic (a small human moment)
  → agent writes the micro-story script + 4–6 scene beats (setup → turn → payoff)
  → generate_project           # warm, intimate cinematic scenes per beat
       aspect_ratio: "9:16"
       image: gpt-image | flux-dev | seedream-5-lite
       consistency: MANDATORY cast lock — do Step 0 first (anchor URL + restate wardrobe; LoRA for a 95% hero)
  → create_media action:"animate"   # ltx-q-i2v / seedance-i2v, resolution:"auto", duration:8–15 — slow, gentle motion sells intimacy
  → create_media action:"tts" model_override:"inworld-tts"   # warm, gentle narrator (chatterbox is clone-from-ref; inworld is the default warm voice)
  → create_media action:"music" lyrics_prompt:"[Instrumental]"   # "warm ambient piano, slow" OR "cinematic strings, hopeful"
  → create_media model_override:"ffmpeg-concat"   # stitch the per-scene clips → base reel
  → create_media model_override:"ffmpeg-mux" audio_fill:"loop"   # VO + looped score under the FULL reel (no silent tail)
  → OPTIONAL: create_media model_override:"ffmpeg-burn-subtitles"   # soft captions in post — skip if the VO carries it
  → OPTIONAL: create_media model_override:"ffmpeg-export"   # only if the mux output isn't already 1080×1920 9:16
```

> **De-engineered happy path = 7 stages** (ref frame → scenes → animate → tts → music → concat → mux). Subtitles and export are finishing extras — add them only when the deliverable needs captions or a re-encode.

## ⚠ Step 0 — Lock your cast (mandatory — this was tested)

This format lives or dies on the **same people across every shot**, and the auto character-anchor is **not enough on its own**. In testing, the default path aged a 7-year-old into a teenager and swapped the grandfather's identity entirely by the final frame. Do this first, in tiers:

1. **Render ONE clean reference frame** of your character(s) — a clear establishing shot (`generate_project` scene 1, or a `create_media` portrait).
2. **Pass it as an explicit anchor on every scene** — `generate_project` with `character_anchor: "<that image URL>"` (a real URL, **not** `"auto"` — auto failed to extract the characters from terse prompts). Verified: this locks **identity, age, and key props** (the child's age, the silver hair, the red bike all held).
3. **Restate the wardrobe in EVERY scene prompt** — "the SAME grandfather in the grey polo", "the SAME girl in the grey shirt with a ponytail". The anchor holds the face; the prompt holds the clothes — without this, wardrobe drifts (measured 0.79–0.9).
4. **For a 95% hero series** — train a per-character `/lora` and apply it to every scene. The anchor recipe above gets you **~85–90% fast**; a LoRA gets you **~95%**.

One hero character → the anchor recipe is reliable. Two+ characters in frame → lock the most important one (expect the secondary to drift more); a LoRA per character is the fix for a polished series.

## Script structure (beats)

1. **Relatable setup (1–2 beats)** — an ordinary, familiar moment. Warmth, not drama yet. "Every morning he made two cups of coffee. He'd made two for forty years."
2. **The turn (1 beat)** — the small shift that changes everything. Quiet, not loud. "This morning, he only needed one."
3. **Emotional payoff (1–2 beats)** — the feeling lands. Don't over-explain — let the image and the pause carry it. "He made two anyway."

## Pacing & aspect

- **Slow and lingering** — 10–14s per scene; let each moment sit. 9:16, 1080×1920.
- Gentle motion (`ltx-q-i2v` / `seedance-i2v`) — a slow drift, a soft push. Never frantic.
- Music carries the emotion: warm piano under the setup, strings swelling on the payoff. The `ffmpeg-mux` step ducks the bed under the narrator and `audio_fill:"loop"` keeps it under the full reel. A held silence before the final line is more powerful than more words.

## What can disappoint (cap ceilings)

i2v can morph a face **mid-clip** even when the keyframe held — keep motions short and gentle, and accept that 8–15s clips carry more drift risk than 5s ones. The `music` cap delivers a mood, not a specific melody; if a precise score matters, bring a licensed track and mux it.

## Watch-outs

- **Render in WAVES — never fire every clip at once (tested the hard way, 2026-06-08).** Animating a whole reel by firing all its scene i2v jobs simultaneously — ×N reels — overwhelmed the SDK `/inference` worker (BYOC video caps have capacity ~2 each behind a single dispatch worker); a ~30-clip burst crashed it and most renders failed. Cap in-flight **video** renders at **≤4–5**, animate **one reel at a time**, and poll each batch to *done* before firing the next. Images are cheaper (cap ~4) but still wave large sets.

- **Character consistency is the make-or-break — see Step 0 (mandatory).** An intimate story dies if "he" looks like three different people; that is the default failure mode, not an edge case. The anchor-URL + restate-wardrobe recipe holds identity/age/props (~85–90%); a per-character `/lora` is the gold standard (~95%). Never ship a narrative on the auto-anchor alone.
- **AI text-in-image is unreliable.** Keep the narrative in narration + post captions (`ffmpeg-burn-subtitles`); don't bake words into the scene.
- **Don't overwrite the emotion.** The format dies when it's manipulative or over-narrated. Trust the pause, the music, the image. Less script, more feeling. Review the payoff: does it earn the emotion, or force it?
- **The multi-link quality ceiling.** A rushed pace, a cheerful-toned narrator, or an upbeat track will flatten an emotional story instantly. Every link must match the tone — review pacing, listen to the TTS, and check the music feels right before stitching.

## Cross-surface (MCP / CLI / Webapp)

- **MCP:** `generate_project` (intimate 9:16 scenes, character anchor URL) → `create_media` (animate `resolution:"auto"`, `inworld-tts`, `music`) → `ffmpeg-concat` → `ffmpeg-mux` (`audio_fill:"loop"`) → `ffmpeg-burn-subtitles` → `ffmpeg-export`.
- **CLI:** `livepeer story "<a small human moment>, 5 emotional beats, warm intimate cinematic"` → motion + TTS + music via `livepeer media create` → `livepeer media export --aspect 9:16 --subtitles captions.srt`.
- **Webapp chat:** `make a 60s emotional micro-story about <moment>` — the agent writes setup→turn→payoff, keeps the character consistent via a reference, and offers gentle finishing as follow-ups.

## Execution mode (autopilot / director / auto)

This is a multi-scene project, so pass a `mode` to `generate_project` to control the **pre-finish gate** (after every scene renders, before delivery): `autopilot` (certain brief → ship in one call, no pause), `director` (always pause for a human pick/steer before delivery), or `auto` (recommended — ship UNLESS the auto-critic is unsure: score < `confidence_threshold` (default 0.8), verdict `iterate`, or a flagged shot). Setting `mode` runs the project async (returns a `job_id`); resume a director/auto pause with `review_checkpoint` + `resume_from_checkpoint` (decision `continue` to deliver / `abort` to discard). Omit `mode` for legacy straight-through behavior.
