---
title: "Storybook City Reel — any story into a cozy hand-drawn street-life reel"
tier: hero
format: short-form-video
theme: illustration | style-capture | storybook | city | social
persona: illustration-loving creator, city/tourism account, lifestyle brand, Xiaohongshu/TikTok aesthetic channel
duration: "1 story → ~50s vertical reel in ~15–30 min"
budget_usd: "$1.50–$5.00 per reel (5 keyframes + 5×10s i2v + 1 music track)"
caps: ["gpt-image", "nano-banana", "ltx-q-i2v", "seedance-i2v", "music", "ffmpeg-concat", "ffmpeg-export", "ffmpeg-audio-mix"]
skills: ["short-form-video"]
showcases: ["/chapters/vancouver-worldcup-storybook-example.html", "/chapters/worldcup-demon-score-example.html"]
status: "playbook + flagship showcase rendered E2E (2026-06-10) — Vancouver World Cup opening week; + DEMON-scored variant (soundtrack generated live on music.daydream.live)"
reliability: 4.4 # 5 − image gen .3 − i2v .6 − music .2 + .5 proven E2E Vancouver flagship (2026-06-10); 5-stage happy path, assemble + gates deterministic ffmpeg/ffprobe
---

# Storybook City Reel — a captured illustration style that turns any story into a cozy animated short

This playbook **captures a specific illustration style** — the "cozy storybook street-life" look you see on viral Xiaohongshu / Instagram city reels (a London street, a Paris corner, all flat 2D folk-art charm) — and turns **any user story** into that style: 5 styled keyframes → 5 gently-animated clips → one ~50s vertical reel with an instrumental jazz bed. The style is locked by a verbatim style block pasted into *every* prompt; the protagonist is locked by a name + fixed wardrobe restated in *every* scene.

## The captured style block (use VERBATIM in every keyframe prompt)

> Flat 2D hand-drawn storybook illustration, folk-art naive style. Clean uniform dark-brown ink outlines; flat color fills with subtle paper-grain texture; NO gradients, NO 3D shading, NO photorealism; at most minimal soft shadows. Warm muted retro palette: brick red / terracotta, cream / ivory, slate teal-blue, sage and olive green, mustard ochre, dusty rose pink, charcoal; warm off-white sky. Dense European-style streetscape: tall narrow brick townhouses with white-trimmed sash windows, dormers and chimneys; ground-floor shopfronts with dark hand-lettered signboards, window flower boxes, awnings; vintage cast-iron street lamps. Many small charming people at naive proportions (heads ~1/5 of body): vintage 1940s–60s attire — long wool coats, cloche hats, fedoras, scarves, satchels; rosy single-dot cheeks, simple dot eyes; everyone engaged in everyday actions (strolling, chatting at the café counter, cycling, carrying coffee or flowers). Busy but orderly composition with overlapping vignettes; larger figures in the foreground, street receding behind. Blossom trees frame the edges with confetti-dot foliage; petals drift through the air and dot the ground. Mood: cozy, nostalgic, gentle — everyday life is beautiful.

The streetscape clause is a *default*, not a cage — swap "Dense European-style streetscape…" details for your city's landmarks (drawn in the same naive style) and the look still holds. The flagship showcase draws Gastown, BC Place and Chinatown this way.

## The animation grammar (use in every motion prompt)

Drifting falling petals · people strolling slowly · subtle coat/scarf sway · steam rising from cups · warm shop windows glowing · gentle parallax. **NO camera shake, NO morphing** — the illustration stays flat-2D while elements move. Add exactly **one concrete character action** per scene (sips the coffee, waves the flag, raises the phone, throws arms up, toasts with friends) so the motion gate has something to measure.

## How the agent should run this (interaction contract)

1. **CONFIRM (one message, ≤1 question):** restate in 1 line ("your story as a 5-scene storybook city reel, 9:16 ~50s, instrumental jazz bed, ~$3.20–4.30, ~15–30 min") + real choices: gpt-image (most obedient flat-2D, ~2–4 min/frame) vs nano-banana (much faster, slightly looser palette) · your city's landmarks vs the default European streetscape. Ask only if the protagonist's name/wardrobe isn't derivable from the story.
2. **PREVIEW CHECKPOINT:** the 5-beat script + named cast with fixed wardrobe ($0), then the 5 keyframes (~$0.30–1) gated for flat-2D style + wardrobe hold BEFORE the ~$2.80 i2v pass fires.
3. **NARRATE:** keyframes ≤3 in parallel (~2–4 min each — say gpt-image is slow up front); i2v ≤4 in flight (~2–4 min each); poll `get_create_media` every 10–15s, 1-liner per wave, never silent >2 min.
4. **FAIL GRACEFULLY:** gpt-image client timeout → re-fetch, don't re-render (the image usually landed on fal); ltx-q-i2v 422 → check `resolution:"auto"` + `num_frames:241` (its `duration` param is IGNORED); 3D-ifying/face-morph at the motion gate → re-roll once with "flat 2D stays flat" emphasized → swap to `seedance-i2v` (which needs `resolution:"1080p"`, NOT "auto" — the schemas disagree). ≤2 retries per clip; ship the best passing takes.
5. **DELIVER:** ffprobe-gated reel (audio ≈ video length, 1080×1920) + one honest line ("wardrobe held ~90%; the crossbody bag drifted in scene N — a per-character LoRA is the 95% path") + ONE next step (reuse the style block verbatim on your next story).

## The cap chain

```
any user story
  → 1. SCRIPT    agent breaks it into 5 beats; protagonist gets a NAME + fixed wardrobe
  → 2. KEYFRAMES gpt-image, image_size:"portrait_16_9" (9:16)      # style block VERBATIM + beat + name/wardrobe restated
  → 3. MOTION    ltx-q-i2v per scene — resolution:"auto", num_frames:241, frames_per_second:24, generate_audio:false  (= 10s)
  → 4. MUSIC     music cap — lyrics_prompt:"[Instrumental]", warm spring jazz style prompt
  → 5. ASSEMBLE  concat → 1080x1920 30fps yuv420p → loop the music under the full reel → ffprobe gate
```

### 1 · SCRIPT — name the cast, fix the wardrobe

Break the story into **5 beats** (arc: opening moment → rising energy → centerpiece → peak emotion → warm close). Define the recurring protagonist with a **NAME and a fixed wardrobe** — e.g. *"Mei, a young Chinese woman in her early 20s, black bob haircut, mustard-yellow wool coat, white sneakers, small sage-green crossbody bag"* — and **restate the full name + wardrobe in every scene prompt**. Named cast = consistency (named → 0.85–1.0 hold, unnamed → drift; same finding as the novel-recap and emotional-microstory tests). Pick wardrobe colors *from the style palette* (mustard ochre, sage green…) so the protagonist belongs to the illustration.

### 2 · KEYFRAMES — gpt-image with the style block verbatim

`gpt-image` (openai/gpt-image-2) has the best stylization fidelity on BYOC — it actually obeys "NO gradients, NO 3D shading" where flux-dev drifts photoreal. **Every scene prompt = the full style block verbatim + the scene beat + the protagonist's name/wardrobe restated**, plus a tail like *"Vertical 9:16 storybook poster composition. Absolutely flat 2D illustration; no photorealism, no 3D rendering."* Pass `image_size: "portrait_16_9"` — **gotcha (verified 2026-06-10):** fal's gpt-image schema takes an *enum* (`portrait_16_9`, `portrait_4_3`, `square_hd`, …), and 422s on pixel-dimension strings like `"1024x1536"`. Expect ~2–4 min per frame; render ≤3 in parallel. **Fallback: `nano-banana`** — also strong on flat illustration, much faster, slightly less obedient on palette.

**Gate each keyframe before animating:** open it and check (a) flat-2D storybook per the block — reject any photoreal/3D drift and regenerate with the style constraints emphasized; (b) the protagonist is recognizable with the SAME wardrobe in all 5.

### 3 · MOTION — ltx-q-i2v, with today's cap gotchas

`ltx-q-i2v` (fal-ai/ltx-2.3-quality/image-to-video) is the right tier — it animates flat illustrations without 3D-ifying them. **Cap gotchas (verified 2026-06-10, put these in your params):**

- `resolution: "auto"` — the fal schema changed; `"1080p"` now 422s on ltx. (This silently broke all ltx i2v until fixed.)
- **`duration` is IGNORED by ltx-q-i2v.** To get 10s clips pass `num_frames: 241, frames_per_second: 24` explicitly.
- `generate_audio: false` — you're laying a music bed; baked-in audio just fights the mix.
- If you swap in `seedance-i2v` as the alternate: it's the *opposite* — seedance **requires** `resolution: "1080p"` (rejects `"auto"`). The two caps' schemas disagree; don't copy params between them.

Motion prompt = the animation grammar above + the scene's one concrete character action. **MOTION GATE:** extract frames at t=2s and t=7s — they must show **visibly different poses/action** (petals moved, the arm raised, the sip taken — not just a zoom), and identity/style must hold (still flat-2D; reject any clip that 3D-ifies or morphs the character).

### 4 · MUSIC — instrumental only

`music` cap (minimax), **`lyrics_prompt: "[Instrumental]"`** — without it you get random vocals over your cozy street. Style prompt to match the reference reels' "spring jazz" vibe: *"warm spring jazz, brushed drums, upright bass, light piano, cheerful café mood"*. Ask for ~the reel length (≈50s); minimax tracks run ~60s, the loop+trim in step 5 makes the exact length a non-issue.

### 5 · ASSEMBLE — concat, conform, loop the bed, gate

Local ffmpeg is fine (or `ffmpeg-concat` + `ffmpeg-export` + `ffmpeg-audio-mix` caps for the server-side path):

```bash
# concat the 5 clips (concat demuxer), then conform + mux with looped music
ffmpeg -f concat -safe 0 -i list.txt -c copy base.mp4
ffmpeg -i base.mp4 -stream_loop -1 -i music.mp3 \
  -vf "scale=1080:1920:force_original_aspect_ratio=increase,crop=1080:1920,fps=30,format=yuv420p" \
  -map 0:v -map 1:a -c:v libx264 -crf 18 -c:a aac -b:a 192k -shortest reel.mp4
```

`-stream_loop -1` on the audio + `-shortest` = the cap-side `audio_fill:"loop"` semantics — the bed runs under the FULL reel, no silent tail. **ffprobe gate before shipping:** audio duration ≈ video duration (±1.5s), ~50s total, 1080×1920, and 3 sampled frames still show the style + the action.

## When to use this playbook

- A city / travel / "week in my life" story that wants the **cozy illustrated reel** aesthetic instead of photoreal.
- Event tie-ins (a festival, an opening week, a season) where charm > realism and brand-safety matters (everything is drawn, generic flags/jerseys are easy, logos are simply never drawn).
- Any brand that wants a *consistent captured style* across many stories — the style block is the asset; reuse it verbatim.

## Honest limits

- **Text-in-image is unreliable.** Keep signboards SHORT and generic ('MAPLE HOUR', 'CAFÉ') — one or two words usually render in this hand-lettered style; full sentences won't. Anything load-bearing goes in post.
- **Style drift on i2v is the #1 risk.** ltx-q occasionally adds depth/lighting that reads as 3D, or morphs the dot-eyes face. The motion gate exists for this — reject and re-roll with "flat 2D illustration stays flat" emphasized. Ship the best passing take; don't ship a morph.
- **Wardrobe ≈ consistency ceiling.** Name + restated wardrobe holds ~85–95% across 5 scenes; tiny accessories (the crossbody bag) drift the most. A per-character LoRA is the 95%+ path if this becomes a series.
- **Crowd scenes dilute the protagonist.** Always say "in the foreground" for your hero, or the i2v model animates the crowd and forgets the character action.
- **gpt-image is slow** (~2–4 min/frame) and occasionally times out client-side while the image still lands on fal — retry the *fetch*, not the render, before paying twice.

## Cost estimate

| Step | Cap | Est. |
|---|---|---|
| 5 keyframes | gpt-image portrait_16_9 | ~$0.30–$1.00 |
| 5 × 10s motion | ltx-q-i2v @ ~$0.056/s | ~$2.80 |
| 1 music track | music (minimax) | ~$0.10–$0.50 |
| assemble | local ffmpeg | $0 |
| **Total** | | **~$3.20–$4.30 / reel** |

## Cross-surface (MCP / CLI / Webapp)

- **MCP:** `create_media` ×5 (`model_override:"gpt-image"`, the style-block prompt, `image_size:"portrait_16_9"`) → `create_media action:"animate" model_override:"ltx-q-i2v"` ×5 (`source_url`, `resolution:"auto"`, `num_frames:241`, `frames_per_second:24`, `generate_audio:false`) → `create_media action:"music"` (`lyrics_prompt:"[Instrumental]"`) → `ffmpeg-concat` → `ffmpeg-export` (1080×1920) → `ffmpeg-audio-mix` / `ffmpeg-mux` with `audio_fill:"loop"`.
- **CLI:** `livepeer media create --model gpt-image --prompt "<style block + beat>" ...` per scene, then `livepeer media animate --model ltx-q-i2v ...`, `livepeer media music --instrumental ...`, finish with `livepeer media export --aspect 9:16`. Wave the i2v calls ≤4 in flight.
- **Webapp chat:** paste the style block + your story and ask for *"a 5-scene storybook city reel, 9:16, instrumental jazz"* — the agent routes the same chain; gate the keyframes on the canvas before saying "animate them".

## Flagship showcase

**[Vancouver World Cup Storybook →](/chapters/vancouver-worldcup-storybook-example.html)** — Mei's life moments in downtown Vancouver during FIFA World Cup opening week (June 2026, BC Place is a host city): Gastown steam-clock coffee → Granville fan march → BC Place at golden hour → fan-zone goal moment → Chinatown evening with bubble tea. 5 × 10s, motion-gated, warm spring-jazz bed, no FIFA marks anywhere (generic flags/jerseys/balls only).
