# Viral-dance short — agent playbook ⭐ flagship

> Paste this whole file (with the BRIEF filled in) into Claude cowork, chat, or Code. **Pick a dance style, pick a cast of characters — get a 10-second TikTok-ready short.** Build a vertical 9:16 reel of original characters performing a viral-style synchronized dance under neon lights, with an energy-matched music beat and a TikTok-safe caption overlay. Ready to upload to TikTok / YouTube Shorts / Instagram Reels / Douyin.

## What you'll get

A 10-second (default) social-share dance short, built end-to-end via the storyboard MCP:

- **One ensemble hero keyframe** (gpt-image, 9:16 vertical) — N characters staged with VERTICAL DEPTH (back/middle/front layering), not horizontally side-by-side. This is the critical composition recipe — see Lesson 1 below.
- **One animated dance clip** (seedance-i2v or kling-v3-i2v, 10 seconds) — synchronized choreography described in abstract movement primitives (sway / hop-step / arm-wave / spin) rather than copyrighted choreography. The agent rewrites any specific viral name into renderable equivalents.
- **One TikTok-ready music bed** (`music`, minimax) — structured brief with explicit BPM + drop timing + genre, passed with `lyrics_prompt: "[Instrumental]"`. See Lesson 3 below.
- **One bold caption overlay** (`hyperframes-caption`) — placed in TikTok's UPPER-MIDDLE SAFE ZONE (y≈300), NOT at the bottom (which TikTok's own UI covers). See Lesson 4 below.
- **A showcase HTML page** with the reel embedded, the cast described, choreography breakdown, pipeline trace.

**Time**: ~5–10 minutes wall clock (i2v queue varies). **Your attention**: ~3 minutes (one approval gate on cast + dance style before fire). **Cost**: ~$2–$3 per reel. *Fastest flagship in the catalogue* — single keyframe + single i2v + single music + single caption.

---

## Top-1% TikTok-quality lessons (HARD-WON, do not skip)

### Lesson 1: VERTICAL COMPOSITION RECIPE — stack characters by depth, never side-by-side

**The trap**: when an agent renders N characters in a 9:16 keyframe, the default instinct is to place them in a horizontal row (left-center-right). This DESTROYS the reel.

**Why**: ALL current i2v models (kling-v3-i2v, seedance-i2v, ltx-i2v, veo-i2v) normalize their output to ~4:3 (1108×828 or 1112×834) regardless of source aspect ratio. When ffmpeg then center-crops 4:3 back to 9:16 for TikTok, the LEFT and RIGHT thirds get sliced off — meaning only the center character survives.

**The fix**: stage characters with VERTICAL DEPTH:
- **Back / top** — tallest character standing in the back, head + upper body in the top third
- **Middle / center** — full body of the centerpiece character, mid-third
- **Front / bottom** — closest-to-camera character (smaller due to perspective), bottom third

All three characters then live in the CENTER COLUMN of the keyframe, which is exactly what survives center-crop. Verified on MONSTER DANCE CREW after the first run failed.

### Lesson 2: i2v model selection — they ALL center-crop to ~4:3, so plan for that

**Reality**: kling-v3-i2v, seedance-i2v, veo-i2v — they all output their own internal ratio (~4:3) regardless of 9:16 source. There is no "preserves 9:16 natively" i2v on BYOC today.

**Implication**: stop fighting it. Compose with the assumption that the final reel is a 1080×1920 center-crop of a 1112×834 video. Anything important must live in the center column. Lesson 1 enforces this.

**Model choice**:
- **kling-v3-i2v**: fast (60–120s), reliable, decent motion. Default choice.
- **seedance-i2v**: better character consistency for cartoon characters; slower (200–300s); occasional long queue waits.
- For dance reels, kling is usually fine — the prompt-following is what matters, not the model.

### Lesson 3: Music brief MUST be structured, NOT vibe-only

**The trap**: a brief like "energetic playful spooky-fun beat" sounds specific but produces generic kitsch.

**The recipe**: every music brief must specify:
1. **BPM**: exact number (120 for shuffle/dance, 128 for trap/hyperpop, 100 for samba)
2. **Genre + reference**: "modern hyperpop / electro-trap", "K-pop point-dance trap", "Latin samba"
3. **Drop timing**: "intro 0–2s, drop at 2s, peak 2–8s, fill 8–9.5s, stop-drop 9.5s"
4. **Instrumentation**: "808 bass + finger-snap claps on off-beats + pluck-synth lead hook"
5. **Anti-guide**: "NOT cute, NOT kitschy, NOT spooky" — name the kitsch trap explicitly

Without the BPM + drop-timing combo, the music doesn't match the dance rhythm and the reel feels disconnected.

### Lesson 4: TikTok caption placement — UPPER-MIDDLE, not bottom

**The trap**: placing the caption at the bottom of the 9:16 frame (y≈1450–1700) is invisible after TikTok's own UI overlays its caption, share buttons, and music icon on top.

**The TikTok safe zone**:
- Top 0–130px → unsafe (TikTok status bar)
- **Upper-middle 200–500px → SAFE — put captions here (y=300 default)**
- Middle 500–1400px → safe but covers the main video
- Bottom 1400–1920px → unsafe (TikTok own caption, share UI, music icon)

**The recipe**: caption PNG with bold Impact / Helvetica-Bold ≈ 92pt, white with heavy black drop-shadow (8 directions × 4px offset for legibility on any background), centered horizontally, anchored at **y=300**.

### Lesson 5: Choreography vocabulary — abstract primitives, never named viral routines

**The legal guard**: do not pass "do the Macarena", "Single Ladies bridge", named choreographer's signatures, etc. The playbook rewrites every specific viral name into generic verbs the AI can render:
- sway side-to-side (hip-shift)
- hop-step forward (with optional twist)
- arm-wave (sequenced left-then-right)
- spin / turn / twist
- starfish pose / V-pose / final pose
- shoulder-shimmy / chest-bump-with-self

Combine 4–6 of these into a 10s choreography. The AI will render group-synchronized versions reliably.

### Lesson 6: NO IP characters, EVER

The playbook explicitly refuses Pixar / Disney / Studio Ghibli / anime-franchise / named brand mascot likenesses. If a user asks for one, the agent offers to **invent an ORIGINAL crew** in the same spirit. The MONSTER DANCE CREW example came from a "do the Monsters Inc dance" brief — the agent refused, invented Bramble + Sprig + Ozzie, and built the reel from those. The result reads as a "buddy monster crew" without copying any specific franchise.

---

## Tell the agent about the cast + dance

```yaml
characters:                   # 2-5 character specs. The agent renders each freshly per run.
                              # Each spec ~30 words, with EXPLICIT depth-layer assignment:
                              #  - name:  Bramble
                              #    body:  tall lavender-purple fuzzy monster
                              #    face:  three small round eyes, floppy ears
                              #    layer: back-top      # back-top | middle-center | front-bottom
                              #    pose:  arms raised in V-shape above head
                              #
                              # The agent enforces vertical-stacking composition (Lesson 1).

dance_style:                  # Movement primitive list (Lesson 5). Examples:
                              #  - "synchronized neon-shuffle: sway side-to-side, then arm-wave,
                              #     then hop-step forward, then final V-pose"
                              #  - "K-pop point dance: synchronized finger-snap arm-cross, hip-shift,
                              #     then point-and-pose finale"
                              # The agent REFUSES copyrighted choreography names and substitutes primitives.

setting:                      # Visual context. Examples:
                              #  - "Neon-lit nighttime city alley, magenta + teal signs, full moon, ankle fog"
                              #  - "Sunset rooftop, golden hour, distant city skyline, hazy"
                              #  - "Arcade floor, CRT screens glowing, retro-pixel aesthetic"

caption:                      # Single hook line, ≤6 words. Bold sans, ≈92pt, white + drop-shadow.
                              # Placed at y=300 (Lesson 4 — UPPER-middle TikTok-safe zone).

music_brief:                  # MUST follow Lesson 3 template — BPM + genre + drop timing + instrumentation
                              # + anti-guide. Example for zombie-shuffle:
                              #  "120 BPM modern TikTok-shuffle hyperpop / electro-pop. Intro 0-2s with
                              #   finger-snap pickup, drop at 2s, peak 2-8s with 808 bass on every kick +
                              #   finger-snaps on off-beats + pluck-synth lead hook, fill 8-9.5s,
                              #   stop-drop 9.5s. NOT cute, NOT kitschy, NOT lo-fi. No vocals."

duration_sec:                 # 10 (default) / 8 / 6.

dance_slug:                   # kebab-case for output. e.g. monster-dance-crew / robot-shuffle / fox-line.
```

---

## Quality controls (auto-applied)

The pipeline guards inherited from THE EMISSARY / THE ROAD / 麻婆豆腐 family plus dance-specific additions:

- **Brief-stage IP refusal** — Pixar/Disney/Ghibli/anime franchise characters refused; agent offers original crew.
- **Brief-stage choreography abstraction** — named viral routines rewritten into generic verbs.
- **Vertical-stacking composition** — keyframe prompt explicitly specifies back-top / middle / front-bottom layer assignment per character.
- **Center-column character placement** — every character's body must be in the center 60% of width, so center-crop preserves the cast.
- **Music brief template** — BPM + genre + drop timing + instrumentation + anti-guide required; vague briefs rejected.
- **Caption TikTok-safe placement** — y=300 (NOT y=1450); upper-middle zone above the TikTok UI overlay area.
- **Single shot** — camera stays still, no mid-clip cuts (i2v doesn't do cuts cleanly).
- **Mid-render review gate** — agent surfaces the keyframe AFTER it lands, before firing the expensive i2v wave, so a wrong-composition keyframe can be redone for $0.20 instead of a wasted $1.10 i2v render.

> **Why this playbook is the fastest flagship.** One keyframe + one i2v + one music + one caption. No multi-shot stitching, no per-scene voice-over, no episode arc. Just a single 10-second bang. The constraint set above is what keeps that single bang from looking janky.

## The pipeline (real MCP verbs — 5 calls)

```
1 — KEYFRAME    create_media(action:"generate", model_override:"gpt-image",
                  prompt:<vertically-stacked cast (Lesson 1), 9:16, setting, neon>)
                → hero keyframe.  ⛔ Mid-render gate: surface it before firing i2v.

2 — ANIMATE     create_media(action:"animate", model_override:"seedance-i2v",
                  source_url:<keyframe>, resolution:"auto", duration:<duration_sec>,
                  prompt:<abstract choreography primitives (Lesson 5), single locked shot>)
                → dance clip (~4:3 internal — that's fine, Lesson 2).
                Fallback on failure: kling-v3-i2v, then ltx-q-i2v. Always resolution:"auto".

3 — MUSIC       create_media(action:"music", model_override:"music",
                  prompt:<structured BPM+drop brief (Lesson 3)>,
                  lyrics_prompt:"[Instrumental]", duration:<duration_sec>)
                → music.mp3.  lyrics_prompt:"[Instrumental]" is REQUIRED or it sings your style brief.

4 — CAPTION     create_media(model_override:"hyperframes-caption",
                  source_url:<dance clip>, text:<caption>)  # bold, y≈300 upper-middle (Lesson 4)
                → captioned clip.  (Or render a caption.png and ffmpeg-overlay it at y=300.)

5 — FINISH      a) create_media(action:"mux_audio", model_override:"ffmpeg-mux",
                     source_url:<captioned clip>, audio_url:<music.mp3>, audio_fill:"loop")
                     → MANDATORY: audio_fill:"loop" makes the beat cover the FULL 10s
                       (a bare mux leaves a silent tail when the i2v clip outlasts the track).
                   b) create_media(model_override:"ffmpeg-export",
                     source_url:<muxed clip>, params:{preset:"tiktok-portrait", crop:"center-crop"})
                     → 1080×1920 9:16 reel, center-crop preserving the center column (Lesson 1+2).
```

The whole thing is 5 `create_media` calls. Step 5a's `audio_fill:"loop"` and Step 5b's center-crop are the two finishing moves that separate a clean reel from a janky one.

## Output structure

```
public/chapters/<dance_slug>/
  keyframe.png                  ← hero shot of the cast (vertically-stacked composition)
  caption.png                   ← TikTok-style caption overlay PNG
  music.mp3                     ← the dance-beat score
  <dance_slug>-short.mp4        ← final 10s vertical reel
public/chapters/<dance_slug>-example.html  ← showcase page
```

The example for this playbook is **MONSTER DANCE CREW** — an original three-monster crew (Bramble, Sprig, Ozzie) doing a synchronized neon-lit shuffle. The first pass was rebuilt after the user feedback exposed three bugs (music too kitsch, characters horizontal-staged so center-crop killed two of three, caption in TikTok UI dead-zone). The lessons above are baked from that rebuild. See `/chapters/monster-dance-crew-example.html`.
