---
reliability: 4.1 # computed: 5.0 − 0.3 (keyframe image batch) − 0.6 (i2v batch); beat map + ffmpeg finishing deterministic/LLM (−0.0); 4 happy-path stages (LoRA train OPTIONAL, would add −0.8); no E2E showcase bonus
---

# Animated music video — agent playbook

> Paste this whole file (with the BRIEF filled in) into Claude cowork, chat, or Code. The agent walks through it with you and stops for your approval at every step that matters. **This is the playbook for indie musicians and Suno users who want an animated MV with a character they recognize — not a "trippy AI visualizer."** Every cut lands on a downbeat. The protagonist holds across every scene. The keyframes ARE the beats. 50 minutes, under $8, exported in 16:9 and 9:16 for Reels and TikTok.

## What you'll get

A music video assembled the way an animation house assembles a real one — beats first, character second, cuts third, color last:

- **Beat map** — every downbeat, accent, and section boundary in the track, timestamped. The video's editorial spine.
- **Narrative arc divided into 4-6 song sections** — intro, verse, chorus, bridge, drop, outro (calibrated to your track's actual structure). Each section gets its own scene.
- **Character anchor** — protagonist locked via LoRA or kontext-edit reference. Same face across every section, every cut.
- **8-12 keyframes locked to beats** — one keyframe per significant beat moment in the song. Drawn from the narrative arc, not from a vague "match the vibes" instruction.
- **Animated clips per keyframe (~2-3s each)** — `seedance-i2v` (resolution `"auto"`) brings each frame into motion, cut to its beat slot.
- **Beat-aligned final cut** — `ffmpeg-concat` with hard cuts on the downbeats, then the full song muxed over the whole video via `ffmpeg-mux` with `audio_fill:"loop"` so the track covers every frame (no silent tail).
- **Two exports** — 16:9 for YouTube + 9:16 for Reels / TikTok. Both ship together.
- **Project archive** — every keyframe URL, every clip URL, the beat map, the section breakdown, all persisted for a re-cut or a sequel video.

**Time**: ~50 minutes wall clock. **Your attention**: about 12 minutes across 7 gates. **Cost**: $4.50–$8.00 end to end. *Goes deeper than the generic music-video playbook* — that one renders to any song; this one targets animation specifically with character lock and beat-locked editorial discipline.

## Tell the agent about the song

```yaml
song_url:                # the track — Suno URL, your .mp3 / .wav, or any URL ffmpeg can read
song_title:              # appears in the export filename + optional title card
artist_name:             # appears in the project credit
protagonist_lora_id:     # OPTIONAL — if set, character anchor active. Leave blank for kontext-edit reference
protagonist_brief:       # 3-5 visual identifiers, used to train LoRA or for reference prompts
narrative_arc_one_line:  # the video's story in one sentence. e.g. "A girl walks through her childhood town as it transforms into the future"
visual_style:            # e.g. "Studio Ghibli pastoral" · "Akira urban-decay" · "Spider-Verse comic-pop" · "ink-wash anime" · "rotoscoped indie"
palette_hex:             # 3-4 hex colors space-separated, e.g. #1A1B26 #C8A165 #E5A4CB #6B2E2B
aspect_ratio:            # 16:9 (default) — both 16:9 and 9:16 exports always ship
mood_arc:                # OPTIONAL — how the mood evolves across sections, e.g. "melancholic → hopeful → triumphant → quiet"
video_slug:              # kebab-case, used in filenames, e.g. childhood-town-mv
```

## How the agent should run this (interaction contract)

1. **CONFIRM (one turn, no render):** restate in 1 line ("animated MV for {song_title}, {visual_style}, {N} keyframes, character {brief}") + the 3 choices that matter: **narrative vs impressionistic**, **anchor source** (reference default — LoRA only if the user volunteers 6+ ref images AND a multi-video future), **keyframe count** (8 vs 12 ≈ $2 swing). Quote ~$4.50–$8 / ~50 min. One question max (narrative vs impressionistic); otherwise assume and start.
2. **PREVIEW CHECKPOINT = the keyframe contact sheet** (Step 4, ~$1.80). Show all 8–12 keyframes + the consistency-critic verdict BEFORE any animation: "reply 'go' to animate (~$2.40, ~12 min) or redo [#]". The beat map (Step 2–3) is a second, free checkpoint — one message, user can shift cuts.
3. **NARRATE:** animation is the long pole — post "clip 4/10 done, ~3 min remaining" per clip via async `job_id` + `get_create_media` polling. Render i2v in waves of ≤4–5; never go silent >2 min.
4. **FAIL GRACEFULLY:** `seedance-i2v` hang/fetch-failed → `ltx-q-i2v` with `resolution:"auto"` → `cosmos-3-i2v` → Ken-Burns zoompan on the keyframe (still beat-aligned, honest note). Keyframe 422 → soften prompt once → `gemini-image`/`nano-banana`. Max 2 retries per clip; deliver an N−1-clip cut with one Ken-Burns slot over nothing.
5. **DELIVER:** both exports + 1 honest line ("cuts are on the mapped beats; clip 7 fell back to Ken-Burns") + ONE next step ("want a 30s teaser cut from the chorus?").

Gate discipline: 4 stops (brief, beat map, keyframes, final). Steps 5–7 STOPs are advisory — only pause if something failed or the critic flagged drift.

## Quality controls (new — auto-applied)


This playbook leverages the closed quality loop. The agent applies these by default; you don't have to ask:

- **Cost upfront.** Before any multi-step batch fires, the agent calls `submit_plan` with `confirm:false` and shows you the proposed steps + total cost. Nothing runs until you approve.
- **Auto-retry on weak shots.** The server default is `quality_threshold: 0.7`. For character-anchored keyframes the playbook step tells the agent to pass `quality_threshold: 0.85` so a low-confidence first try gets one auto-retry before you see it.
- **Critic verdict on the result.** When the multi-keyframe step finishes, `critique_shot` runs across the batch with `focus: "consistency"` to catch face drift before animation fires. Animation amplifies drift — catching it pre-animation saves $1-2 per redo.
- **One-click recovery.** If a step fails, `get_plan` returns a `recovery` hint pointing at `replan({plan_id})` — the agent classifies the failure and proposes a recovery plan you can approve.
- **End-of-project scorecard.** The final wrap-up calls `get_scorecard({project_id})` with succeeded count, mean quality score, critic verdict, total cost, and wall time.

You can override any of these per step: pass `quality_gate: false` on cost-sensitive tries, `quality_threshold: 0.6` for stylized/abstract work where strict adherence isn't the point, or `submit_plan({plan_id, confirm:true})` to skip the cost-preview pause when you trust the chain.

## How this works

Music videos live and die by editorial timing. A cut that lands a half-beat early reads as amateur; a cut that lands on the downbeat reads as professional. Every existing AI music-video tool fails at this because they don't know where the beats are — they "auto-cut every 3 seconds" and call it good. This playbook fails differently: beat-extract first, score the narrative against the beats, key-frame to the strongest beat moments, animate to the beat-aligned cut points. The keyframes ARE the beats.

The second failure mode of AI music videos is character drift — Scene 1 girl looks nothing like Scene 4 girl. Solved here with LoRA or anchor reference, locked before any keyframe renders.

Seven gates. The beat map and character lock are the foundation; the rest is mechanical assembly.

## The steps

### Step 0 — Confirm the brief

The agent reads back the BRIEF, previews `song_url` (waveform if available), and asks one editorial question:
- "Is this a narrative MV (story unfolds linearly) or an impressionistic MV (mood progression without literal plot)?"

The answer drives how rigidly the keyframes hew to a story vs. evoke a feeling.

**STOP**: "Brief looks right? (approve / edit [field] [new value])"

### Step 1 — Character anchor lock (~5 min if LoRA preset, ~15 min if training)

If `protagonist_lora_id` is set: `apply_lora({lora_id})`, confirm load, done.

If not: **default to kontext-edit reference** — it's the no-extra-cost, no-extra-failure-mode path and holds well to ~8 keyframes. Only offer LoRA training (OPTIONAL, +15 min, +1 failure mode) when the user has 6+ reference images AND plans more than one video with this character; don't block the run on the question — state the default and continue.

> **For the agent**: if training, `submit_lora_train` with the reference set, poll `get_lora_train` every 30s. If reference-only, pick the strongest reference as the `quality_anchor_url` for every subsequent keyframe.

**STOP**: "Character anchor locked. Proceed to beat extraction? (approve / change anchor)"

### Step 2 — Beat-extract the song (~30s, ~$0.05)

**Honesty note: there is no audio-beat-detection cap on production.** The agent estimates the beat map — ffprobe duration + LLM analysis of the track's known/declared BPM and structure (or asks the user for the BPM, which Suno and any DAW display). The map is editorially useful, not sample-accurate; cuts land "on the beat" to within ~100–200ms, which reads fine at MV pacing. The agent returns:
- Full beat timestamp list (every downbeat + accent)
- BPM
- Section boundaries (intro / verse / chorus / bridge / drop / outro — whichever the track has)
- Recommended cut points: typically 8-12 beats spaced for narrative pacing, biased to downbeats and section transitions

For a 3-4 minute track the agent picks 8-12 strong cut points. The user can adjust before keyframes fire.

> **For the agent**: surface the section breakdown as a list ("intro 0:00-0:15 → verse 0:15-0:45 → chorus 0:45-1:15 → ...") with the recommended cut points overlaid. If a section is unusually short or long, surface that — the user may want to skip or extend.

**STOP**: "Beat map + section breakdown shown. Approve the cut points, or shift? (approve / shift [#] [direction])"

### Step 3 — Narrative arc to sections (~1 min, ~$0.05)

The agent maps `narrative_arc_one_line` across the song's sections from Step 2. Example: a 4-section track with the arc "girl walks through her childhood town as it transforms" → intro: town at present; verse: childhood memories layering in; chorus: full transformation visible; outro: girl looks back. Each section gets a one-sentence beat description that drives its 2-3 keyframes.

If `mood_arc` is set, the agent layers mood progression onto the section breakdown.

**STOP**: "Arc-to-sections mapping shown. Approve, or shift a section's beat? (approve / edit [section] [direction])"

### Step 4 — Generate 8-12 beat-locked keyframes (~10 min, ~$1.80)

The agent fans out the 8-12 cut-point beats as parallel `create_media` calls. Each prompt combines:
- The section's narrative beat description from Step 3
- The character anchor (LoRA or reference)
- `visual_style` injected explicitly
- `palette_hex` colors named in the prompt
- Section-specific mood from `mood_arc` if set
- Aspect ratio from `aspect_ratio` (and a flagged 9:16 variant for the vertical cut later)

> **For the agent**: use `flux-dev` with LoRA if `protagonist_lora_id` set; `kontext-edit` otherwise with anchor as `image_url`. Pass `quality_threshold: 0.85` on every keyframe. After all land, call `critique_shot` with `focus: "consistency"` across the batch — surface drift verdict before animation. Animation amplifies drift; redoing a $0.15 keyframe is cheaper than redoing a $0.20 animated clip.

**STOP**: "8-12 keyframes + consistency critic shown. Approve all, or redo specific? (approve / redo [#] [direction])"

### Step 5 — Animate keyframes to ~2-3s clips (~12 min, ~$2.40)

The agent calls `submit_plan` for parallel `create_media({ action: "animate", model_override: "seedance-i2v", source_url: <keyframe_url>, resolution: "auto", duration: <beat_slot_seconds> })` jobs. Each clip is sized to its beat slot from Step 2 — typically 2-3s for verse beats, 3-4s for chorus beats, longer for sustained sections. Motion prompts match the keyframe's narrative beat ("slow walk forward", "wind through leaves, character still", "explosive forward motion on the drop").

> **For the agent**: `seedance-i2v` is the default (cinematic ≤15s). For a more fluid camera on hero beats, `grok-imagine-video`; for cheaper high-fidelity, `cosmos-3-i2v`. Always pass `resolution: "auto"` (any fixed resolution like `"1080p"` 422s on the ltx tier). Pass `quality_threshold: 0.8` per clip. Render 16:9 here; the 9:16 export comes from re-cropping in Step 7 via `ffmpeg-export` (scale+crop). Surface total cost in `submit_plan` preview before firing — this is the biggest line item.

**STOP**: "Animated clips shown. Approve all, or redo a clip? (approve / redo [#] [direction])"

### Step 6 — Beat-aligned concat + audio mix (~2 min, ~$0.10)

The pipeline fires as one approved plan:

1. `create_media({ model_override: "ffmpeg-concat", ... })` on all animated clips with the cut timestamps from Step 2 — `transition: "cut"` is the default; for ballad-y / dreamy pacing, use `crossfade-300`.
2. `create_media({ model_override: "ffmpeg-mux", source_url: <concat_url>, audio_url: <song_url>, audio_fill: "loop" })` — the full song muxed over the whole stitched video. `audio_fill: "loop"` guarantees the bed covers the entire video with no silent tail; the cuts already sit on the beat map.

> **For the agent**: `ffmpeg-concat` with 8+ clips needs the direct-SDK fallback pattern (the MCP envelope drops array params over a certain length). If the song is shorter than the video, `audio_fill: "loop"` repeats it; if longer, the mux trims to video length.

**STOP**: "Concat preview shown. Approve, or shift a cut? (approve / shift [cut #] [direction])"

### Step 7 — Dual export (16:9 + 9:16) (~3 min, ~$0.15)

Two parallel `create_media({ model_override: "ffmpeg-export", source_url: <muxed_url>, ... })` calls:

- **16:9 / YouTube** — `size: landscape_16_9`, 1080p, 30fps
- **9:16 / Reels-TikTok** — `size: portrait_9_16`, 1080p, 30fps. Crop via `ffmpeg-export` scale+crop (center-weighted toward the upper third). Do NOT use `opencv-smart-crop` — it is not registered on production.

Both files ready to upload.

> **For the agent**: surface the cost preview via `submit_plan` before firing the export; if the user only wants 16:9, they can skip the 9:16 ($0.075 saved).

### Final step — scorecard wrap-up

Call `get_scorecard({project_id: ...})` — paste output to your channel.

```
Scorecard for cjob_xxx
N/N keyframes shipped — critic: SHIP
Title: {song_title} MV
Steps: N/N succeeded (100%)
Quality: mean X.XX · M/N passed · K retries
Critic verdict: SHIP — character face consistency 0.XX across all keyframes
Cost: $X.XX actual (est $Y.YY) · Ts wall time
Viewer: <project URL>
```

## When you're done

The agent prints:

```
✅ Music video (16:9 / YouTube): <URL>
✅ Music video (9:16 / Reels-TikTok): <URL>
✅ 8-12 keyframes: <URL × N>
✅ Animated clips: <URL × N>
✅ Beat map JSON: <URL>
✅ Project viewer: <project URL>

Total spent: $X.XX
Total wall-clock: MM:SS
Critic verdict: SHIP — character face consistency 0.XX across all keyframes
```

## What can disappoint (cap ceilings)

- **Beat alignment is estimated, not detected.** No audio-beat-detection cap exists on production — the beat map is BPM-math + LLM structure analysis. Cuts read beat-aligned at MV pacing but won't survive frame-accurate scrutiny in a DAW.
- **i2v drifts from stylized keyframes.** seedance/ltx pull stylized 2D ("Spider-Verse", "ink-wash") toward photoreal over 3s; keep clips ≤3s and motion prompts minimal. The keyframes look better than the moving frames — that's the current ceiling.
- **fal video latency is real.** Budget for 1–2 hung clips per 10; the fallback chain (ltx-q-i2v → cosmos-3-i2v → Ken-Burns) is the plan, not the exception.
- **Cross-scene face consistency is reference-grade, not LoRA-grade**, unless you trained one. Expect minor drift the critic will flag.

## What to do next

You just made an animated music video where the cuts land on the beats and the protagonist's face doesn't shift between scenes. That's the bar Aphex Twin music videos cleared with a six-figure budget and a 16-week timeline. You cleared it in 50 minutes for under $8.

The character LoRA + the visual style guide are now persistent project assets. The next single off your EP gets a music video that visibly belongs to the same world — same protagonist, same palette, same beat discipline — for $4 instead of $8 because the foundation already exists. By the time you ship four singles, you have a recognizable visual identity that fans pattern-match to your music. That's how indie acts who've never been on radio build cultural footprints.

If you're a Suno user, this playbook turns "song I generated" into "released single with a music video" — the final 5% of the work that takes a hobby into a release calendar. The 16:9 cut goes to YouTube, the 9:16 cut goes to Reels and TikTok with the song's hook in the first 3 seconds. TikTok's algorithm punishes vertical cuts that aren't beat-aligned; this one is, frame-by-frame.

If you're a band looking for a visualizer to play behind a live set, run the playbook with `visual_style: "abstract impressionistic"` and a `narrative_arc_one_line` of "instrumental visualization, no character" — the playbook still respects beats and palette, but skips the protagonist work and ships in 30 minutes for $3. That's the format that loops behind your set without the audience reading literal narrative into it.
