---
reliability: 3.8 # computed: 5.0 − 0.3 (9-keyframe image batch) − 0.6 (grok i2v wave) − 0.6 (kling i2v wave) − 0.2 (minimax-music); ffmpeg finishing −0.0; 5 stages, no overflow; +0.5 proven E2E showcase (OPUS 1)
---

# Beat-driven opera — agent playbook ⭐ flagship

> Paste this whole file (with the BRIEF filled in) into Claude cowork, chat, or Code. **This is the single-protagonist showcase that proves cast-lock + premium i2v on every shot.** A 90-second beat-locked ballet, one identity held across 9 shots, two premium caps mixed deliberately. The example is [OPUS 1](/chapters/opus-1-example.html) — Elena alone in an empty Beaux-Arts studio at first light.

## What you'll get

A flagship-grade single-protagonist piece, paced by music, anchored by cast-lock:

- **9 keyframes** in your locked visual style — every frame holds the same character identity, same wardrobe, same expression baseline.
- **9 real premium i2v animations** — no Ken-Burns. Two caps mixed: `grok-imagine-video` (premium fluid camera) for the "camera dances" shots, `kling-v3-i2v` (reliable workhorse) for the "body moves" shots.
- **Music-locked editing** — every shot timed to a musical beat (10s = ~2 piano bars at 3/4). The reel finishes exactly when the score does.
- **Cinema 1920×1080 reel** with blurred-fill side bars for vertical-source clips.
- **Optional 9:16 social teaser** from the strongest 3 shots.
- **Cast bible** preserved as a re-pasteable 8-axis spec for future episodes.

**Time**: ~45 minutes wall clock. **Your attention**: ~10 minutes across 2 audit gates. **Cost**: $7–$10 e2e. *This is the premium showcase format* — every shot real i2v, no fallback to stills.

## Tell the agent about the piece

```yaml
piece_title:           # e.g. "OPUS 1" — the artwork's title. Used in cards + reel poster overlay.
piece_subtitle:        # 1-sentence about the piece. e.g. "A 90-second ballet at first light."
protagonist_name:      # The single character's name. Used in cast spec.
protagonist_age:       # Approximate age (e.g. "28") — appears in identity_anchor.
protagonist_role:      # One short noun phrase (e.g. "principal ballerina", "concert pianist", "blacksmith").

# 8-axis cast spec — re-pasted verbatim into every shot prompt.
# Axes 1, 6, 7, 8 are LOAD-BEARING — if you skip any of them, expect drift.
cast_identity_anchor:  # 1 short sentence (≤14 words) — the agent's seed for every prompt.
                       # e.g. "28-year-old Russian-European fair-skinned principal ballerina"
cast_face:             # hair colour + style, eye colour, distinguishing facial features.
                       # e.g. "dark espresso hair pulled into a tight classical low bun, soft blue-grey eyes, calm focused expression"
cast_wardrobe:         # head-to-toe outfit. THIS IS THE COSTUME ACROSS THE WHOLE PIECE.
                       # e.g. "pale dove-grey long-sleeve leotard, soft pink tights, pink pointe shoes with ribbons"
cast_change_policy:    # always "SAME COSTUME across every shot" unless the piece has a wardrobe reveal moment.
cast_expression:       # default emotional state. e.g. "inward-focused, quietly luminous"
cast_silhouette:       # posture/proportions cue. e.g. "classical-trained · long neck · turned-out feet"

# ── Axes 6-8: anti-drift locks (added 2026-06-03 after the OPUS 1 first-pass drift) ──
cast_ethnicity_lock:   # explicit phenotype, with negation. The 6th axis. REQUIRED for non-Anglo cast,
                       # STRONGLY RECOMMENDED for any cast where ambiguity drifts to a wrong reading.
                       # e.g. "Russian-European, fair-skinned with rosy cheekbones, soft blue-grey eyes,
                       #       dark espresso hair. NO non-Western features."
cast_anatomy_lock:     # explicit limb / body-part count, with negation. The 7th axis. REQUIRED for
                       # any pose involving full-body motion (leap, turn, kick, arabesque, split).
                       # e.g. "a single complete human body — EXACTLY TWO LEGS, two arms, ten fingers,
                       #       no extra limbs, no duplicate body parts. Anatomically proportioned."
cast_subject_framing_lock:  # explicit "this is a photograph of a SUBJECT, not a depiction" clause.
                       # The 8th axis. REQUIRED when the style brief contains period-art references
                       # ("Belle-Époque painterly", "Vermeer atelier", "Renaissance", "neoclassical") —
                       # those style cues compete with the subject and the model can drift to
                       # "render a painting of X" instead of "render X in the style of period painters."
                       # e.g. "CINEMATIC PHOTOGRAPHIC STILL of a single human ballerina · NOT A PAINTING ·
                       #       NOT VERMEER · NO pigment / NO easel / NO atelier subject matter."

# Style + setting
visual_style:          # e.g. "Belle-Époque painterly", "Vermeer atelier", "Hopper-light interior",
                       # "Studio Ghibli watercolour", "1950s noir charcoal".
setting:               # 1-2 sentence world description — the LOCATION the piece lives in.
                       # e.g. "an empty Beaux-Arts ballet studio with tall arched windows and faded parquet floor"
light_direction:       # e.g. "north window left, low warm dawn angle, golden beam diagonal across floor"

# Music + pacing
music_brief:           # 1-sentence brief for minimax-music. KEEP IT SPECIFIC.
                       # e.g. "wistful contemplative solo piano in the spirit of Satie's Gymnopédies — sparse, lyrical, 3/4 time"
beats_per_shot:        # 10 seconds = the safe default. Yields 9 × 10s = 90s.
total_shots:           # 9 (default). 6 minimum (60s), 12 max (120s).

# Cap mix (the playbook's secret sauce)
camera_dances_shots:   # array of shot indices where the camera moves dramatically.
                       # e.g. [1, 4, 7]. These get grok-imagine-video.
                       # Every other shot gets kling-v3-i2v (camera holds, body moves).

# Output
aspect:                # cinema 16:9 (default 1920×1080) / vertical 9:16 (1080×1920).
piece_slug:            # kebab-case for filenames. e.g. opus-1.
```

---

## How the agent should run this (interaction contract)

1. **CONFIRM (one turn, no render):** restate in 1 line ("90s beat-locked {protagonist_role} piece, {visual_style}, 9 shots, 3 grok camera-dances") + the 3 choices that matter: **total_shots** (9 default; 6 cuts cost ~$3), **camera_dances_shots** (each grok shot ≈ $1.47 vs $0.53 kling), **aspect**. Quote ~$7–$10 / ~45 min. If axes 6–8 of the cast spec are blank, fill them with sensible defaults and SAY so — don't block; that's the one thing worth a question only when the cast is non-Anglo or full-body motion is involved.
2. **PREVIEW CHECKPOINT = the Wave-1 keyframe audit grid** (9 stills, ~$0.57). All i2v money waits behind this gate: "reply 'go' to animate ($7+) or redo [#]". This playbook's whole economics: a $0.06 keyframe redo vs a $1.50 i2v redo.
3. **NARRATE:** the i2v waves run ~15–25 min — post "S4 (grok) done, S5–S7 rendering, ~8 min left" per shot via async polling; music renders in parallel (say so). Never silent >2 min.
4. **FAIL GRACEFULLY:** grok stall → `kling-v3-i2v` (camera-dance becomes a held shot — say which shots downgraded). kling fail → `ltx-q-i2v` `resolution:"auto"`. minimax-music fail → retry once → `sonilo-t2m` (same brief). Max 2 retries per shot; an 8-shot reel re-timed to the score beats no reel.
5. **DELIVER:** the 1920×1080 reel + 1 honest line ("S7's leap shows minor anatomy wobble — within current i2v ceiling") + ONE next step ("9:16 teaser from S1/S4/S7?").

Gate discipline: exactly the 2 audit gates (Wave 1 keyframes, Wave 2 anims) + the upfront confirm. Nothing else pauses.

## Quality controls (auto-applied)


This is the flagship for cast-lock at single-character scale. The agent applies these by default:

- **Cost upfront.** Before the 9 keyframes fire, the agent shows you: 9 × krea-2-large (~$0.57) + 3 × grok-imagine-video (~$4.41) + 6 × kling-v3-i2v (~$3.18) + 1 × minimax-music (~$0.30) = $8.46 estimated. Nothing runs until you approve.
- **Audit gate after keyframes (Wave 1).** All 9 keyframes laid out in a grid against the cast bible + style brief + body-proportion checklist. The agent flags any face drift, costume drift, or anatomical break BEFORE any i2v fires (each i2v failure = $0.50–$1.50 wasted).
- **Cast spec re-pasted verbatim.** The 8-axis cast spec is appended to every single keyframe AND every single i2v prompt — not summarized, not paraphrased.
- **B9 motion budget on every shot.** Each i2v prompt has exactly ONE allowed motion (one blink, one breath, one ribbon-stir, one held leap). Multi-motion prompts cause cast drift; this is the strict rule.
- **Audit gate after animations (Wave 2).** All 9 anims reviewed against cast-lock invariant — flagged: same face? same costume? same eye-line? Nothing concatenated until pass.
- **Beat-locked finishing.** Every shot trimmed to the exact beat count. Music plays through, fades in over 0.4s, fades out over 2s. No narration competes with the score.

> **Why two caps instead of one:** grok-imagine-video is ~$0.15/s and does fluid camera moves no other cap can. kling-v3-i2v is ~$0.05/s and does held shots reliably. The right cap per shot type costs ~$1 less per shot than running grok everywhere — saving ~$5 across the 6 held shots — without sacrificing cinematic quality.

---

## The 9-shot beat sheet (canonical)

| Shot | Role | Beat | Cap | Motion budget |
|------|------|------|-----|---------------|
| S1 | OVERTURE | wide establishing — camera enters | `grok-imagine-video` | camera push-in, light brightens 5% |
| S2 | WARM-UP | mid two-shot — at the work surface | `kling-v3-i2v` | one extension + return |
| S3 | CARE | macro — hands or tool detail | `kling-v3-i2v` | one slow stroke, one glint |
| S4 | LIFT | wide — the hero held pose | `grok-imagine-video` | camera arcs around held pose |
| S5 | TURN | mid — body moves through frame | `kling-v3-i2v` | body crosses, camera holds |
| S6 | BREATH | portrait — pause at window or threshold | `kling-v3-i2v` | one slow blink across 10s |
| S7 | LEAP | wide — the hero motion shot | `grok-imagine-video` | camera floats with motion |
| S8 | REVERENCE | mid — the controlled finish | `kling-v3-i2v` | arms unfold, no body shift |
| S9 | CODA | portrait — held final | `kling-v3-i2v` | one breath, light brightens |

Substitute role labels for your piece type — OVERTURE/CARE/LEAP/CODA work for ballet, music, craft, sport. The shape is: open wide, get intimate, hit a hero shot mid-arc, finish held.

---

## Reproduce via MCP

The pipeline:
1. `brand_kit_create` — palette + style keywords. (Optional but stabilizes look across re-runs.)
2. `generate_project({ project_id, scenes: 9, style: visual_style })` — creates project shell.
3. 9 × `create_media({ action: "generate", model_override: "krea-2-large", prompt: STYLE + CAST_SPEC + SHOT_BEAT })` parallel.
4. Audit gate (your attention).
5. For each shot index in `camera_dances_shots`: `create_media({ action: "animate", model_override: "grok-imagine-video", source_url: keyframe_url, prompt: MOTION_BUDGET + CAST_SPEC, resolution: "auto", duration: 10 })`.
6. For all other shots: same but `model_override: "kling-v3-i2v"`. (Always pass `resolution: "auto"` — fixed resolutions 422 on the ltx fallback tier.)
7. `create_media({ action: "music", prompt: music_brief, lyrics_prompt: "[Instrumental]" })` parallel with step 5/6. **Critical:** beat-driven-opera scores are instrumental — pass `lyrics_prompt: "[Instrumental]"` explicitly. NEVER omit it; minimax-music sings whatever is in `lyrics_prompt` verbatim. (If you DO want vocals, pass the actual libretto lyrics in `lyrics_prompt` with [Verse]/[Chorus] markup. See `public/playbooks/music-video-mtv.md` §"The lyrics-vs-style bug".)
8. Audit gate (your attention).
9. Finish: `create_media({ model_override: "ffmpeg-concat", ... })` to stitch the 9 beat-trimmed shots (vertical-source shots blurred-fill to 1920×1080 via `ffmpeg-export`), then `create_media({ model_override: "ffmpeg-mux", source_url: <concat_url>, audio_url: <music_url>, audio_fill: "loop" })` so the score plays edge to edge (no silent tail). Fade in 0.4s / out 2s. → final 1920×1080 cinema mp4.

---

## Honest limits — what can disappoint (cap ceilings)

- **i2v anatomy + identity drift is the ceiling, not a bug in your prompt.** Full-body motion shots (LEAP, TURN) are where extra limbs and face morphing appear — that's exactly why axes 6–8 and the one-motion budget exist. Even with them, expect ~1 in 9 shots to need a redo.
- **"Beat-locked" means trimmed-to-beat-math, not audio-detected.** No beat-detection cap is live; shots are trimmed to `beats_per_shot × 60/BPM` from the music brief's declared tempo. At 10s shots over solo piano, nobody can tell.
- **Single protagonist only.** Two-character beat-driven pieces (a pas de deux, a duet) increase cast-lock complexity by ~3x and need a different playbook (episodic-detective-pilot.md).
- **Cap warm-up matters.** grok-imagine-video is reliable as of 2026-06-02; if it stalls in your run, the FALLBACK_CHAINS route to kling-v3-i2v automatically — you lose some camera-dance quality on the affected shots but never lose the reel.
- **Music duration is a soft cap on shots.** minimax-music reliably generates 90s+ even when you pass `duration: 15`, but if your piece runs >120s consider stitching two music calls.
