---
reliability: 4.2 # 5 − image gen .3 − i2v .6 − music .2 − 1 stage over 5 (.2) + proven E2E showcase (the-robe-coat, 25/25 first-try) .5
---

# Fusion garment runway — agent playbook ⭐ flagship

> Paste this whole file (with the BRIEF filled in) into Claude cowork, chat, or Code. **One hero garment silhouette. N cultural ornament fusions. A 60s runway reel.** Build a fashion / costume bible where the SAME garment silhouette (a robe, a coat, a gown, a tunic) is rendered through N curated east-meets-west cultural fusion lenses — silhouette held constant, ornament + textile + accessories variable — and then animated into a continuous runway reel with score. Useful for fashion line development, costume design pitches, and cross-cultural editorial shoots.

## What you'll get

A costume bible structured the way a high-fashion atelier presents a capsule collection — but with cross-cultural breadth no single atelier could produce in a season:

- **A single garment silhouette held constant** — same cut, same length, same drape, same model proportion across every cell. The silhouette is the through-line; the cultural ornament + textile + accessories are the variable.
- **N fusion cells** (default 10 — runway-sized, neither short nor exhausting) — each cell pairs one Eastern textile/ornament tradition with one Western, with a defensible fusion grammar (which element gets blended where: silhouette vs textile vs ornament vs accessories vs setting).
- **Vertical 9:16 portrait** — full-figure model, runway-shaped. Pass `aspect_ratio: "9:16"` to `generate_project` so every look is composed native-vertical, not center-cropped.
- **Runway reel** — each look gets a slow turn via `seedance-i2v` (`resolution: "auto"`, 6s, "model very slowly turns from three-quarter view to face camera, garment drape stays still"), `ffmpeg-concat`-stitched into a 60s runway sequence, then `ffmpeg-mux` with `audio_fill: "loop"` so the score covers the WHOLE reel.
- **One unified score** — runway-tier instrumental via `music` ([Instrumental] default — low pulsing bass + bright string flourish on each transition), the score of a Paris/Tokyo crossover capsule launch.
- **Showcase HTML page** — single shareable page with the 10-look grid + runway reel + cell-by-cell fusion breakdown + pipeline trace + cost summary.

**Time**: ~14–18 minutes wall clock. **Your attention**: ~7 minutes across 2 gates. **Cost**: $3 stills-only / $10 stills + reel + score for 10 looks. *This is the editorial-reel flagship* — the playbook for fashion-week-tier cross-cultural capsule presentations.

## Tell the agent about the garment + the fusions

```yaml
garment_archetype:         # e.g. "floor-length robe coat" / "structured a-line gown" / "wide-sleeved tunic + sash" / "fitted high-collar dress"
garment_silhouette:        # The verbatim block repeated in every cell prompt to hold silhouette + drape constant. Example:
                           #   "Floor-length robe coat, open front, wide bell sleeves to wrist, vertical drape with column silhouette,
                           #    fastens at high collar. Modeled on a slim female figure, age mid-20s, dark hair pulled back simply,
                           #    full-figure runway pose, three-quarter view to camera."
composition_constant:      # Verbatim composition rule across all cells. Example:
                           #   "Editorial full-figure runway shot, 85mm equivalent, model occupies 80% of frame height. Soft single
                           #    cool key light from upper-front at 20°, neutral shadow at feet. Studio cyclorama background slightly
                           #    soft. Painterly photoreal high-fashion editorial quality. NO readable text anywhere in frame."
cell_count:                # 8 (tight) / 10 (default) / 12 (extensive). 1 cell per fusion pairing.
cells:                     # Ordered list of N east-meets-west fusion pairings. Each item:
                           #   - east_lens:        Ming dynasty silk damask (cinnabar + phoenix + cloud)
                           #     west_lens:        Vermeer Dutch tonal warmth (cream + ochre + muted gold)
                           #     fusion_grammar:   Silk damask textile × Vermeer warm palette + cream cyclorama
fusion_quality_guards:     # Default ON:
                           #   - No readable text on labels / signage in frame (text rendering failure)
                           #   - No "cyberpunk samurai kimono" stock-mashup phrasing
                           #   - Same silhouette + same model proportion + same pose across every cell
                           #   - Single-model frames only (no companion figures, no crowd extras)
                           #   - Fabric construction plausible (no floating shoulders, no impossible drape)
include_reel:              # true (default) — fire N × seedance-i2v slow turns + 1 music track, total ~$10 for 10 looks
                           # false           — stills-only, total ~$3
music_brief:               # Score mood arc. Example:
                           #   "High-fashion runway instrumental, low pulsing bass + bright string flourishes on each look transition,
                           #    building tension across 60s. No vocals."
runway_slug:               # kebab-case for output filenames + showcase page. e.g. the-capsule / the-robe-coat / the-tunic.
```

---

## How the agent should run this (interaction contract)

1. **CONFIRM (one message, ≤1 question):** restate the brief in 1 line ("10-look <garment> east-meets-west runway, 9:16") + the real choices — stills-only ~$3/~6 min · stills+reel+score ~$10/~15 min · lean 8 cells. Ask only if `cells`/`garment_silhouette` are missing; otherwise state assumptions and start.
2. **PREVIEW CHECKPOINT:** the Wave-1 contact sheet (~$2 of stills) IS the gate — show all N looks and invite `director_re_render` re-fires BEFORE any seedance i2v fires. Never animate unreviewed stills.
3. **NARRATE:** 1 line + ETA per wave ("Wave 1: 10 gpt-image cells, ~3–5 min, cjob_…"); poll `get_creative_job`/`get_create_media` every 10–15s; post an update at least every 2 min while jobs run.
4. **FAIL GRACEFULLY:** gpt-image 422 → soften the cultural phrasing once → `nano-banana`; seedance-i2v hang/"fetch failed" → retry once → `ltx-q-i2v` (`num_frames: 241`, resolution "auto"). Each failure = 1 line (WHAT/WHY/fallback). ≤2 retries per cell; a 9-look reel with a named dropped cell beats no reel.
5. **DELIVER:** muxed reel + contact sheet + one honest quality line ("look 7's drape went boxier in the turn — i2v fabric-fidelity ceiling") + ONE next step ("re-fire look 7, or publish the showcase page").

## What can disappoint (cap ceilings)

- **Silhouette drift** — the verbatim block biases but doesn't lock; expect 1–2 cells per 10 to need a re-fire.
- **i2v turns alter drape/face slightly** — seedance preserves ~90% of the still; flowing fabric is the weak spot.
- **No readable text, ever** — labels/tags in frame will garble; the guard exists because the failure is guaranteed.
- **Music loop seam** — a 15–30s bed looped to 60s can have an audible seam at the joins.

## Quality controls (auto-applied)

The closed quality loop that separates this from "a folder of fantasy fashion":

- **Brief approval first.** Before any MCP fires, the agent shows you the garment silhouette spec + the N cell pairings + the constants for review.
- **Silhouette + proportion constant.** The `garment_silhouette` block is repeated verbatim in every cell prompt, biasing gpt-image toward shape + drape stability. Without it, each cell drifts to a different garment.
- **Composition + pose constant.** Same focal length, same model pose, same camera angle across all cells. The viewer reads the gallery as one garment in 10 ornament treatments, not 10 unrelated dresses.
- **Curated pairings.** Each cell uses "X textile × Y palette + Z accessory" fusion grammar. "Ming silk damask × Vermeer warm palette" is curated; "Asian × European fashion" is kitsch.
- **No readable text guard.** Defense against AI text-rendering failures on labels / tags / signage in frame.
- **Fabric plausibility.** Prompts request fabric that drapes the way the named textile drapes (heavy silk damask = column drape; sheer chiffon = floating drape). No floating-shoulder impossibilities.
- **gpt-image as the renderer.** Best stylization fidelity for named textile traditions and named-painter palette references.
- **Mid-render check.** After Wave 1 lands, the agent surfaces a contact-sheet for your review. You can re-fire any off-silhouette or implausible-drape cell before the i2v wave fires.
- **Slow-turn only on i2v.** Turn prompts say "model turns slowly only, garment drape stays mostly still, no aggressive fabric motion". Wind/billow animation would break the editorial aesthetic.
- **Audio covers the full reel.** The 60s reel is muxed with `ffmpeg-mux` + `audio_fill: "loop"` so a 15–30s score repeats under the whole runway — never a silent tail.

## How the agent builds it (real verbs)

1. **Cells (Wave 1):** `generate_project({ aspect_ratio: "9:16", model_override: "gpt-image" })` — one scene per cell, each prompt = `garment_silhouette` + `composition_constant` + the cell's fusion grammar verbatim. Poll `get_creative_job` to `done`.
2. **Mid-render gate:** surface the contact sheet; re-fire any off-silhouette cell with `director_re_render({ project_id, scene_index, prompt })`.
3. **Reel (Wave 2, only if `include_reel`):** for each look, `create_media({ action: "animate", model_override: "seedance-i2v", source_url: <look.png>, resolution: "auto", duration: 6, prompt: "model very slowly turns from three-quarter view to face camera, garment drape stays still" })`. **Cap in-flight i2v at ≤4–5; poll to `done` before firing the next batch** (BYOC video caps are capacity ~2 behind one dispatch worker).
4. **Score:** `create_media({ model_override: "music", prompt: music_brief })` — `[Instrumental]` is the default; only add a `lyrics_prompt` if you want vocals.
5. **Finish:** `create_media({ model_override: "ffmpeg-concat", clips: [<look-01.mp4>, …] })` → `create_media({ model_override: "ffmpeg-mux", source_url: <concat.mp4>, audio_url: <music.mp3>, audio_fill: "loop" })` → `create_media({ model_override: "ffmpeg-export", source_url: <muxed.mp4> })` for final 9:16 1080×1920.

**Cross-surface:** CLI runs the same chain via `livepeer story … --aspect 9:16`, per-look `livepeer media animate --model seedance-i2v`, then `livepeer media concat`/`mux --audio-fill loop`/`export`. Webapp chat: "build a 10-look east-meets-west robe-coat runway, 9:16, with a runway score."

> **Why one-silhouette-ten-worlds works:** fashion development pitches benefit from seeing the SAME silhouette explored across cross-cultural ornament traditions — it shows the silhouette has structural strength while the ornament + textile vocabulary is being interrogated. Doing this by hand requires deep textile + ornament training in every culture — weeks per look. The fusion playbook collapses it to one afternoon: pick the silhouette, pick the pairings, run.

## Output structure

```
public/chapters/<runway_slug>/
  look-01-{east}-{west}.png            ← 10 keyframe editorial portraits
  ...
  look-01.mp4                          ← (optional) 10 i2v slow turns
  ...
  music.mp3                            ← (optional) unified runway score
  <runway_slug>-runway.mp4             ← (optional) 60s runway reel
public/chapters/<runway_slug>-example.html  ← showcase page
public/case-studies/<runway_slug>.html      ← optional showcase entry
```

The example for this playbook is **THE ROBE COAT** — a single floor-length robe-coat silhouette rendered through 12 east-meets-west royal-court fusion lenses. **Live example shipped 2026-05-30**: [`/chapters/the-robe-coat-example.html`](https://storyboard.daydream.monster/chapters/the-robe-coat-example.html) — 12 vertical 9:16 runway looks + 60s slow-turn runway reel with low-pulse-bass score, 25/25 jobs succeeded first try, ~$9.64 end-to-end.
