---
tier: hero
reliability: 3.9 # story-driven narrative (12-16 shots) + locked product/character + i2v coverage + grade/score/mix + multi-format cutdowns; -.x long-form drift, hero human performance — gate every shot
---

# Brand Film — an agency-grade commercial set, any brand ⭐

> Paste this whole file (with the BRIEF filled in) into Claude cowork, chat, or Code with the Storyboard MCP attached. **Load the `brand-film-director` skill first** — it designs the story, the style lane, and the length set; this playbook produces them. Hand it a product photo + a brand and it returns a **hero brand film (60–90s) + a 30s broadcast cut + 6s/15s social, multi-aspect, multi-language** — product locked, consistency-gated, sound-designed, endings resolved.

## What you'll get

- A **60–90s hero brand film** with a real story arc (one human truth → tension → turn → payoff → brand stamp).
- A **30s broadcast cut** (re-cut for the format) + **15s / 6s** social, **16:9 / 9:16 / 1:1**, in **N languages**.
- The product **provably identical** across every shot; brand text **legible by construction**; **resolved** audio.
- A speed / cost / adaptiveness report for the room.

**The wow:** this is the tier where AI commercials win — emotion and narrative, not features. The pipeline is brand-agnostic: only the brief, product photo, brand kit, and style lane change.

## The caps

| Step | Cap |
|---|---|
| Lock the product (real photo) | the brand's photo → `place_subject` / `kontext-edit` |
| Lock the hero character | `gpt-image` / `nano-banana` → reused via `kontext-edit` |
| Coverage shots | `seedance-i2v` (≤15s beats, cost-efficient) / `kling-v3-turbo-i2v` |
| Hero / premium beats | `ray-32-i2v` (Ray 3.2 HDR) — optional premium tier |
| Score (licensed) | `sonilo-t2m` |
| Voiceover (multilingual) | `gemini-tts` |
| Spokesperson lip-sync | `talking-head` (for VO-on-camera moments) |
| Grade · overlay · stitch · mix | `ffmpeg` + PIL (deterministic, legible text) |

## Tell the agent about the film

```yaml
mode:           # express (vague idea → e2e, least involvement) | director (options + surprises + iteration). Default: express.
brand:          # e.g. "Nike"
product:        # e.g. "running shoe"
product_photo:  # URL of the real product (the pipeline locks it); omit to generate a hero
style_lane:     # athletic-kinetic | luxury-intimate | tech-minimal-wonder | warm-human  (express auto-picks)
human_truth:    # the feeling the brand owns (the story core)  (express auto-picks)
tagline:        # OPTIONAL original line (NOT a trademarked slogan); else the agent writes one
hero_length:    # 60 | 75 | 90 (seconds)  (express defaults to 75)
markets:        # languages, e.g. [English, Spanish, Japanese, Chinese]
tier:           # standard (kling/seedance) | premium (Ray 3.2 hero beats)
launch_slug:    # e.g. nike-keep-going
```

> **Express needs only `brand` + `product` (or even just a product noun).** Everything else has a strong default the agent fills in. Director can set any field and will be offered options for the rest.

## Two modes — pick your involvement

**Express (least involvement — how the Nike demo was made).** Give it a brand + a product (a photo, or just "a running shoe"). The agent makes every creative call with strong defaults and runs **end-to-end** to the full set. It pauses for exactly one thing — the **consistency gate** — and otherwise just delivers, then reports speed / cost / adaptiveness.

> *"Nike, a running shoe, make me a hero film."* → a full set, no other questions.

**Director (raise the bar — for pros with options).** At each decision point the agent proposes **2–3 distinct directions plus one *surprise*** (a bolder idea you didn't ask for) and waits for your pick or edit before spending. **Iterate freely** — re-roll any beat, swap the lane, re-cast the voice, re-time a cut, fork a variant. Involvement is **dial-able**: answer every prompt, or say *"you choose from here"* to drop back to Express mid-run.

> Decision points (express auto-picks each; director offers options + a surprise at each): **story angle · lane · cast (voice+music) · length set · per-beat shots · grade · ending.** Same gate and same quality bar in both modes — only the number of human checkpoints changes.

## How the agent should run this (interaction contract)

1. **DESIGN (via the director skill).** Write the one human truth → arc → beat sheet (8–16 beats) → VO script → shot list (subject, motion, lens, emotion). Lock the style lane. *Express:* pick the strongest and proceed. *Director:* propose 2–3 angles + one surprise; wait for the pick.
2. **LOCK.** Lock the real product photo + a hero-character reference. Register a brand_kit (palette + wordmark) so type/colour stay on-brand. **Decide the branding treatment now and write it down** (real licensed logo for authorized work; consistent original unbranded design for demos) — it must be identical in every shot.
3. **KEYFRAMES.** Generate each shot's keyframe: product beats via `kontext-edit` from the locked product; human beats via `kontext-edit` from the locked character. **For a WORN/small product (jewelry, watch, eyewear, shoe) the exact piece will NOT survive a text description** — do one of: (a) **anchor the close beat on the product image** (base = product, build the ear/wrist/face around it, frame tight); or (b) **deterministic composite** — `bg-remove` the product → overlay the cut-out onto the body part (gravity-vertical) + soft shadow → animate that still with **Ken-Burns, not i2v** (i2v drifts a small worn product). Apply the lane's lens/grade cues. **Best-of-N + taste-judge** on hero beats.
4. **CONSISTENCY GATE (mandatory — the #1 quality factor; the one checkpoint Express also stops for).** Grid **the locked product + every standalone product shot + every shot where the product is worn/used** (not just the close-ups — the classic miss) + the locked character + every human placement. Verify TWO things in all of them: **(a) colour/material** matches the locked source, **(b) branding/logo** is consistent — present-or-absent the same way everywhere (no logo on the worn product but blank on the hero shot, and vice-versa). Fix drifters at the cheap image stage before any video. Scan for AI-mangled text.
5. **ANIMATE.** i2v each keyframe (≤5–8s, restrained motion that serves the beat). `seedance-i2v` for coverage; `ray-32-i2v` for hero beats if premium. Re-check end frames (drift).
6. **ASSEMBLE THE HERO.** Cut to the arc; beat-cut to the music. Color-grade pass. Legible overlay (wordmark / line / CTA). `sonilo-t2m` score + `gemini-tts` VO, ducked mix, **resolved ending** + brand card.
7. **CUTDOWNS (purpose-built).** Re-cut a 30s (not a trim) + 15s/6s social; conform 16:9 / 9:16 / 1:1; localize VO per market, each cut sized so the voice finishes.
8. **REPORT + DELIVER.** Speed (wall-clock), cost (itemized cap calls × rates), adaptiveness note. Ship the set + the consistency proof + one honest line.

## What can disappoint (cap ceilings) — guardrails

- **Consistency breaks two ways — colour/material AND branding/logo.** This is the #1 thing to get right. Lock ONE product image; edit-not-regenerate into every beat; gate the product **as worn / in context**, not just the close-ups; decide the branding once and apply it identically (real licensed logo for authorized work, consistent unbranded original for demos — never reproduce a trademark). If you fix one (e.g. recolour to volt-green) re-gate for the other (logo present on the worn shoe but not the product shot).
- **Worn / small products are the hard case — get the EXACT piece, not a similar one.** Text-described worn products always render a *lookalike*; naive multi-image swaps the person; gpt-image-edit is flaky. PRIMARY: product-anchored tight close-ups (generate the beat FROM the product image, scene built around it). Deterministic composite is a last resort and only with precise attach-point detection + erasing the original (blind coordinates fail — earring on the cheek, old piece showing through); then Ken-Burns, not i2v. **Gate worn beats against the locked product, not the prompt** — "similar" is a fail.
- **Long-form drift** — keep shots ≤5–8s; lock product + character; gate every shot; let the story spine (not raw continuity) carry it.
- **Human performance** — lip-sync VO moments; locked refs elsewhere; composite a real talent plate for the absolute hero shot.
- **Voice longer than the cut** — size each language's video to its VO + ~1.5s.
- **Text** — overlay only, never generated.
- **Ray cost** — premium; prototype one hero beat before fanning.

## Realistic cost (per brand, full set)

Itemize the actual calls (see brand-film-director §7). A 90s hero + 30s + 6s + 4 languages, **standard tier**, lands roughly **$15–35** of compute (image keyframes + seedance/kling coverage + 90s licensed score + multilingual VO + local finishing); **premium (Ray hero beats)** adds ~$10–15. ~1–2 hours wall-clock — vs weeks and six figures for a shoot.

## When you're done

```
✅ Campaign: public/launches/{launch_slug}.html
✅ Hero (60–90s) · 30s · 15s/6s · multi-aspect · N languages + consistency proof
✅ Story-driven · product locked & gated · legible text · audio resolved
Speed: ~X min · Cost: ~$Y itemized · Adaptiveness: same pipeline, new brand
```

## What to do next
1. **Next brand** — swap brief + product photo + kit + lane; rerun.
2. **Add markets** — one `gemini-tts` pass per language.
3. **Final-mile for a paying brand** — real talent plate for the hero, bespoke score, brand-exact product plates.

---

## Notes for the agent (only read if a step fails)

**Consistency.** Lock ONE product image + ONE character image; `kontext-edit` each into every shot ("keep identical; restage only the scene"); build + look at the grid before animating.

**Longer beats.** `seedance-i2v` supports longer durations (cost-efficient coverage); reserve `ray-32-i2v` (`duration` STRING `'5s'`) for hero beats. Heavy caps → direct-SDK `/inference` with `timeout:700`.

**Endings.** `afade` audio out ~1.5s before end; size each language cut to VO+1.6s; land on a held brand card with a short video fade. Verify the final 0.4s is near-silent.

**Overlay/grade.** Legible text via PIL PNGs + ffmpeg `overlay`+alpha `fade` (reliable) or `hyperframes-render`. Color grade with ffmpeg `curves`/`eq`/`colorbalance` per the lane.
