⭐ Storyboard MCP · Fusion flagship

THE GATE one building. eight alt-histories.

A single ceremonial gateway — three openings, three stories tall, same silhouette + scale + camera angle + golden-hour lighting held constant — rendered through eight curated east-meets-west cross-cultural alt-history fusion lenses. Forbidden City silhouette meets Roman triumphal carving. Japanese torii meets Gothic cathedral tracery. Mughal jali meets Venetian palace mosaic. The location bible AI does that humans cannot — that no concept-art lead could produce in a sprint.

cells
8 east-meets-west
aspect
16:9 cinema
reel
40.0 seconds
through-line
silhouette + camera + light
variable
ornament + material + setting
The walkthrough · 1920×108040.0s · 8 cells × 5s · slow pan + atmospheric score

The 8 cells — click to view full resolution

Every cell uses the same building silhouette spec, the same 35mm establishing-shot framing, the same golden-hour key light from upper-right at 35°. The SILHOUETTE is the through-line; the cultural ornament + material + adjoining setting are the variable. The viewer reads the gallery as one building seen through 8 alt-history cultural lenses — defensible game concept art, not random AI fantasy.

Why one-silhouette-eight-worlds is the load-bearing creative choice. Game design + film concept art often need "what if this culture had built it" alt-history pitches. By hand that requires deep architectural training in every culture — months per cell. This collapses it to one afternoon: pick the silhouette, pick the pairings, run. The same building reads as the same place across radically different ornament + material vocabularies, which makes the pitch defensible to a creative director rather than read as random AI fantasy.

Fusion grammar — what gets blended where

01
Forbidden City × Roman
Yellow glazed tile + red lacquer column silhouette meets carved Roman triumphal-arch relief panels in Carrara marble.
02
Japanese torii × Gothic
Vermillion cypress crossbeam silhouette meets pointed-arch tracery, flying buttresses, rose-window oculus.
03
Mughal jali × Venetian
Carved white-marble jali screens meet pink Verona-marble pylons + Byzantine gold-tesserae mosaic spandrels.
04
Khmer Angkor × Greek
Stone-warm bas-relief apsara carvings + naga balustrade meet Doric columns + triangular pediment + entablature metopes.
05
Tibetan stupa × Romanesque
Golden dharmachakra wheel medallion + maroon/saffron pylons meet rounded Romanesque arches + carved tympanum sculpture.
06
Persian iwan × Italian Renaissance
Towering pointed iwan + cobalt Persian tilework meets Brunelleschi-tier classical orders + terracotta brick spandrels.
07
Korean iljumun × English Tudor
Hipped-roof dougong bracket complex + green tile roofing meets half-timbered oak framing + Tudor-arched openings.
08
Ottoman muqarnas × German Baroque
Stalactite muqarnas honeycomb-vaulting + Iznik turquoise tilework meets copper-clad onion domes + carved Baroque scrollwork.

Pipeline trace — what the storyboard MCP actually did

Wave 1 — silhouette constant + 8 fusion establishing shots 8 × create_media action=generate, model=gpt-image, aspect_ratio=16:9 8/8 done in ~130-140s each, ZERO moderation flags Each prompt = same silhouette constant + same 35mm golden-hour composition + fusion-specific block. Shape-stability bias works.
Wave 2 — slow pans + civic-monumental score 8 × create_media action=animate, model=kling-v3-i2v, duration=5, prompt="very slow horizontal pan, structure entirely still, atmospheric haze drifts" 8/8 done in ~55-100s each 1 × create_media action=music, model=minimax-music, prompt="slow organ pad + soft percussion + distant single-bell tolls, civic-monumental arrival" done
Wave 3 — local stitch ffmpeg normalize: kling 1108×828 → 1920×1080 cinema (scale + crop) ffmpeg concat: 8 clips → single 40s timeline (concat demuxer, no re-encode) ffmpeg mix: score @ 0.95 vol + clip ambient ducked to 0.22, afade in/out output: the-gate-reel.mp4 (18MB, 40.0s)
Quality controls applied ✓ Building silhouette constant repeated VERBATIM in every cell (shape-stability bias) ✓ Composition constant (35mm, golden-hour upper-right) verbatim in every cell ✓ Same scale + camera angle + time-of-day across all 8 cells (reads as ONE building, 8 worlds) ✓ Curated "X silhouette × Y ornament + Z setting" fusion grammar — no kitsch ✓ NO readable text guard on every prompt (defense against AI text-rendering failures on inscriptions/plaques) ✓ Architectural plausibility constraint (load-bearing structure where it goes) ✓ Slow pan only on i2v (no aggressive zoom, no roof-tile motion, no banner-flutter) 17/17 jobs succeeded first try · ZERO moderation fallbacks · ZERO infra errors

Cost — total MCP run end-to-end

Wave 1 keyframes
~$1.60
8 × gpt-image @ ~$0.20
Wave 2 pans
~$4.56
8 × kling-v3-i2v @ ~$0.57
Music score
~$0.40
1 × minimax-music
Total
~$6.56
under the $9 flagship target

Reproduce this — the playbook is open

This reel was built from the fusion-architecture-walk playbook — the same template can produce ANY single-building bible: a harbor lighthouse across 8 maritime cultures, a town-square clocktower across 8 civic traditions, a citadel approach across 8 fortification eras. Two sister playbooks: fusion-character-portrait (one face, N cultural lenses — see THE EMISSARY) and fusion-garment-runway (one silhouette, N ornament fusions — see THE ROBE COAT).