⭐ Storyboard MCP · Local-food-short flagship

麻婆豆腐One face. One dish. One language. Thirty seconds.

A native-Mandarin food short, voiced in a soft Chengdu inflection, captioned in Chinese, scored with a single guzheng — built end-to-end via the storyboard MCP in roughly ten minutes for under ten dollars. Ready to upload to Douyin, TikTok, YouTube Shorts, or Instagram Reels with no further editing. This is what AI does that a single creator with a phone cannot: produces a regional-cuisine reel in the audience's own language, at documentary-tier quality, on demand.

runtime
30.0 seconds
language
zh-CN · Mandarin
aspect
9:16 vertical · 1080×1920
dish
麻婆豆腐 · Chengdu, 1862
voice
gemini-tts · 软, 慢, 温
The short · 1080×192030.0s · 6 shots × 5s · VO + 字幕 + 音乐
软的,慢的,温的。 这不是一份食谱,也不是一篇科普——是一位老朋友,坐在你对面,温柔地说起她家的味道。AI能做的,正是这个:用你的语言,讲一道你熟悉的菜,讲得不急不躁,讲得就像有人在替你想念家。 Soft, slow, warm. Not a recipe and not an explainer — a friend across the table, telling you tenderly about the taste of her home. What AI does here: tell a dish in your language, unhurried, like someone remembering home on your behalf.

Voice-over script — in Mandarin and English

The script was written natively in Mandarin (not translated from English) — opens with sensory invitation, uses metaphor over enumeration, ends with "China's gentlest side". Voice direction sent to gemini-tts: warm female Mandarin narrator, slight Chengdu Sichuan inflection, conversational documentary pace, soft warm intimate tone, slight smile in the voice, no rush, breath room between phrases.

中文原文 · Native Mandarin

你闻到了吗?那一缕麻,一丝辣,从成都的小巷飘来。
一百六十年前,陈麻婆的灶台上,多了这一锅红——
豆瓣酱的醇厚,花椒的麻香,碎肉的酥脆,还有那一块块,颤巍巍的嫩豆腐。
没有华丽的辞藻,只有一个家的味道。
一勺米饭,一口豆腐——
这就是中国,最温柔的那一面。

English translation

Can you smell it? That hint of numbing, that whisper of spice, drifting from the alleys of Chengdu.
A hundred and sixty years ago, on Chen Mapo's stove, a new pot of red appeared —
the richness of doubanjiang, the numbing aroma of Sichuan peppercorn, the crispness of minced meat, and those trembling cubes of silken tofu.
No fancy words — just the taste of home.
A spoonful of rice, one bite of tofu —
this is China's gentlest side.

The 6 shots — 5 seconds each, same warm key-light direction across all

The hero bowl
01 · 0–5s你闻到了吗?
Ingredient flat-lay
02 · 5–10s成都 · 1862
Wok action
03 · 10–15s陈麻婆的一锅红
Tofu slide
04 · 15–20s颤巍巍的嫩豆腐
Chopsticks lift
05 · 20–25s家的味道
The tasting moment
06 · 25–30s中国最温柔的一面

Quality controls applied — the review-task that ran before Wave 2 fired

The reel is the second pass. The first script draft was declarative and encyclopedic ("陈麻婆于1862年所创,凝聚六个字…"); the review-task replaced it with a sensory, conversational opener ("你闻到了吗?") and an intimate close ("中国最温柔的那一面"). The first i2v prompt draft pushed for visible cooking motion; the review-task softened them to gentle steam drift + slow chopsticks lift + the breath of a closed-eyes tasting smile. Both decisions are baked into the playbook so the next dish — a Hanoi pho, a Roman carbonara, a Bangkok som tum — gets the same warmth-first review-gate before any expensive render fires.

Pipeline trace — what the storyboard MCP actually did

Wave 1 — 6 keyframes (gpt-image, 9:16 vertical) 6 × create_media action=generate, model=gpt-image, aspect_ratio=9:16 6/6 done in ~112–135s each, ZERO moderation flags Same warm key-light upper-left at 30° in every prompt. NO readable text guard.
Review gate — refine script + i2v prompts for warmth Script v1 (declarative) v2 (sensory, intimate). Caption blocks softened. i2v prompts say "gentle", "soft", "tender" — no aggressive cooking motion.
Wave 2 — 6 i2v + native voice + warm music score 6 × create_media action=animate, model=kling-v3-i2v, duration=5, gentle motion only 6/6 done in ~55–62s each 1 × create_media action=generate, model=gemini-tts, prompt=Mandarin script + voice direction done in 21s, 31.92s narration 1 × create_media action=music, model=minimax-music, prompt=warm guzheng + hand-percussion done in 55s
Wave 3 — local stitch (ffmpeg + Pillow) ffmpeg normalize: kling 1108×828 → 1080×1920 portrait (scale + crop) ffmpeg concat: 6 clips → 30s timeline (concat demuxer, no re-encode) Pillow renders 6 CJK caption PNGs (Heiti SC 78pt, white with dark drop-shadow) ffmpeg overlay: 6 captions on video with `enable='between(t,a,b)'` timing ffmpeg mix: voice @ 1.0 + music ducked to 0.22, afade in/out output: mapo-tofu-short.mp4 (21MB, 30.0s)
All 14 MCP jobs succeeded first try · ZERO moderation fallbacks · ZERO infra errors

Cost — total run end-to-end

Wave 1 keyframes
~$1.20
6 × gpt-image @ ~$0.20
Wave 2 i2v
~$3.42
6 × kling-v3-i2v @ ~$0.57
Voice + music
~$0.50
gemini-tts + minimax-music
Total
~$5.12
under the $6 flagship target

Reproduce in any language — the playbook is open

This reel was built from the local-food-short playbook. The same template works for any famous local dish in any language gemini-tts supports — Pho Bo voiced in northern Vietnamese, Carbonara voiced in Roman Italian, Bibimbap voiced in Seoul Korean, Pad Thai voiced in Bangkok Thai, Tagine voiced in Moroccan Arabic. Paste the playbook into Claude with your YAML config (language + dish_name + key_flavors + voice_brief), and the same 3-wave + review-gate pipeline runs again — different dish, different tongue, same warmth.