A native-Mandarin food short, voiced in a soft Chengdu inflection, captioned in Chinese, scored with a single guzheng — built end-to-end via the storyboard MCP in roughly ten minutes for under ten dollars. Ready to upload to Douyin, TikTok, YouTube Shorts, or Instagram Reels with no further editing. This is what AI does that a single creator with a phone cannot: produces a regional-cuisine reel in the audience's own language, at documentary-tier quality, on demand.
The script was written natively in Mandarin (not translated from English) — opens with sensory invitation, uses metaphor over enumeration, ends with "China's gentlest side". Voice direction sent to gemini-tts: warm female Mandarin narrator, slight Chengdu Sichuan inflection, conversational documentary pace, soft warm intimate tone, slight smile in the voice, no rush, breath room between phrases.
你闻到了吗?那一缕麻,一丝辣,从成都的小巷飘来。
一百六十年前,陈麻婆的灶台上,多了这一锅红——
豆瓣酱的醇厚,花椒的麻香,碎肉的酥脆,还有那一块块,颤巍巍的嫩豆腐。
没有华丽的辞藻,只有一个家的味道。
一勺米饭,一口豆腐——
这就是中国,最温柔的那一面。
Can you smell it? That hint of numbing, that whisper of spice, drifting from the alleys of Chengdu.
A hundred and sixty years ago, on Chen Mapo's stove, a new pot of red appeared —
the richness of doubanjiang, the numbing aroma of Sichuan peppercorn, the crispness of minced meat, and those trembling cubes of silken tofu.
No fancy words — just the taste of home.
A spoonful of rice, one bite of tofu —
this is China's gentlest side.






The reel is the second pass. The first script draft was declarative and encyclopedic ("陈麻婆于1862年所创,凝聚六个字…"); the review-task replaced it with a sensory, conversational opener ("你闻到了吗?") and an intimate close ("中国最温柔的那一面"). The first i2v prompt draft pushed for visible cooking motion; the review-task softened them to gentle steam drift + slow chopsticks lift + the breath of a closed-eyes tasting smile. Both decisions are baked into the playbook so the next dish — a Hanoi pho, a Roman carbonara, a Bangkok som tum — gets the same warmth-first review-gate before any expensive render fires.
create_media action=generate, model=gpt-image, aspect_ratio=9:16 → 6/6 done in ~112–135s each, ZERO moderation flags
Same warm key-light upper-left at 30° in every prompt. NO readable text guard.
create_media action=animate, model=kling-v3-i2v, duration=5, gentle motion only → 6/6 done in ~55–62s each
1 × create_media action=generate, model=gemini-tts, prompt=Mandarin script + voice direction → done in 21s, 31.92s narration
1 × create_media action=music, model=minimax-music, prompt=warm guzheng + hand-percussion → done in 55s
mapo-tofu-short.mp4 (21MB, 30.0s)
This reel was built from the local-food-short playbook. The same template works for any famous local dish in any language gemini-tts supports — Pho Bo voiced in northern Vietnamese, Carbonara voiced in Roman Italian, Bibimbap voiced in Seoul Korean, Pad Thai voiced in Bangkok Thai, Tagine voiced in Moroccan Arabic. Paste the playbook into Claude with your YAML config (language + dish_name + key_flavors + voice_brief), and the same 3-wave + review-gate pipeline runs again — different dish, different tongue, same warmth.