One master spot. Every market. Same spokesperson, re-voiced and lip-synced into each language — transcribed, translated, scored, delivered.
Master
1 spot
Markets
EN · ES · JA
Lip-synced
3 / 3
Audio
licensed
Wall time
~18 min
Cost
~$4
Same spokesperson, three languages, lip-synced — side by side.
Each market, in its own voice
The same talent, re-voiced per language and lip-synced to match. Play each — the mouth moves to the new language, not the original.
EN — MASTER
"Discover AURELIA. Radiance, naturally."
ES — ESPAÑOL
"Descubre AURELIA. Resplandor, naturalmente."
JA — 日本語
"オーレリア。自然に輝く肌へ。毎日に、美しさを。"
The pipeline
nemotron-asr transcribed the master in 1.9s → "Discover AURELIA, radiance naturally, skincare that works as beautifully as it feels." → the base for translation.
spokesperson (gpt-image)→gemini-tts (EN · warm ♀ voice)→talking-head → EN master→nemotron-asr → transcript→translate + gemini-tts (ES/JA · same voice)→talking-head ×2 → lip-synced→sonilo bed + ffmpeg compare
Honest verdict
✓ What held
3/3 languages lip-synced to the same spokesperson — the mouth follows each translated VO.
One warm, on-brand female voice (gemini-tts) matched to the spokesperson and held consistent across all three markets — set the delivery with a plain-language tone direction.
nemotron-asr transcribed accurately in ~2s (40-language capable).
Multilingual TTS + licensed Sonilo bed = a publish-ready, rights-clear set.
⚠ What to watch
talking-head runs ~5 min/clip — rendered via the direct-SDK path with a long timeout (the slow-cap fix); parallelize per market.
Lip-sync fidelity varies by language phonemes — gate on the first market; subtitles are the safety net.
TTS pronounces invented brand names loosely ("AURLEA") — spell critical names phonetically in the script.