A failing starship. A 4% chance. Three voices, the klaxons, the score — every second of the sound you're about to hear came from a single text prompt, generated in one pass.
▶ The motion-comic — 6 krea-2-os panels cut to the Seed Audio track. Unmute for the full mix.
seed-audio (bytedance/seed-audio-1.0) on sdk.daydream.monster. No per-line TTS. No separate music. No foley. No mixing pass.This is the entire input. Speaker, emotional direction in parentheses, the lines, then the ambience and the score — one text block. Seed Audio cast three distinct voices, performed the emotional arc, placed the sound effects, scored it, and mixed it.
No voice actors, no reference recordings — each voice was described in words and Seed Audio built it. (It can also clone a voice from up to 3 reference clips, or derive one from a character image.)









| The old way | Seed Audio 1.0 |
|---|---|
| Cast + direct 3 voice actors, book a booth | 3 voices from one prompt, each with tone direction |
| Hire a foley artist for klaxons / sparks / room tone | SFX described in the same prompt, placed in the mix |
| Commission a composer for the underscore | Original score that builds + resolves, same pass |
| Pay a mixing engineer to balance it | Already mixed — one master track out |
| Days of work + a studio budget | One pass · ~one minute · pennies |
seed-audio → a single mixed audio_url. The whole point.krea-2-os portraits + beat panels (~3s each).ffmpeg → a watchable reel. Optional: lip-sync a close-up with sync-lipsync-v3.Why it matters. Audio was the last part of a scene you couldn't fake — voices, performance, sound design and score each needed a person. Seed Audio collapses all four into one prompt, one pass, already mixed. A game studio scripts a cutscene; a novelist hears their chapter; a teacher builds a two-character dialogue — in the time it takes to write it.
Make your own. Open the One Prompt, a Whole Scene playbook, fill in a SCENE, and run it.