Four ingredients on a single reference sheet — a coral sneaker, a runner, a rooftop at dusk, a duffel — composited into one cinematic ad scene that keeps each one on-model. Not generic text-to-video: the sheet says what appears, the prompt says what happens. Real assets, end-to-end on Livepeer's ltx-q-ingredient cap.
Left: the reference sheet (built with gpt-image) — four labeled cells, neutral backgrounds, consistent scale. Right: the generated scene from ltx-q-ingredient with the action prompt. The sheet locks identity; the prompt directs the shot.
Prompt: "the runner in the teal windbreaker laces up the coral sneaker on the city rooftop at golden-hour dusk, the gym duffel beside them; slow push-in as the sneaker catches the last warm light." The sneaker, the jacket, the rooftop, and the bag all carry through from the sheet.
pillow-grid.)create_media(action:"animate", source_url:<sheet>, model_override:"ltx-q-ingredient", prompt:<action + style>). The sheet is the source; the prompt is the direction.sonilo-v2m (or music), ffmpeg-concat multiple scenes, ffmpeg-export to aspect, ffmpeg-overlay a logo.Verdict: ships. One reference sheet → a coherent, on-model ad scene in ~2 minutes, where the sneaker, character, environment, and prop all survive into the video. The identity-preservation that generic text-to-video can't give you.
Pair it with the reference-sheet-to-ad playbook and the ltx-ingredient skill for the full recipe + reference-sheet craft.