---
title: "Ever Wondered… — a surprising science question → reveal → the science → mind-blown closer"
tier: hero
format: short-form-video
theme: science | curiosity | edutainment | wonder
persona: science creator, STEM educator, curiosity-channel host, museum social
duration: "1 topic → 45–90s vertical short in ~3–6 min"
budget_usd: "$0.40–$3.00 per short (5–7 scene images/clips + TTS + music + finishing)"
caps: ["gpt-image", "seedream-5-lite", "flux-dev", "cosmos-3-image", "ltx-q-i2v", "seedance-i2v", "inworld-tts", "music", "ffmpeg-burn-subtitles", "hyperframes-caption", "ffmpeg-concat", "ffmpeg-mux", "ffmpeg-export"]
skills: ["short-form-video"]
showcases: []
status: "playbook (2026-06-08) — ported from the Pixelle-Video format study; runs on existing Storyboard caps. Showcase render pending."
reliability: 3.7 # 5 − image gen .3 − i2v .6 − tts .2 − music .2; caption/subtitle/finishing chain deterministic; captions + fact overlays OPTIONAL on a first cut; no proven E2E showcase yet
---

# Ever Wondered… — a science curiosity hit in 90 seconds

Point this at a surprising science question — "why is the sky blue at noon but red at sunset?", "what's inside a black hole?", "how do octopuses taste with their arms?" — and it builds an **"ever wondered" curiosity short**: a question, a reveal, the actual science, and a mind-blown closer. Striking macro/space/micro imagery, energetic delivery. Runs entirely on Storyboard's existing Livepeer caps.

## What you'll get

> **A finished reel = vertical 9:16 + voiceover + music + ~45–90s.** Pass `aspect_ratio: "9:16"` to `generate_project` (it now threads to every scene), set the i2v `duration` to 8–15s **on EVERY scene** and animate **all** of them, add `inworld-tts` narration (not the chatterbox default voice) + `music`, and **mux the bed with `audio_fill:"loop"`** so the score covers the whole reel. A silent landscape montage — or one whose music cuts out partway — is a draft, not a deliverable. See the `short-form-video` skill's *definition of done*.

A 45–90s **vertical (9:16)** curiosity short: 5–7 striking scenes (macro, space, microscopic, or natural-world imagery), a curious energetic narrator, a wonder/ambient-build score, captions, and a finished MP4.

## The cap chain

```
topic (a "did you ever wonder…" question)
  → agent writes the curiosity script + 5–7 scene beats (question → reveal → science → closer)
  → generate_project           # striking macro / space / micro imagery per beat
       image: gpt-image | seedream-5-lite | flux-dev | cosmos-3-image (cosmic/world scenes)
  → (optional) motion          # ltx-q-i2v or seedance-i2v, resolution:"auto", for hypnotic motion (drifting nebula, pulsing cell)
  → inworld-tts                # curious, energetic narrator (NOT the chatterbox default voice)
  → music  lyrics_prompt:"[Instrumental]"   # "ambient wonder, building, awe, cosmic" — a 15s bed is fine; you loop it
  → hyperframes-caption        # punch-in number/fact overlays
  → ffmpeg-burn-subtitles      # spoken-word captions in post
  → ffmpeg-concat                                      # stitch the per-scene clips
  → ffmpeg-mux  audio_fill:"loop"                      # lay the looped, ducked score under the FULL reel
  → ffmpeg-export                                      # 9:16, 1080×1920
  # audio_fill:"loop" repeats a short bed over the whole reel — else the score cuts out, last third silent.
```

## How the agent should run this (interaction contract)

1. **CONFIRM (one message, ≤1 question):** restate in 1 line ("60s 9:16 curiosity short on <question>, 6 beats, inworld VO + wonder score, ~$1–3, ~10–15 min") + real choices: 5 vs 7 beats · flux-dev/seedream (fast) vs gpt-image (cleanest) · stills-draft vs full animated reel. If the question is clear, state assumptions and start.
2. **PREVIEW CHECKPOINT:** show the SCRIPT (question → reveal → science → closer, $0) and the still keyframes (~$0.30) BEFORE any i2v fires. Fact-check the closer here — that's the line people screenshot; a wrong "fun fact" is this playbook's worst failure.
3. **NARRATE:** per wave: "animating beats 1–4 (~2–4 min each), job …"; poll `get_create_media` every 10–15s, ≤4–5 i2v in flight, post a 1-liner at least every 2 min — never silent >2 min.
4. **FAIL GRACEFULLY:** gpt-image 422/timeout → seedream-5-lite or flux-dev (cosmos-3-image for cosmic scenes); seedance hang/"fetch failed" → retry once → `ltx-q-i2v` (`num_frames: 241`, resolution `"auto"`) → Ken-Burns still (drifting zoom suits cosmic/micro imagery anyway). ≤2 retries per scene; a reel with one Ken-Burns beat beats no reel.
5. **DELIVER:** final MP4 + one honest line ("imagery is artist's impression, not real data; closer fact verified against <source>") + ONE next step ("add the fact overlay + captions, or run the next question").

## Script structure (beats)

1. **Question hook (0–4s)** — the wonder, posed directly. "Ever wonder why you can't tickle yourself?"
2. **Reveal (1 beat)** — a teaser of the answer, just enough to hold them. "Your brain saw it coming."
3. **The science (2–4 beats)** — the actual mechanism, escalating from intuitive to surprising, one scene each.
4. **Mind-blown closer (last beat)** — the biggest, most shareable fact last. "…which means your brain is constantly predicting your own future."

## Pacing & aspect

- **Energetic but with breath on the reveal** — 7–10s per scene, a beat of silence before the closer. 9:16, 1080×1920.
- Lean on awe: drifting cosmic or pulsing micro motion (`ltx-q-i2v` / `seedance-i2v`) sells wonder.
- Score builds toward the closer; punch in a fact overlay (`hyperframes-caption`) on the mind-blown beat.

## Watch-outs

- **Render in WAVES — never fire every clip at once (tested the hard way, 2026-06-08).** Animating a whole reel by firing all its scene i2v jobs simultaneously — ×N reels — overwhelmed the SDK `/inference` worker (BYOC video caps have capacity ~2 each behind a single dispatch worker); a ~30-clip burst crashed it and most renders failed. Cap in-flight **video** renders at **≤4–5**, animate **one reel at a time**, and poll each batch to *done* before firing the next. Images are cheaper (cap ~4) but still wave large sets.

- **Scientific accuracy.** "Mind-blowing" + wrong = misinformation. Verify the closer fact especially — that's the one people screenshot. The model will happily generate a confident, false "fun fact."
- **Imagery is illustrative, not real data.** A generated "black hole" or "cell" is an artist's impression, not a photograph or scan. Don't imply it's real footage; say "illustration" if there's any ambiguity.
- **No text in the imagery.** Don't bake the question or the fact into the generated image — AI text rendering garbles it. Use `hyperframes-caption` / `ffmpeg-burn-subtitles` in post.
- **The multi-link quality ceiling.** A weak closer wastes a strong build; a flat narrator kills the wonder. Review the script's payoff and listen to the TTS once before stitching.

## Cross-surface (MCP / CLI / Webapp)

- **MCP:** `generate_project` (striking imagery) → `create_media` (motion, TTS, music) → `hyperframes-caption` fact overlay → `director_export`.
- **CLI:** `livepeer story "ever wondered: <question>, 6 curiosity beats, striking imagery"` → motion + TTS + music via `livepeer media create` → `livepeer media export --aspect 9:16 --subtitles captions.srt`.
- **Webapp chat:** `make an 'ever wondered' science short about <question>` — the agent drafts the question→reveal→science→closer arc and offers finishing as follow-ups.

## Execution mode (autopilot / director / auto)

This is a multi-scene project, so pass a `mode` to `generate_project` to control the **pre-finish gate** (after every scene renders, before delivery): `autopilot` (certain brief → ship in one call, no pause), `director` (always pause for a human pick/steer before delivery), or `auto` (recommended — ship UNLESS the auto-critic is unsure: score < `confidence_threshold` (default 0.8), verdict `iterate`, or a flagged shot). Setting `mode` runs the project async (returns a `job_id`); resume a director/auto pause with `review_checkpoint` + `resume_from_checkpoint` (decision `continue` to deliver / `abort` to discard). Omit `mode` for legacy straight-through behavior.
