Audience question → animated visual answer
Turn your approved answer and visual beats into a narrated vertical explainer with stock footage, timed diagrams, and synchronized captions.
Run it
AI Agent
Terminal
New here? Quick start
brew install pipe2-ai/tap/pipe2 go install github.com/pipe2-ai/pipe2-cli/cmd/pipe2@latest echo "$PAT" | pipe2 auth login --token - Change the result
| Input | Type | default | Description |
|---|---|---|---|
--answer | string | defaultClouds contain tiny water droplets. Those droplets fall so slowly that rising air can keep them suspended. When droplets grow large enough, they can fall through that rising air as rain. A cloud isn't weightless; its droplets are just small. | Your fact-checked answer. Keep question plus answer between 20 and 60 words. |
--footage_query | string | defaultvertical clouds sky | Concrete portrait stock-footage subject matching your answer. |
--question | string | defaultWhy don't clouds fall? | Audience question, read aloud and shown on screen. Maximum 80 characters. |
--visual_beats | string | defaultUse these fractions of total duration: 0–12%: show the question immediately on two balanced lines. 12–24%: TINY WATER DROPS heading and six cyan dots in a loose cluster, gently drifting down. 24–51%: RISING AIR / HOLDS THEM UP. Yellow arrows repeatedly rise on either side of the tiny dots; the dots stay nearly suspended. 51–80%: BIGGER DROPS / FALL AS RAIN. Three waves of three larger blue drops fall; remove the upward arrows. 80–100%: NOT WEIGHTLESS. / JUST TINY. Compare small suspended dots with an upward arrow on the left and two larger falling drops on the right. These are illustrative diagrams, not measured sizes or a claim that clouds contain no ice. | Approved labels and simple diagrams, with timing as percentages of narration duration. Supply matching beats when changing the question or answer. |
-
Measures the narration and creates the transcript used for caption timing.
Why don't clouds fall? Clouds contain tiny water droplets. Those droplets fall so slowly that rising air can keep them suspended. When droplets grow large enough, they can fall through that rising air as rain. A cloud isn't weightless. Its droplets are just small -
Finds stock footage; selects two distinct clips long enough to cover the narration and records their sources.
-
Assembles the two clips in vertical framing with narration, using the measured audio length.
-
Animates the approved explanatory labels, arrows, and shapes over the footage.
-
Adds bold yellow phrase captions below the diagrams, synchronized to the answer.
Why it works and what to change
Why it works
The narration is the edit’s clock. Measuring the actual spoken answer avoids guessing its duration from a word count, while the same transcript supplies caption timing. Two real stock clips provide visual context without asking a video model to invent evidence for your explanation.
The opening question gives way to an illustrated answer: short labels, moving arrows and simple shapes follow your approved visual beats. In the cloud example, tiny droplets are held up by rising air before larger drops fall as rain. These diagrams illustrate the mechanism, not measured sizes or speeds. Bold yellow captions sit below them. These are phrase-timed captions, not word-level animation.
You own the answer. The recipe does not research facts, scrape comments or pretend to be the person who asked the question. Use a science explanation, a destination FAQ or a product-care answer with footage that genuinely matches the subject. MiniMax Speech 2.8 Turbo reads the supplied English text with its expressive narrator voice. Text Card plans and renders the graphics over the stock footage. No image or video generation model is involved.
The knobs that matter
questionsupplies the spoken opening and on-screen title. Ask one specific question in at most 80 characters; a short line reads best on a phone.answersupplies the approved explanation, including a complete ending. Keep the combined question and answer between 20 and 60 words.footage_querynames something visible, such as clouds or a train station, rather than an abstract feeling or the whole question.visual_beatsdescribes the labels, simple diagrams and timing as percentages of the spoken answer. Change it with the question or answer: the recipe stops before spending credits if a new topic still has the default cloud diagrams.
The manifest lists the complete input contract.
Notes
Only native 9:16 HD clips qualify: a portrait export setting alone can leave landscape footage letterboxed. Include “vertical” in your footage query. Stock search order can change. Entries explicitly labeled AI-generated are excluded, but metadata alone cannot verify footage provenance. Review the two selected clips and retain the source links and attribution printed during the run. Generic B-roll illustrates an explanation; it does not prove a specific historical event or product claim. Check media usage rights before posting, and disclose synthetic narration where required. This recipe exports video; it does not post or create native reply stickers on a social platform.
The default answer follows the USGS explanation of cloud droplets and updrafts. Review the final speech, diagram timing, captions and crop before sharing any variation. The same inputs preserve the workflow, not identical stock-search results or an identical AI-planned layout.