Long video → multiple captioned clips, in one command

Slice any long video into N captioned shorter clips. Transcribe once, auto-pick the moments with AI, trim + caption each in parallel. Optionally reformat to vertical (9:16) for TikTok / Reels / Shorts, square (1:1) for Instagram, or portrait (4:5); the caption anchor auto-adjusts.

  • Video
  • 6 pipelines
  • ≥ 11.6 credits

AI Agent

Terminal

New here? Quick start
01
Homebrew brew install pipe2-ai/tap/pipe2
Go go install github.com/pipe2-ai/pipe2-cli/cmd/pipe2@latest
02
Get an API token ↗ echo "$PAT" | pipe2 auth login --token -

Change the result

--inputrequired
asset_urldefaultSource video: a YouTube or social URL, a direct media URL, a local file, or an existing pipe2 asset. Remote and local sources are uploaded automatically; use --asset <id> or --no-fetch when the asset is already available.
--reformat
enumdefaultOptional output aspect ratio. Leave empty to preserve the source's native aspect (default: fastest, cheapest). Set to 9:16 for TikTok/Reels/Shorts, 1:1 or 4:5 for Instagram, 16:9 for horizontal YouTube cards from a vertical source.
· 9:16 · 1:1 +2
9:161:14:516:9
--highlights-count
intdefault5How many moments to pick when highlights runs (auto mode). Ignored if --clips is set.
--highlights-style
stringdefaultNatural-language steer for the highlights picker: e.g. "the funniest moments", "the strongest arguments". Empty uses the picker's default.
--clips
stringdefaultOptional path to a JSON file overriding the auto-picker, shaped [{"context": "...", "start_sec": 42.5, "end_sec": 78.0}, ...]. When set, the highlights step is skipped. Leave empty to let the highlights pipeline pick automatically.
--corrections
stringdefaultComma-separated word-boundary substitutions applied to the transcript, in the form "from=to,from=to". Use it when the recognizer mis-hears the same word the same way every time (a name, an acronym, a domain term). To preserve a phrase that contains a substring you also want to rewrite, declare the longer phrase first as a no-op ("phrase=phrase,word=replacement"): longer matches win, so the phrase is shielded before the bare-word rule fires. Case-sensitive.
--lang
stringdefaultenISO 639-1 transcription language, or 'auto'.
--no-watermark
booldefaultfalseShip the output unbranded. Skips the watermark step entirely.
--parallel
intdefault4Max number of clips to process in parallel.
--position
enumdefaultautoVertical caption position. "auto" (default) keeps text away from the main subject when possible and otherwise places it at the bottom. Set a position explicitly to override.
auto · top · middle +1
autotopmiddlebottom
--preset
enumdefaultserif-editorialCaption styling preset for every clip.
tiktok-bold-yellow · minimal-white · subtle-drop +3
tiktok-bold-yellowminimal-whitesubtle-dropkaraoke-gradientbig-serifserif-editorial
--watermark-scale
intdefault20Watermark width as a percentage of the video width.
--watermark-url
asset_urldefaultOverride the default Pipe2 logo with your own image (URL or local path). Empty means use the bundled Pipe2 watermark.
--watermark-variant
enumdefaultlightWhich bundled Pipe2 logo to use when --watermark-url is empty. "light" for dark videos, "dark" for bright ones; coloured variants match the clip palette.
amber · aqua · crimson +4
amberaquacrimsondarkemeraldindigolight
How it works 6 pipelines
  1. Transcribes the full source once and reuses that transcript for every selected clip.

    Input source: OpenClaw Creator, Why 80% of Apps Will Disappear
    Text
    00:00:00,180  Today, I'm sitting down with Peter Steinberger, the creator of OpenClaw, the open source personal AI agent that has completely taken over the internet.
    00:00:09,060  The GitHub repo exploded to over 160,000 stars practically overnight.
    00:00:14,130  The community has built countless projects, like Malt Book, where bots talk among themselves.
    00:00:19,560  And now, the bots are even renting humans to do tasks in the real world.
    00:00:24,220  In our conversation, we discuss his aha moment, his contrarian development philosophies, and what this means for builders in 2026.
    00:00:32,740  Let's dive in.
    00:00:38,980  So good to see you, man.
    00:00:39,960  Hey, what's up?
    00:00:40,760  Um, so you've made something people want.
    …
  2. Reads the transcript and picks N editorial moments: the auto-pick path. Skipped when --clips supplies a manual JSON list.

    JSON
    [{"context":"Peter explains OpenClaw's core differentiator: local execution gives it access to everything the user can do, unlike cloud-based AI.","desired_seconds":32,"start_sec":95.24000000000001,"end_sec":127.64},{"context":"The vivid moment Peter realized OpenClaw's creative problem-solving: it autonomously transcribed a voice message using ffmpeg and OpenAI's API without being explicitly programmed to do so.","desired_seconds":106,"start_sec":515.182,"end_sec":621.252},{"context":"Peter's contrarian prediction: 80% of apps disappear because personal AI agents manage data and tasks more naturally than purpose-built applications.","desired_seconds":55,"start_sec":641.792,"end_sec":696.312},{"context":"Peter's philosophy on building: minimize friction by using Unix tools and CLIs instead of inventing new abstractions, letting the model handle creative problem-solving.","desired_seconds":54,"start_sec":1178.264,"end_sec":1232.584},{"context":"Peter articulates why swarm intelligence mirrors human society: individuals alone can't build iPhones or go to space, but groups specializing together achieve anything.","desired_seconds":46,"start_sec":258,"end_sec":303.5}]
  3. Cuts each selected window on sentence boundaries and returns a matching clip transcript for captions.

    Video
  4. video-reframe only if --reformat

    When --reformat is set, reframes each clip to the requested aspect ratio, keeps the active speaker in view, and changes framing cleanly at shot boundaries. Otherwise the source aspect ratio is preserved.

    Video
  5. Adds the matching transcript to each clip in the chosen caption style and position.

    Video
  6. Overlays the Pipe2 logo (or your own --watermark-url) at the top-left corner so the clip carries attribution across reposts. Pass --no-watermark to skip.

    Video