Long video → multiple captioned clips, in one command
Slice any long video into N captioned shorter clips. Transcribe once, auto-pick the moments with AI, trim + caption each in parallel. Optionally reformat to vertical (9:16) for TikTok / Reels / Shorts, square (1:1) for Instagram, or portrait (4:5); the caption anchor auto-adjusts.
Run it
AI Agent
Terminal
New here? Quick start
brew install pipe2-ai/tap/pipe2 go install github.com/pipe2-ai/pipe2-cli/cmd/pipe2@latest echo "$PAT" | pipe2 auth login --token - 5 clips from one episode
Change the result
| Input | Type | default | Description |
|---|---|---|---|
--input | asset_url | default— | Source video: a YouTube or social URL, a direct media URL, a local file, or an existing pipe2 asset. Remote and local sources are uploaded automatically; use --asset <id> or --no-fetch when the asset is already available. |
--reformat | enum | default | Optional output aspect ratio. Leave empty to preserve the source's native aspect (default: fastest, cheapest). Set to 9:16 for TikTok/Reels/Shorts, 1:1 or 4:5 for Instagram, 16:9 for horizontal YouTube cards from a vertical source.· 9:16 · 1:1 +29:161:14:516:9 |
--highlights-count | int | default5 | How many moments to pick when highlights runs (auto mode). Ignored if --clips is set. |
--highlights-style | string | default | Natural-language steer for the highlights picker: e.g. "the funniest moments", "the strongest arguments". Empty uses the picker's default. |
--clips | string | default | Optional path to a JSON file overriding the auto-picker, shaped [{"context": "...", "start_sec": 42.5, "end_sec": 78.0}, ...]. When set, the highlights step is skipped. Leave empty to let the highlights pipeline pick automatically. |
--corrections | string | default | Comma-separated word-boundary substitutions applied to the transcript, in the form "from=to,from=to". Use it when the recognizer mis-hears the same word the same way every time (a name, an acronym, a domain term). To preserve a phrase that contains a substring you also want to rewrite, declare the longer phrase first as a no-op ("phrase=phrase,word=replacement"): longer matches win, so the phrase is shielded before the bare-word rule fires. Case-sensitive. |
--lang | string | defaulten | ISO 639-1 transcription language, or 'auto'. |
--no-watermark | bool | defaultfalse | Ship the output unbranded. Skips the watermark step entirely. |
--parallel | int | default4 | Max number of clips to process in parallel. |
--position | enum | defaultauto | Vertical caption position. "auto" (default) keeps text away from the main subject when possible and otherwise places it at the bottom. Set a position explicitly to override.auto · top · middle +1autotopmiddlebottom |
--preset | enum | defaultserif-editorial | Caption styling preset for every clip.tiktok-bold-yellow · minimal-white · subtle-drop +3tiktok-bold-yellowminimal-whitesubtle-dropkaraoke-gradientbig-serifserif-editorial |
--watermark-scale | int | default20 | Watermark width as a percentage of the video width. |
--watermark-url | asset_url | default | Override the default Pipe2 logo with your own image (URL or local path). Empty means use the bundled Pipe2 watermark. |
--watermark-variant | enum | defaultlight | Which bundled Pipe2 logo to use when --watermark-url is empty. "light" for dark videos, "dark" for bright ones; coloured variants match the clip palette.amber · aqua · crimson +4amberaquacrimsondarkemeraldindigolight |
-
Transcribes the full source once and reuses that transcript for every selected clip.
Input source: OpenClaw Creator, Why 80% of Apps Will Disappear00:00:00,180 Today, I'm sitting down with Peter Steinberger, the creator of OpenClaw, the open source personal AI agent that has completely taken over the internet. 00:00:09,060 The GitHub repo exploded to over 160,000 stars practically overnight. 00:00:14,130 The community has built countless projects, like Malt Book, where bots talk among themselves. 00:00:19,560 And now, the bots are even renting humans to do tasks in the real world. 00:00:24,220 In our conversation, we discuss his aha moment, his contrarian development philosophies, and what this means for builders in 2026. 00:00:32,740 Let's dive in. 00:00:38,980 So good to see you, man. 00:00:39,960 Hey, what's up? 00:00:40,760 Um, so you've made something people want. … -
Reads the transcript and picks N editorial moments: the auto-pick path. Skipped when --clips supplies a manual JSON list.
[{"context":"Peter explains OpenClaw's core differentiator: local execution gives it access to everything the user can do, unlike cloud-based AI.","desired_seconds":32,"start_sec":95.24000000000001,"end_sec":127.64},{"context":"The vivid moment Peter realized OpenClaw's creative problem-solving: it autonomously transcribed a voice message using ffmpeg and OpenAI's API without being explicitly programmed to do so.","desired_seconds":106,"start_sec":515.182,"end_sec":621.252},{"context":"Peter's contrarian prediction: 80% of apps disappear because personal AI agents manage data and tasks more naturally than purpose-built applications.","desired_seconds":55,"start_sec":641.792,"end_sec":696.312},{"context":"Peter's philosophy on building: minimize friction by using Unix tools and CLIs instead of inventing new abstractions, letting the model handle creative problem-solving.","desired_seconds":54,"start_sec":1178.264,"end_sec":1232.584},{"context":"Peter articulates why swarm intelligence mirrors human society: individuals alone can't build iPhones or go to space, but groups specializing together achieve anything.","desired_seconds":46,"start_sec":258,"end_sec":303.5}] -
Cuts each selected window on sentence boundaries and returns a matching clip transcript for captions.
-
When --reformat is set, reframes each clip to the requested aspect ratio, keeps the active speaker in view, and changes framing cleanly at shot boundaries. Otherwise the source aspect ratio is preserved.
-
Adds the matching transcript to each clip in the chosen caption style and position.
-
Overlays the Pipe2 logo (or your own --watermark-url) at the top-left corner so the clip carries attribution across reposts. Pass --no-watermark to skip.