Captions

From 0.2 credits

Burn styled captions into any video. The speech is transcribed for you, or attach your own SRT. Pick a preset and ship a ready-to-post MP4.

Complete to generate

Captions API

Call this pipeline from your own code. One request dispatches a run; the model is an input field, not a separate endpoint.

Get an API token
inputs
4
required
1
Pricing
from 0.2 credits
Language
Language
bun add @pipe2-ai/sdk
Run a pipeline
import { createClient } from "@pipe2-ai/sdk";

const client = createClient(process.env.PIPE2_TOKEN!);

const { run_pipeline } = await client.RunPipeline({
  pipeline_slug: "captions",
  input: {
    source_asset_id: "a1b2c3d4-0000-4000-8000-000000000001",
  },
});

pip install pipe2
Run a pipeline
import asyncio
import os

from pipe2 import Pipe2Client

client = Pipe2Client(token=os.environ["PIPE2_TOKEN"])

run = asyncio.run(client.run_pipeline(
    pipeline_slug="captions",
    input={
        "source_asset_id": "a1b2c3d4-0000-4000-8000-000000000001",
    },
))

go get github.com/pipe2-ai/sdk-go
Run a pipeline
client := pipe2.NewClient(os.Getenv("PIPE2_TOKEN"))

input := json.RawMessage(`{
  "source_asset_id": "a1b2c3d4-0000-4000-8000-000000000001"
}`)

run, err := pipe2.RunPipeline(context.Background(), client, "captions", input)
if err != nil {
    log.Fatal(err)
}

brew install pipe2-ai/tap/pipe2
Run a pipeline
pipe2 pipelines run \
  --pipeline captions \
  --input-json '{"source_asset_id": "a1b2c3d4-0000-4000-8000-000000000001"}' \
  --wait

Full API reference CLI reference

Frequently Asked Questions

What file formats are supported?
Video: MP4, MOV, WebM, AVI, MKV, and more. Transcript: .srt (standard subtitle format with timestamps). Output is always MP4 for universal platform compatibility.
Do I need a transcript?
No. Leave it empty and the speech is transcribed automatically; the price then includes transcription. To use your own text, attach an SRT from the Transcription pipeline or any other tool.
Are the captions burned into the pixels or overlaid on top?
Burned into the pixels. The captions travel with the video regardless of where you upload it, no separate .srt to manage afterwards, no risk of the platform stripping or restyling them.
Can I edit the transcript before burning it in?
Yes. Run Transcription, correct the SRT in your library, then attach it here. You can also upload an edited .srt directly.
Can I reuse one transcript across multiple caption styles?
Yes. Attach the same SRT each time and pick a different preset, much cheaper than re-running transcription, and useful for A/B testing captions on short-form.
How long does the caption burn take?
Typically 30-60% of the source duration. A 1-minute clip takes 20-40 seconds; a 10-minute clip takes 3-6 minutes. Automatic transcription adds some time for the speech.

Captions: Burn Styled Captions Into Any Video

Attach a video, pick a style, and get back an MP4 with captions burned directly into the pixels. The speech is transcribed for you. No CapCut, no Premiere, no caption extension to install, just a ready-to-post file sized for TikTok, Reels, and Shorts.

Want to fix names or jargon before they are burned in? Run the Transcription pipeline, correct the SRT in your library, and attach it here. Captions then uses your text exactly as written.

How It Works

  1. Attach your video: upload from your device or pick one from your library.
  2. Attach an SRT (optional): leave it empty to transcribe the speech automatically, or attach a corrected or existing .srt.
  3. Pick a caption style: six production-ready presets tuned for different audiences.
  4. Download the result: a new MP4 with captions burned in, plus a poster frame for thumbnails.

Caption Style Presets

  • TikTok Bold Yellow: heavy weight, black stroke, TikTok-safe positioning. The default everyone copies.
  • Minimal White: clean sans-serif with a soft shadow. Works for talking-head podcasts and interviews.
  • Subtle Drop Shadow: low-contrast caption for cinematic edits where the visuals carry the scene.
  • Karaoke Gradient: active word highlighted in a purple→pink gradient as it's spoken. High retention on short-form.
  • Big Serif: editorial-feel serif for brand docs, case studies, and long-form educational clips.
  • Serif Editorial: restrained serif type on a translucent panel for podcasts, explainers, and thought-leadership clips.

Why Attach Your Own SRT

  • Edit the transcript before burning: fix names, jargon, or speaker labels in the SRT instead of living with whatever STT gave you.
  • Reuse one transcript across many styles: burn the same SRT with six different presets for A/B testing without re-transcribing.
  • Bring your own SRT: if you already have a transcript from another tool, attach it and skip the transcription cost entirely.
  • TikTok-safe positioning: captions clear the bottom-third UI overlays on every major platform.

Who It's For

  • Short-form creators shipping daily Reels / Shorts / TikToks
  • Podcasters repurposing long-form episodes into shareable clips
  • Course creators and educators adding accessibility to tutorials
  • Agencies producing branded social content at scale
  • Anyone who has an SRT already and just wants it rendered cleanly into the video

Use it on its own for a ship-ready clip, or feed the output straight into the Video Reel pipeline for multi-clip compilations.

Explore more pipelines

Video Generator
from 9.5
Video Generator
Image Generator
0.2-6.7
Image Generator
Audio Generator
0.002-40.0
Audio Generator
Music Generator
2.5-22.6
Music Generator