← Articles
How-to

How to Add Subtitles to a Video with an SRT File

Create or upload an SRT, choose a caption style and position, and burn subtitles into a video as a ready-to-share MP4.

By Pipe2.ai · Updated July 10, 2026

The reliable way to add subtitles to a video is to create an SRT transcript, correct it, then render it with the video. In Pipe2.ai, run Transcription first if you need an SRT, then attach the video and SRT to the Captions pipeline, choose a style and position, and export the captioned MP4.

Add subtitles with the two-step SRT workflow

Pipe2.ai separates transcription from caption rendering. That gives you a reviewable subtitle file between speech recognition and the final video.

  1. Create or upload an SRT. Run the source through Transcription, or use an SRT from another tool.
  2. Review the transcript. Fix names, technical terms, punctuation, and timing before text becomes part of the picture.
  3. Attach the video and SRT. Both are required inputs for the Captions pipeline.
  4. Choose a style and position. Preview the combination against the actual framing.
  5. Render and inspect the MP4. Watch the whole result rather than checking only a still frame.

Transcription and caption rendering are separate runs with separate pricing. If you already have a good SRT, skip transcription and go directly to the caption burn.

Choose burned-in subtitles or a separate caption track

The Captions pipeline creates burned-in subtitles: the words are rendered into the video pixels. They remain visible in every player and retain the chosen font, color, and placement. Viewers cannot turn them off.

A separate SRT track behaves differently. A compatible player can let viewers show or hide it, and the text may be available to accessibility tools, but the platform controls its appearance. The two formats can be used together: publish the captioned MP4 for consistent presentation and retain the SRT for platforms that accept a separate accessibility track.

Choose intentionally. Burned-in subtitles are useful for social feeds where videos often begin muted; a separate track gives the viewer more control.

Edit the SRT before rendering

An SRT is a plain-text file containing numbered subtitle blocks, start and end timestamps, and the words shown during that interval. Small transcript errors become much more visible after they are baked into the picture.

Review at least these details:

  • Names, product terms, acronyms, and uncommon vocabulary
  • Lines that appear too early, too late, or remain after the speaker stops
  • Blocks that contain too much text to read comfortably
  • Punctuation that changes the intended meaning
  • Speaker labels that you do not want displayed on screen

Keep the corrected SRT as the reusable source of truth. You can render it more than once with different visual presets without re-running speech-to-text.

Pick a caption style and position

The current Captions pipeline provides six style presets: TikTok Bold Yellow, Minimal White, Subtle Drop Shadow, Karaoke Gradient, Big Serif, and Serif Editorial. It also offers Auto, Bottom, Top, and Middle positions.

Use the visual hierarchy of the footage to decide:

  • Bottom is a conventional starting point when the lower frame is clear.
  • Top can protect a speaker, product, or on-screen graphics in the lower half.
  • Middle is available but often competes with the subject, so preview it carefully.
  • Auto uses an upstream subject-position hint when one is available; without that hint, it resolves to Bottom.

If a landscape clip will be converted to vertical, run Video Reframe before captions. The final frame composition should determine where the text sits.

Review readability and the final production order

Watch the output at the size and on the type of device your audience will use. Check contrast over changing backgrounds, verify that every line stays inside the visible frame, and make sure captions do not collide with app controls.

For a finished social clip, a sensible order is reframe, captions, then watermark. Captioning after reframing fits the text to the final canvas; watermarking last protects the logo from later crops. See how to add a watermark to a video for the final branding pass.

The caption render is performed server-side and returns a new MP4. Preserve the original video and corrected SRT alongside it so you can revise styling or timing without rebuilding the transcript from scratch.

Frequently asked questions

What is the simplest way to add subtitles to a video?

Create an SRT transcript, attach it with the source video in the Captions pipeline, choose a style and position, and render the result. If you do not have an SRT, run the Transcription pipeline first and use the SRT asset it produces.

Do I need an SRT file to use the Captions pipeline?

Yes. The Captions pipeline requires both a video asset and an SRT transcript asset. The SRT can come from Pipe2.ai's Transcription pipeline or from another transcription or subtitle editor.

Are Pipe2.ai subtitles burned into the video?

Yes. The captions are rendered into the video pixels and the output is an MP4, so the styling travels with the file. Keep the SRT too if you also want a selectable caption track on platforms that support one.

Can I edit or reuse a transcript before adding subtitles?

Yes. Correct names, jargon, wording, and timing in the SRT before rendering. You can attach the same edited SRT to multiple Captions runs to compare styles without paying to transcribe the source again.

Try these pipelines

Related articles