Audio Generator

From 0.002 credits

Turn text into natural speech with dozens of voices, 70+ languages, and custom style directions for tone, accent, and pacing.

Create your first voiceover

Write what you want to say, or listen to an example.

Gemini 3.8 Flash TTS
Gemini 3.8 Flash-Lite TTS
Eleven v4
Complete to generate

Enter the exact words to speak; Gemini reads them as written. Add sounds or pauses where they happen in angle brackets, such as <laugh>, <sigh>, or <short pause>, and capitalize a word to stress it.

Examples

Meet Pyra: voiceover

Gemini 3.8 Flash TTS. Pyra's own story, read in a warm storyteller voice: a voiceover you can drop straight into any video.

How to use Audio Generator

A compact decision guide for placing this pipeline in a workflow.

Best for

  • Narration and voiceover in 70+ languages
  • Text-to-speech with natural language style control (e.g. 'speak warmly', 'whisper', 'excited tone')
  • Multi-voice content with dozens of expressive voices
  • Multilingual scripts, the language is auto-detected from the text

Tips

  • Write narration in short, clear sentences for better pacing
  • Use punctuation to control pauses: periods for full stops, commas for brief pauses, ellipsis for long pauses
  • Leave the voice on Auto and the AI picks one that matches the style instructions, or pin a specific voice if you want a recurring character
  • Use natural language style hints: 'speak slowly and dramatically', 'cheerful and upbeat', 'calm and reassuring'

Audio Generator API

Call this pipeline from your own code. One request dispatches a run; the model is an input field, not a separate endpoint.

Get an API token
inputs
5
required
1
Pricing
0.002 to 40.0 credits
Language
Language
bun add @pipe2-ai/sdk
Run a pipeline
import { createClient } from "@pipe2-ai/sdk";

const client = createClient(process.env.PIPE2_TOKEN!);

const { run_pipeline } = await client.RunPipeline({
  pipeline_slug: "audio-generator",
  input: {
    text: "Welcome back to the show. Today we're talking about something a little strange.",
    model: "gemini-3-8-flash-tts",
  },
});

pip install pipe2
Run a pipeline
import asyncio
import os

from pipe2 import Pipe2Client

client = Pipe2Client(token=os.environ["PIPE2_TOKEN"])

run = asyncio.run(client.run_pipeline(
    pipeline_slug="audio-generator",
    input={
        "text": "Welcome back to the show. Today we're talking about something a little strange.",
        "model": "gemini-3-8-flash-tts",
    },
))

go get github.com/pipe2-ai/sdk-go
Run a pipeline
client := pipe2.NewClient(os.Getenv("PIPE2_TOKEN"))

input := json.RawMessage(`{
  "text": "Welcome back to the show. Today we're talking about something a little strange.",
  "model": "gemini-3-8-flash-tts"
}`)

run, err := pipe2.RunPipeline(context.Background(), client, "audio-generator", input)
if err != nil {
    log.Fatal(err)
}

brew install pipe2-ai/tap/pipe2
Run a pipeline
pipe2 pipelines run \
  --pipeline audio-generator \
  --input-json '{"text": "Welcome back to the show. Today we're talking about something a little strange.", "model": "gemini-3-8-flash-tts"}' \
  --wait

Full API reference CLI reference

Frequently Asked Questions

What voices are available?
Dozens of expressive voices spanning bright, deep, warm, firm, and dramatic characters. Each responds to plain-language style instructions for tone and delivery.
What languages are supported?
70+ languages with automatic detection. From major languages like English, Chinese, Japanese, Spanish to regional languages including Cebuano, Konkani, and Luxembourgish.
How do instructions work?
Describe how you want the voice to sound in plain English. For example: 'A 40-year-old British journalist, speaking slowly with dramatic pauses on key points.' AI interprets your direction and crafts the vocal performance.
What audio format is the output?
MP3, ready to download or chain into pipelines like Video Reel as a narration track.
Which model should I pick?
Gemini 3.8 Flash-Lite TTS is the fastest and cheapest, good for drafts and short clips. Gemini 3.8 Flash TTS gives studio-quality narration in more than 130 languages. Eleven v4 is the most expressive: it laughs, sighs, or whispers where you mark it. MiniMax Speech 2.8 Turbo and HD are a fast and a higher-fidelity alternative. Each returns an MP3 in seconds.
How many credits does it cost?
0.002 to 40.0 credits, depending on the model and the length of your text. Gemini models charge only for the audio they actually produce. Results are ready in seconds.

AI Text-to-Speech Generator

Turn any text into natural-sounding speech in dozens of voices and 70+ languages. Control tone, accent, pacing, and emotion through plain-language instructions, production-ready voiceover in seconds.

What you can make

  • Documentary narration: authoritative, professional voices
  • Podcast intros, outros, dialogue: multi-character delivery
  • Multilingual voiceovers: cover the same script in 70+ languages
  • Audiobook narration: emotional pacing with mid-sentence direction
  • Accessibility: make text content available as audio

How it works

  1. Paste your text: what you want spoken
  2. Pick a voice or leave on Auto: Auto picks the best voice based on your text and instructions
  3. Add a style hint (optional), describe accent, pace, emotion: "warm British accent, slow pace with dramatic pauses"
  4. Pick a model: Gemini 3.8 Flash-Lite TTS for fast, low-cost drafts, Gemini 3.8 Flash TTS for studio-quality narration, Eleven v4 for the most expressive delivery, or MiniMax Speech 2.8

Style tips

  • Write narration in short, clear sentences for better pacing
  • Use punctuation to control pauses, periods for full stops, commas for brief pauses, ellipsis for long pauses
  • Style hints take plain English: "speak slowly and dramatically," "cheerful and upbeat," "calm and reassuring"

Explore more pipelines

Video Generator
from 9.5
Video Generator
Image Generator
0.2-6.7
Image Generator
Music Generator
2.5-22.6
Music Generator
Video Editor
from 9.5
Video Editor