Audio Generator
Turn text into natural speech with dozens of voices, 70+ languages, and custom style directions for tone, accent, and pacing.
Or pick a specific voice...
Runs on
Examples
Meet Pyra: voiceover
Gemini 2.5 Flash TTS. Pyra's own story, read in a warm storyteller voice: a voiceover you can drop straight into any video.
How to use Audio Generator
A compact decision guide for placing this pipeline in a workflow.
Best for
- Narration and voiceover in 70+ languages
- Text-to-speech with natural language style control (e.g. 'speak warmly', 'whisper', 'excited tone')
- Multi-voice content with dozens of expressive voices
- Multilingual scripts, the language is auto-detected from the text
Tips
- Write narration in short, clear sentences for better pacing
- Use punctuation to control pauses: periods for full stops, commas for brief pauses, ellipsis for long pauses
- Leave the voice on Auto and the AI picks one that matches the style instructions, or pin a specific voice if you want a recurring character
- Use natural language style hints: 'speak slowly and dramatically', 'cheerful and upbeat', 'calm and reassuring'
Audio Generator API
Call this pipeline from your own code. One request dispatches a run; the model is an input field, not a separate endpoint.
Get an API token- inputs
- 5
- required
- 1
- Pricing
- 0.002 to 40.0 credits
| Input | Type | default | Description |
|---|---|---|---|
text | string | default— | Text to convert to speech. |
enhance_prompt | boolean | defaulttrue | Improve the prompt for the selected model before generation. Turn off to send your prompt unchanged. |
instructions | string | default— | Describe tone, accent, pacing, emotion. The AI crafts a rich style prompt. |
model | enum | defaultgemini-2-5-flash-tts | Voice engine to use.gemini-2-5-flash-tts · gemini-2-5-pro-tts · minimax-speech-2-8-turbo +1gemini-2-5-flash-ttsgemini-2-5-pro-ttsminimax-speech-2-8-turbominimax-speech-2-8-hd |
voice | string | defaultauto | Voice name from dozens available, or auto to let the AI pick. |
Point any MCP client at pipe2 and your agent gets these tools. It finds this pipeline, reads the same schema above, then runs it.
list_pipelinesget_pipeline_schemarun_pipelineget_pipeline_run_statusrequest_upload
https://mcp.pipe2.ai/mcp {
"mcpServers": {
"pipe2ai": {
"url": "https://mcp.pipe2.ai/mcp",
"headers": { "Authorization": "Bearer YOUR_TOKEN" }
}
}
}Recipes using this pipeline
Show all recipesFrequently Asked Questions
What voices are available?
What languages are supported?
How do instructions work?
What audio format is the output?
Which model should I pick?
How many credits does it cost?
AI Text-to-Speech Generator
Turn any text into natural-sounding speech in dozens of voices and 70+ languages. Control tone, accent, pacing, and emotion through plain-language instructions, production-ready voiceover in seconds.
What you can make
- Documentary narration: authoritative, professional voices
- Podcast intros, outros, dialogue: multi-character delivery
- Multilingual voiceovers: cover the same script in 70+ languages
- Audiobook narration: emotional pacing with mid-sentence direction
- Accessibility: make text content available as audio
How it works
- Paste your text: what you want spoken
- Pick a voice or leave on Auto: Auto picks the best voice based on your text and instructions
- Add a style hint (optional), describe accent, pace, emotion: "warm British accent, slow pace with dramatic pauses"
- Pick a model: Flash TTS for fast drafts, Pro TTS for studio-grade prosody on final delivery
Style tips
- Write narration in short, clear sentences for better pacing
- Use punctuation to control pauses, periods for full stops, commas for brief pauses, ellipsis for long pauses
- Style hints take plain English: "speak slowly and dramatically," "cheerful and upbeat," "calm and reassuring"