Text to Speech
Convert written text into natural, expressive speech using AI voices. Choose from dozens of voices across multiple languages, control speaking rate, pitch, and emphasis, and export in professional audio formats.
Overview
The Text to Speech service transforms written content into spoken audio using advanced neural voice synthesis. The resulting speech sounds natural — with appropriate rhythm, intonation, and emphasis — rather than robotic or monotone. It supports a wide range of languages and voice types to match any content requirement.
Key Capabilities
- Over 20 AI voices across 6 languages
- SSML (Speech Synthesis Markup Language) support for fine-grained control
- Adjustable speaking rate, pitch, and volume
- Export to MP3, WAV, OGG, and M4A formats
- Up to 5,000 characters per generation request
- Pause insertion and emphasis marking
Available Voices
Voices are organized by type to help match the tone and context of your content. All voices use neural synthesis for natural-sounding output.
| Voice Type | Characteristics | Best For | Available Languages |
|---|---|---|---|
| Professional Male | Clear, authoritative, mid-range tone | Corporate narration, explainer videos, news-style content | EN, ES, FR, DE, ZH, JA |
| Professional Female | Warm, clear, articulate | E-learning, customer service, product demos | EN, ES, FR, DE, ZH, JA |
| Conversational Male | Relaxed, friendly, natural pace | Podcasts, social media content, casual narration | EN, ES, FR |
| Conversational Female | Approachable, energetic, varied intonation | Marketing content, tutorials, lifestyle video narration | EN, ES, FR |
| Neutral | Even-toned, minimal inflection | Technical documentation, legal content, accessibility | EN, DE, ZH |
| Character | Distinct personality, expressive range | Children's content, entertainment, games | EN only |
Supported Languages
| Language | Code | Voice Types Available | SSML Support |
|---|---|---|---|
| English | EN | All voice types (6 voices) | Full |
| Spanish | ES | Professional, Conversational (4 voices) | Full |
| French | FR | Professional, Conversational (4 voices) | Full |
| German | DE | Professional, Neutral (3 voices) | Partial |
| Mandarin Chinese | ZH | Professional, Neutral (3 voices) | Partial |
| Japanese | JA | Professional (2 voices) | Basic |
Speech Settings
The following parameters let you fine-tune the generated speech. All settings are optional — the defaults produce natural-sounding output without any adjustments.
| Setting | Range | Default | Effect |
|---|---|---|---|
| Speaking Rate | 0.5× – 2.0× | 1.0× | Controls the pace of speech. Values below 1.0 slow delivery; values above 1.0 speed it up. |
| Pitch | −10 to +10 semitones | 0 | Shifts the fundamental frequency of the voice. Positive values raise pitch; negative values lower it. |
| Volume | Silent / x-soft / soft / medium / loud / x-loud | medium | Sets the output loudness level without post-processing normalization. |
| Emphasis | strong / moderate / reduced / none | none | Applies SSML emphasis tags to marked words, increasing expressiveness at those points. |
| Pause Duration | 100ms – 10,000ms | Automatic | Inserts a silence of the specified length at a marked position in the text. |
| Audio Quality | 22 kHz / 44.1 kHz / 48 kHz / 96 kHz | 44.1 kHz | Determines the sample rate of the output audio file. Higher values produce larger files with better fidelity. |
SSML for Precise Control
For fine-grained control over pronunciation, pauses, and emphasis, use SSML markup directly in the text input field. Wrap text in speak tags and use standard SSML elements such as break tags for pauses, emphasis tags for stressed words, and phoneme tags for pronunciation corrections.
How to Generate Speech
After clicking Generate, a preview player appears allowing you to listen before downloading. Use the preview to check pacing and emphasis, then adjust settings if needed and regenerate. The final download button exports the file in your chosen format.
Character Limit
Each generation request supports up to 5,000 characters. For longer texts, split the content into segments — each segment will be a separate audio file. The Audio Enhancement tool can be used to merge and normalize multiple segments after generation if needed.
Output Formats & Quality
| Format | Compression | Best For | Relative File Size |
|---|---|---|---|
| MP3 | Lossy | Web, podcasts, general-purpose delivery | Small |
| WAV | Lossless | Post-production, professional editing workflows | Large |
| OGG | Lossy (open format) | Web applications, games (broad browser support) | Small |
| M4A | Lossy (AAC) | Apple ecosystem, iTunes, mobile apps | Small–Medium |