TotalApp Docs

Text to Speech

Convert written text into natural, expressive speech using AI voices. Choose from dozens of voices across multiple languages, control speaking rate, pitch, and emphasis, and export in professional audio formats.

Overview

The Text to Speech service transforms written content into spoken audio using advanced neural voice synthesis. The resulting speech sounds natural — with appropriate rhythm, intonation, and emphasis — rather than robotic or monotone. It supports a wide range of languages and voice types to match any content requirement.

Key Capabilities

  • Over 20 AI voices across 6 languages
  • SSML (Speech Synthesis Markup Language) support for fine-grained control
  • Adjustable speaking rate, pitch, and volume
  • Export to MP3, WAV, OGG, and M4A formats
  • Up to 5,000 characters per generation request
  • Pause insertion and emphasis marking

Available Voices

Voices are organized by type to help match the tone and context of your content. All voices use neural synthesis for natural-sounding output.

Voice Type Characteristics Best For Available Languages
Professional Male Clear, authoritative, mid-range tone Corporate narration, explainer videos, news-style content EN, ES, FR, DE, ZH, JA
Professional Female Warm, clear, articulate E-learning, customer service, product demos EN, ES, FR, DE, ZH, JA
Conversational Male Relaxed, friendly, natural pace Podcasts, social media content, casual narration EN, ES, FR
Conversational Female Approachable, energetic, varied intonation Marketing content, tutorials, lifestyle video narration EN, ES, FR
Neutral Even-toned, minimal inflection Technical documentation, legal content, accessibility EN, DE, ZH
Character Distinct personality, expressive range Children's content, entertainment, games EN only

Supported Languages

Language Code Voice Types Available SSML Support
English EN All voice types (6 voices) Full
Spanish ES Professional, Conversational (4 voices) Full
French FR Professional, Conversational (4 voices) Full
German DE Professional, Neutral (3 voices) Partial
Mandarin Chinese ZH Professional, Neutral (3 voices) Partial
Japanese JA Professional (2 voices) Basic

Speech Settings

The following parameters let you fine-tune the generated speech. All settings are optional — the defaults produce natural-sounding output without any adjustments.

Setting Range Default Effect
Speaking Rate 0.5× – 2.0× 1.0× Controls the pace of speech. Values below 1.0 slow delivery; values above 1.0 speed it up.
Pitch −10 to +10 semitones 0 Shifts the fundamental frequency of the voice. Positive values raise pitch; negative values lower it.
Volume Silent / x-soft / soft / medium / loud / x-loud medium Sets the output loudness level without post-processing normalization.
Emphasis strong / moderate / reduced / none none Applies SSML emphasis tags to marked words, increasing expressiveness at those points.
Pause Duration 100ms – 10,000ms Automatic Inserts a silence of the specified length at a marked position in the text.
Audio Quality 22 kHz / 44.1 kHz / 48 kHz / 96 kHz 44.1 kHz Determines the sample rate of the output audio file. Higher values produce larger files with better fidelity.

SSML for Precise Control

For fine-grained control over pronunciation, pauses, and emphasis, use SSML markup directly in the text input field. Wrap text in speak tags and use standard SSML elements such as break tags for pauses, emphasis tags for stressed words, and phoneme tags for pronunciation corrections.

How to Generate Speech

1. Open Text to Speech
2. Enter Text
3. Select Voice & Language
4. Adjust Settings
5. Generate & Export

After clicking Generate, a preview player appears allowing you to listen before downloading. Use the preview to check pacing and emphasis, then adjust settings if needed and regenerate. The final download button exports the file in your chosen format.

Character Limit

Each generation request supports up to 5,000 characters. For longer texts, split the content into segments — each segment will be a separate audio file. The Audio Enhancement tool can be used to merge and normalize multiple segments after generation if needed.

Output Formats & Quality

Format Compression Best For Relative File Size
MP3 Lossy Web, podcasts, general-purpose delivery Small
WAV Lossless Post-production, professional editing workflows Large
OGG Lossy (open format) Web applications, games (broad browser support) Small
M4A Lossy (AAC) Apple ecosystem, iTunes, mobile apps Small–Medium