TotalApp Docs

Audio Creator

Text-to-speech, voice cloning, music generation, and sound effects — all AI-powered.

Overview

Audio Creator bundles four separate audio generation tools into one screen. Each tool occupies its own tab and handles a distinct type of audio output. All generated audio is downloadable as MP3 or WAV and can be integrated with Video Creator for full audiovisual productions.

Text to Speech

Convert written text into natural-sounding speech. 30+ voices, adjustable speed and pitch. Preview before downloading.

Voice Cloning

Upload a 30-second audio sample and create a custom voice that can then be used for TTS generation.

Music Generation

Describe a mood, genre, tempo, and duration. AI generates original background music — royalty-free and unique each time.

Sound Effects

Describe a sound and AI generates it. From thunder cracks to office ambience to sci-fi soundscapes.

Text to Speech

The TTS tab lets you convert any written text into audio using one of 30+ AI voices. This is the fastest path from script to audio file — no recording equipment required.

Voice Selection

Browse voices by:

  • Language — English, Spanish, French, German, Turkish, Japanese, and more.
  • Gender — Male, Female, Neutral.
  • Style — Conversational, Formal, News Anchor, Narration, Energetic, Calm.

Click a voice name to hear a short preview before selecting it for your conversion.

Adjustments

  • Speed — 0.5× (very slow) to 2.0× (fast). 1.0× is the natural pace for most voices.
  • Pitch — shift pitch up or down by up to 20 semitones. Use sparingly — extreme pitch shifts can degrade naturalness.
  • Pauses — add markup tags like <break time="1s"/> in your text to insert deliberate pauses for dramatic effect or clarity.

Long Text

For narrations longer than 500 words, split the text into paragraphs and generate each paragraph separately. This gives you finer control over pacing and allows you to regenerate individual sections without redoing the whole script.

Voice Cloning

Voice Cloning creates a custom AI voice from an audio sample of a real person speaking. The cloned voice can then be used as a TTS voice for any text.

Requirements for a Good Clone

  • Minimum 30 seconds of clear speech. More sample = higher fidelity clone. 2–3 minutes is ideal.
  • Clean audio — minimal background noise, no music, no other speakers. A quiet room recording is best.
  • Natural speech — the sample should include varied sentence structures and natural pacing, not a monotone reading.
  • Format — MP3, WAV, or M4A. Sample rate 16kHz minimum; 44.1kHz preferred.

Consent Required

Only clone voices for which you have explicit consent from the speaker. Cloning a voice without consent may violate privacy laws in your jurisdiction. TotalApp logs all voice clone creation events for audit purposes.

Once cloned, the voice appears in your Personal Voices list in the TTS tab and can be used exactly like any built-in voice.

Music Generation

Describe the music you need and the AI composes an original track. Generated music is royalty-free and unique — no two generations of the same prompt are identical.

Useful descriptors to include in your music prompt:

  • Genre — jazz, classical, lo-fi hip hop, cinematic orchestral, electronic ambient, acoustic folk.
  • Mood — uplifting, tense, nostalgic, mysterious, playful, melancholic, energetic.
  • Tempo — "slow and relaxed", "medium tempo", "fast-paced and driving", or a BPM number like "120bpm".
  • Instrumentation — "piano and strings", "solo acoustic guitar", "full orchestra", "synthesiser and drums".
  • Duration — 15, 30, 60, or 120 seconds.

Example prompt: "Upbeat jazz trio, piano with double bass and brushed drums, 90bpm, warm and playful, 30 seconds, suitable for a product demo video background".

Sound Effects

Generate one-off sound effects by describing what you want to hear. The AI is good at both recognisable real-world sounds and abstract or fantastical effects.

Example prompts:

  • "Thunder crack followed by rumbling that fades over 4 seconds"
  • "Busy coffee shop ambience, espresso machine, background chatter, soft jazz in the distance"
  • "Sci-fi door hiss opening, low mechanical whir"
  • "Notification ding, short and pleasant, mobile app sound design style"
  • "Footsteps on gravel, medium pace, 10 seconds"

Export and Integration

All audio from Audio Creator is exported as:

  • MP3 — compressed, smaller file size. Good for web use and social media.
  • WAV — lossless, larger file size. Use for professional production where quality is critical.

To pair audio with a video from Video Creator, download the audio from Audio Creator, then import it into the Video Creator timeline editor in the Audio Track slot. The Audio Track supports one background music track and one voiceover track simultaneously.

Frequently Asked Questions

Can I use a cloned voice for any text I want in Text to Speech?
Yes — once a voice is cloned, it appears in your Personal Voices list inside the Text to Speech tab and works exactly like any built-in voice, including speed, pitch, and pause markup adjustments. Just remember that voice cloning requires explicit consent from the speaker whose voice you're cloning; TotalApp logs all clone creation events for audit purposes.
What's the best way to generate audio for a long narration or script?
Split text longer than 500 words into separate paragraphs and generate each one individually rather than converting the whole script in one pass. This gives you finer control over pacing between sections and lets you regenerate a single paragraph if it doesn't sound right, instead of redoing the entire narration.
Is generated music and sound effects royalty-free?
Yes. Both Music Generation and Sound Effects produce original, royalty-free output — no two generations of the same prompt are identical, so you always get a unique result safe to use in commercial projects without licensing concerns.
How do I combine Audio Creator output with a video?
Download your audio as MP3 or WAV from Audio Creator, then import it into the Video Creator timeline editor's Audio Track slot. The Audio Track supports one background music track and one voiceover track running simultaneously, so you can layer narration over generated music.