AI term

What is TTS (Text-to-Speech)?

AI that converts written text into natural-sounding spoken audio, used for voiceovers, narration and accessibility.

Text-to-speech (TTS) converts written text into spoken audio. Modern neural TTS models produce voices that sound close to a human recording, with natural pacing, intonation, and emotion, a big step up from the robotic voices of older systems. You typically pick a voice from a library (or clone one), paste or stream text in, and get an audio file or a live audio stream out. Better systems handle punctuation, acronyms, and numbers sensibly, and some let you adjust speed, tone, and pronunciation with markup. Pricing is usually per character or per minute of generated audio. Common uses include video voiceovers, audiobook narration, podcast production, IVR phone systems, in-app assistants, and accessibility features like reading articles aloud. Latency matters for real-time uses; quality and language coverage matter more for produced content.

Example

A YouTube creator pastes a 900-word script into ElevenLabs, selects a warm narrator voice, and downloads a finished voiceover MP3 instead of recording and editing it in a studio.

Why it matters

When picking a TTS tool, compare voice realism, language and accent coverage, per-character pricing, and streaming latency. A voice that is fine for an internal demo may not be good enough for a public podcast or ad. Browse the AI tools directory or the model leaderboard to put it into practice.

Related AI terms

All 36 →