Text-to-speech (TTS)
Text-to-speech (TTS) is the generation of spoken audio from written text, producing the voice a caller hears when an AI agent replies.
Also called: speech synthesis, synthetic voice, neural TTS
TTS is the mouth of a voice agent. The language model writes a reply; the TTS engine turns it into sound, with a chosen voice, accent, pace and intonation. Modern neural TTS is close enough to human speech that the voice itself has become a brand decision rather than a technical one.
For real-time conversation the engine has to stream, start speaking the first words while the rest of the sentence is still arriving, or every reply begins with a pause. It also has to pronounce the awkward things correctly: numbers, times, addresses, names and abbreviations, which are exactly what a receptionist says most.
Voice selection is the user-facing side of TTS: choosing which of the available voices represents the business. Voice cloning goes further and builds a new voice from recordings of a specific person.
Related
aiReceptionist does this: Voice selection
