Speech-to-text (STT)
Speech-to-text (STT), also called automatic speech recognition (ASR), is the conversion of spoken audio into written words so that software can process what was said.
Also called: ASR, automatic speech recognition, speech recognition
STT is the ears of a voice agent. Every word the caller says is transcribed, usually in a stream as they speak, and the text is what the language model actually reads. The model never hears audio; if the recognizer hears "fifty" as "fifteen", the model books the wrong time with full confidence.
Recognition quality varies with accent, vocabulary, background noise and audio bandwidth, and telephone audio is the hard case on all four. Domain words, street names, product names, medical or legal terms, are where general-purpose recognizers stumble, and better systems can be primed with a business’s own vocabulary.
STT is also what produces call transcripts after the fact. Word error rate is the standard measure of its accuracy, and it is worth asking how a product handles the words it is unsure of rather than assuming perfect hearing.
