Endpointing

Endpointing is detecting where a spoken utterance begins and ends in a stream of audio, so that speech recognition and the reply are triggered at the right moment.

Also called: end-of-speech detection, utterance segmentation

Speech arrives as a continuous stream. Endpointing draws the boundaries, this is where the caller started, this is where they stopped, so the recognizer knows what to transcribe and the agent knows when to reply. It is usually based on a silence threshold: after a certain number of milliseconds without speech, the utterance is declared over.

The threshold is a trade-off. A short one makes the agent responsive but prone to cutting in during a caller’s natural pause; a long one is patient but leaves dead air after every sentence. Adaptive endpointing shortens or lengthens the wait depending on whether the words so far sound like a complete thought.

For a buyer the term matters because it is the root cause of two common complaints about voice agents, "it interrupts me" and "it takes forever to answer", which are the same setting pulled in opposite directions.