Latency (time to first token)

In voice AI, latency is the delay between the caller finishing an utterance and the agent starting to speak in reply. Time to first token (TTFT) is the language-model portion of that delay, measured to the first word generated.

Also called: time to first token, TTFT, response latency

A spoken reply passes through several stages, endpointing decides the caller has stopped, speech-to-text finishes transcribing, the language model produces a response, and text-to-speech renders the first audio. Each adds to the gap the caller hears. Time to first token measures the model stage alone: how long until it emits the first piece of its answer.

On a phone call the whole gap is what matters. Human conversation tolerates only a short pause before silence starts to feel like a fault, so voice systems stream every stage, transcribing while the caller speaks, generating while the model thinks, speaking while the rest of the sentence is still being written, to make the first word arrive early even if the full reply takes longer.

Latency is a fair thing to measure on a demo call. It is also a trade-off against reasoning depth and against tool use: an agent that looks something up mid-call takes longer to answer than one that does not, and good products cover the wait conversationally rather than with silence.