Diarization
Diarization is dividing an audio recording by speaker, working out who spoke when, so that a transcript can attribute each passage to the right person.
Also called: speaker diarization, speaker separation, who-spoke-when
A raw transcript is just words in order. Diarization adds the labels: this line was the caller, that line was the agent, this part was a third person who joined. On a two-party call with separate audio channels it is trivial; on a single mixed recording, or when a colleague picks up an extension, it has to be inferred from the voices themselves.
It matters wherever a transcript is read or analyzed. A summary that attributes the caller’s complaint to the receptionist is worse than useless, and sentiment analysis depends on knowing whose sentiment is being measured.
Diarization is a property of transcription and recording, not of the live conversation; the agent does not need it to hold the call, but everything downstream of the call does.
