What is diarization?

Diarization
Diarization is the labeling of a transcript by speaker, so each line of the conversation is attributed to whoever said it rather than appearing as one block of text.

What diarization means in practice

Without it a transcript is a wall of sentences. Readable, but you cannot tell who promised what.

On a two-party phone call it is comparatively easy, because each side arrives on its own audio channel.

Conference calls and speakerphones are where it struggles. Several voices on one channel, and overlapping speech that belongs to two people at once.

Errors cluster at the handover points, where one person's last few words get attached to the next speaker.

Labels are also only as useful as the names attached to them. Speaker one and speaker two is technically diarized and no help at all when you come back to the record a month later.

What people get wrong

Who said $450? A quote dispute in four lines

Say a customer calls your garage door company at 9:12 on a Monday. Three weeks later he disputes his invoice and insists you quoted him $450. You open the record, and the undivided version reads: "So what would that run me. Most jobs like that are around 450. Okay so 450. It depends on the spring, we'd confirm on site."

Read as a block, it looks like a firm quote. With speaker labels, the middle lines split. Your side said "around 450", the customer said "Okay so 450", and your side came back with "it depends on the spring". He repeated a number, and your side qualified it. That's a very different conversation to have with him, and you can have it politely because the labeled record shows each side's words.

Questions that reveal how a vendor labels speakers

Ask first whether the recording is dual channel. On a dual-channel call each party is captured on a separate track, so labeling is close to free and rarely wrong. A mono recording mixes both voices together, and software then has to guess who is speaking from the sound of each voice. Expect that guess to get shaky when your caller and your assistant sound alike or talk at once.

Next, ask what happens on a transfer or a three-way call. Once a third voice joins, even a dual-channel setup has two people sharing one track. You should know whether the labels say "Speaker 3" or quietly fold that person into the caller.

Checking all this yourself takes one call. Share a speakerphone with a coworker, interrupt each other twice on purpose, and read what comes back. Give extra attention to the sentences around each interruption, since a few words there usually end up under the wrong name.

How GreetKeeper handles it

GreetKeeper transcripts separate the caller from the assistant, which is the attribution that matters on an inbound call.

The summary sits on top of that, so you get the outcome without reading the whole exchange.

A call transferred to a person moves out of the assistant's record, so the transcript covers the part it handled.

Diarization questions

Does it work on a speakerphone?

Less well. Several voices sharing one channel is the hard case, and background speech gets attributed to whoever spoke last.

Is it the same as transcription?

No. Transcription produces the words. Diarization decides who said them, and a system can do the first well and the second badly.

Why do the boundaries slip?

Because people talk over each other. Overlapping speech is genuinely ambiguous, and the labels drift by a few words around it.

Hear it take one of your calls

Two minutes, your own scenario, no card.