What is conversational latency?

Conversational latency
Conversational latency is the delay between a caller finishing what they say and the assistant beginning its reply, measured as the silence the caller actually experiences.

What conversational latency means in practice

Four stages stack up inside that gap. Deciding the caller has finished, transcribing, working out a reply, and synthesising speech.

The first stage is often the largest and the least discussed, because waiting to be sure someone has stopped talking takes real time.

Human conversation runs on gaps of about two hundred milliseconds. Anything much past that starts reading as hesitation.

Streaming techniques hide some of it by starting to speak before the whole reply is ready, which is why two systems with the same total delay can feel very different.

Network conditions add their own share on top. A caller on a weak mobile signal experiences a longer gap than the same call from a landline, and nothing in the assistant changed.

What people get wrong

Where one second goes

Here's an illustrative budget for one turn in a phone conversation with a voice assistant. These numbers are made up to show the shape of the delay, and they aren't measurements of GreetKeeper or any other product.

The caller stops talking, and the system waits 500 milliseconds to be sure she's finished. Finalizing the transcript takes another 150. The language model needs 400 to produce the first words of a reply. Speech synthesis takes 150 to turn those into the first audio. Add about 100 for the phone network carrying sound there and back. That totals 1.3 seconds of silence on the caller's end.

Look at what dominates. The biggest single item is the deliberate wait at the start, and nothing else can begin until it's over. A vendor could halve the model's thinking time and the caller would gain a fifth of a second. Trim the opening wait by the same amount and the assistant starts cutting people off.

Measuring it with two phones

You don't need lab equipment to measure this. Tell the vendor you're recording, put the demo call on speaker, and set a second phone beside it running a voice memo app. Have a normal conversation of ten turns, then hang up.

Open the recording in any free audio editor that shows the waveform. For each turn, find where your voice ends and where the assistant's begins, and read the gap off the timeline.

Report two figures: the median and the worst. The median tells you what the call feels like. Your worst gap shows how often a caller will wonder whether the line dropped. It's often the turn where the assistant had to look something up, such as open calendar slots.

Human receptionists often take a second or more on hard questions too, but they fill it with sounds like let me see. The gap that bothers callers is the empty one.

Run the test from a cell phone in your own service area, since that's the network your callers use.

How GreetKeeper handles it

GreetKeeper publishes no latency figure. We have not benchmarked it on real calls, and our rule is that a number we have not measured stays off the page.

That is a genuine gap in what we can tell you, and the only honest substitute is hearing it.

A demo runs your own scenario, and the pauses are the thing to listen for rather than the voice.

Latency questions

What delay do callers notice?

Past roughly a second, most people register a pause. Past two, they start wondering whether the line dropped.

Why do vendors quote sub-second figures?

Usually by measuring a single component under ideal conditions. Ask what the number includes and where it was measured.

Does a filler phrase help?

A short acknowledgement covers the gap and stops it feeling dead. Overused, it becomes its own tell.

Hear it take one of your calls

Two minutes, your own scenario, no card.