What is speech recognition accuracy?

Speech recognition accuracy
Speech recognition accuracy is the share of spoken words a system transcribes correctly, measured against what a human listener agrees was actually said.

What speech recognition accuracy means in practice

Headline figures come from clean test recordings. Quiet rooms, good microphones, speakers who finish their sentences.

A phone call has none of that. Compressed audio, background noise, two people talking over each other, and a caller walking past a lorry.

The words that matter most are the ones it gets wrong most. Names, street addresses and email spellings are unusual strings with no context to guess from.

Accuracy also varies by speaker in ways that are worth testing rather than assuming. Accent, speed and phone quality all move it.

What people get wrong

Twenty-two words, two errors, one wrong house

Say a caller tells a plumbing company: Hi, this is Dana Schaefer, I'm at 1415 Alder Court. My water heater has been leaking into the garage since this morning.

The transcript comes back with Dana Shafer at 1450 Alder Court, and everything else is perfect. Twenty of 22 words are right, which is 91%, and the plumber drives to the wrong house.

Fifteen and fifty are a classic confusion, because the endings sound alike on a phone line and nothing in the sentence says which is right. Her surname is the same story: several spellings, one sound.

Notice which words survived: water heater, leaking and garage all came through, since they're common words with a sentence around them to help. The parts that failed are the parts a dispatcher needs most. A score that treats all 22 words equally says 91%, and a score counting only the name and address says zero for two.

Score the fields, and fix them in the conversation

When you test any voice system, grade it on fields and not on words. Make a sheet with four columns: name, callback number, address, and reason for calling. Run ten test calls using real customer details, with different people calling from a cell, a car and a noisy kitchen. Mark each field right or wrong. That's forty fields, and you want to know how many are usable without a callback.

Most of the repair happens in how the conversation is run. A few habits help a human receptionist and an AI one equally. Ask for numbers in small groups, and have callers spell a surname. For anything the job depends on, read it back and wait for a yes.

Text is the other backstop, since a confirmation message showing the address and time lets the customer catch an error hours before a truck rolls. Ask any vendor to show you confirmation on a live call with a difficult name, and bring your own worst one.

How GreetKeeper handles it

GreetKeeper publishes no accuracy figure. We have not benchmarked one on real calls, and a number copied from a vendor's test set would tell you nothing about your callers.

The assistant reads the name, number and time back before it books. A misheard digit gets corrected by the caller on the line instead of surfacing as a wrong callback.

Transcripts show you the errors rather than hiding them, which means you can judge the tool on your own calls within a week.

Accuracy questions

Why are names transcribed badly?

Because they carry no context to fall back on. A model can guess a common word from the sentence around it, and a surname gives it nothing to work with.

Does a bad line make it worse?

Considerably. Compression and packet loss remove detail a human ear fills in from expectation, and the model cannot do that trick as well.

How should I test it?

Have five real callers ring in, then read the transcripts against what you remember. Ten minutes of that beats any published benchmark.

Hear it take one of your calls

Two minutes, your own scenario, no card.