What is word error rate?

Word error rate
Word error rate is the standard measure of transcription quality, counting substituted, deleted and inserted words against a reference transcript and dividing by the number of words spoken.

What word error rate means in practice

Three error types are counted equally. A wrong word, a missing word, and a word the system invented.

Equal weighting is the measure's main flaw. Dropping the word the is scored exactly like turning yes into no.

It can also exceed 100%, which surprises people. A system that hallucinates extra words adds insertions on top of everything else it got wrong.

Comparing two vendors' published rates is close to meaningless. Different test sets, different audio, different rules about punctuation and numbers.

There is also no weighting for importance. A transcript that loses a street number scores better than one that mangles four filler words, and only one of those two errors will cost you a customer.

What people get wrong

Working one out by hand

Take a caller who says ten words: I need to move my appointment to Thursday at four. The transcript reads: I need to move my appointment Thursday at for.

Line them up and the word to before Thursday is missing, which is one deletion. Four became for, which is one substitution. Nothing was added, so insertions are zero. Two errors over ten spoken words is a word error rate of 20%.

Now look at what each error did. The missing to changes nothing, since anyone reading it still understands. Four turning into for is harmless to a person and possibly fatal to software trying to pull a time out of the sentence. Both count as one.

Insertions are how the rate can pass 100%. If a caller says yes please, and background TV chatter gets transcribed as six extra words, that's six errors over two spoken words, or 300%.

The scoring rules change the score

Before a rate is calculated, both transcripts get tidied by a set of rules called normalization. Those rules move the result more than most people expect.

Does 4 match four? Is Dr. the same as doctor, or OK the same as okay? What about punctuation, and are um and uh stripped first? A strict scorer marks every one of those as an error, and a lenient one marks none. The same transcript can score 6% under one rulebook and 14% under another, and published figures rarely say which was used.

You can measure your own in about an hour. Pick ten real calls, type out exactly what was said in each, and compare with the system's version line by line. Decide your rules first and write them down.

Then count a second number beside it: errors in names, numbers, dates and addresses only. Researchers call this an entity or keyword error rate. For a business taking bookings by phone it's the figure that matters, and it's often far worse than the headline rate.

How GreetKeeper handles it

We do not publish one. A number that meant anything would need a labeled set of real GreetKeeper calls, and no such set exists yet.

What we do instead is show you the transcript for every call, so the errors are visible rather than averaged away.

Where a specific word keeps failing, adding it to what the assistant knows about your business is usually the practical fix.

Word error rate questions

What counts as good?

Under 5% is strong on clean audio and rare on phone calls. Treat any figure without its test conditions as marketing.

Are all errors equally bad?

The metric says yes and reality says no. A dropped article costs nothing; a flipped negative changes the meaning of the call.

Can it be improved for my business?

Giving the system your service names, staff names and common local places helps most. Those are the words it has least reason to expect.

Hear it take one of your calls

Two minutes, your own scenario, no card.