What is a synthetic voice?

Synthetic voice
A synthetic voice is speech produced by software rather than recorded from a person, generated from text at the moment it is needed.

What synthetic voice means in practice

The old approach stitched recorded fragments together. It worked for phone menus and fell apart on anything it had no clips for.

Modern systems generate the waveform directly, which is why they can say a surname they have never seen and still sound like speech.

Quality now varies more by sentence than by system. A short reply sounds excellent, and a long one with unusual punctuation can wander.

The improvement created the legal question. When software sounds close enough to a person, some states want the caller told.

What people get wrong

Five lines to audition a voice with

Say you run Kowalczyk Heating on Schuylkill Avenue. No demo reel contains either of those words, so write your own test. Open with your greeting, business name included. Second comes your street address, read the way you'd give it to a driver. Third is a phone number, because some voices rush ten digits into one long blur.

Your fourth line is a price with cents, such as $189.50. Finish with a question that should rise at the end, like "Would Thursday at 2 work for you?" Play all five through a phone speaker and skip the laptop. A phone line strips out much of the detail that made the voice sound rich in your headphones. Pick whichever voice gets your name right and keeps the question sounding like a question, even if another one sounded warmer on the reel.

How it differs from a recorded greeting

A recorded greeting is a person reading fixed lines into a microphone once. It sounds exactly right and it can't say anything new. Change your hours and you're back at the microphone, or more likely the old hours stay on your line for six months.

Synthetic speech is made from whatever text arrives at that moment. That's what lets an assistant say your caller's name or a street it has never seen. What you give up is control. With a recording you approved every syllable. Synthesis lets you approve the words while the engine decides the delivery each time, so one sentence can come out slightly differently on Monday and on Friday. That's the reason testing matters more than it did with recordings.

People also mix it up with voice cloning. Cloning copies one real person. A stock synthetic voice belongs to nobody, which keeps consent paperwork off your desk.

How GreetKeeper handles it

You choose from the voices available and write every word the assistant says. The script decides more of the impression than the voice does.

GreetKeeper does not claim to be indistinguishable from a person, and we would not want to. Maine's law turns on exactly that kind of confusion.

The assistant can announce itself at the start of the call, and Utah requires that verbally for licensed professions such as medicine and law.

Synthetic voice questions

Can callers tell?

Some do and many do not, and it usually depends on what they called about rather than on the voice. Complicated problems reveal the seams fastest.

Does a better voice reduce the need to disclose?

The opposite. The closer it gets to a person, the more likely a caller is to be misled, which is the thing the Maine statute addresses.

Why do some words sound wrong?

Unusual names and local place names are the usual culprits. Writing them phonetically in the script fixes most of them.

Hear it take one of your calls

Two minutes, your own scenario, no card.