Why an AI voice starts to mumble

An AI receptionist gets the last word of a sentence wrong more often than the rest, and after about two minutes of conversation it suddenly gets much worse. We looked into why.

Sep 28, 2026 · 5 min read

‘Waarmee kan ik u helpen?’ (How can I help you?) Our AI receptionist answers calls in Dutch, and this is the sentence it says most often. Yet it is the last word that sometimes goes wrong: ‘helpen’ (to help) then sounds like ‘lopen’ (to walk), ‘noken’ or ‘melken’ (to milk). We measured how often this happens and what causes it. Some of what we found surprised us too.

Thinking and speaking in one breath

Our AI receptionist uses a speech-to-speech model. A model like this produces sound directly: it works out what to say and says it in one go, without a text step in between. That is exactly why a conversation with it sounds so natural. The model hears how someone says something, responds in the right tone and does not go quiet while a separate voice first reads out a sentence.

There is one quirk, though. At the end of a sentence the model sometimes seems to be in a hurry: the last syllable gets swallowed, or the word sounds like a different word. We see this in every speech-to-speech model we have tested. So it is not a problem of one vendor, but of this way of producing speech.

The last word

To measure how often it happens, we have a second AI model judge the last word of every sentence: does it sound like the word that was intended? We do the same for a sample of words in the middle of a sentence, for comparison.

In a selection from our production traffic, just over 1,900 calls on a single day, 8.0% of last words went wrong, against 1.5% in the middle of a sentence. The word that went wrong most often was ‘helpen’ (to help), the very word the standard opening question ends on. Among other things, the judging model heard ‘lopen’ (to walk), ‘open’, ‘noken’, ‘melden’ (to report) and ‘melken’ (to milk).

Notably, ‘natuurlijk’ (of course) and ‘doorverbinden’ (to put through) almost never went wrong, even though ‘doorverbinden’ ends in -en just like ‘helpen’. So it is not a simple rule about one word ending.

‘Helpen’ (to help) is the word that fades most often, right at the end of the standard opening question.
‘Helpen’ (to help) is the word that fades most often, right at the end of the standard opening question.

After two minutes, things change

The biggest surprise was in the timing. In test calls the rate stays at around 1 to 2% for two minutes. After that it jumps five to eight times higher. The longer the call lasts, the more often the last word fades.

For the first two minutes it almost always goes well. After that the last word fades more often.
For the first two minutes it almost always goes well. After that the last word fades more often.

You can see this within a single call too. If the receptionist says ‘helpen’ several times, the last time goes wrong about one and a half times as often as the first.

What it was not

We had suspects. To tell them apart, we ran 560 test calls with six variants of the same receptionist. The variants took turns, so the time of day played no role. In each variant we changed one thing:

  • without the background knowledge about the company that the receptionist gets in advance: no difference
  • without the extra pronunciation rules, our main suspect: no difference
  • with a different voice: no difference
  • with instructions in a language other than English, the model's base language: no improvement, if anything slightly worse

That last point confirms something we were already doing: we write the receptionist's instructions in English, even when the call is in Dutch.

What did make a difference

One variant made a real difference: the receptionist without the tools it uses during a call, such as looking up information or putting a caller through. Late in the call the number of errors dropped by about 37%. The pattern stayed the same, though: still a jump after two minutes, just a smaller one.

Without tools a receptionist cannot put anyone through, of course, so this is not a solution. It is a clue: everything that piles up during a call seems to make the model's job heavier.

Is it the model or the line?

Finally, we compared the same words in two recordings: the sound as the model produced it, and the sound as it reached the caller over the phone line. About 72% of the errors are already in the model's own audio. The rest arises along the way, between our platform and the phone line.

Why we still choose speech-to-speech

You could say: then take the old route, text first and a separate voice after that. But that costs you what callers notice most. A conversation with a speech-to-speech model is much more natural: faster, with the right intonation, and the receptionist simply lets itself be interrupted. That gain is a lot bigger than a swallowed syllable.

So we work around it. The part that arises along the way, we tackle ourselves. For the part in the model, we are looking for ways to keep the last word intact, even when a call runs longer.

And time is on our side. Speech-to-speech models are improving fast, and each new generation will bring improvement here anyway. Because we now know exactly how often the last word goes wrong, we have a yardstick. When a new model comes out, we run the same test and know within a day whether it is better, and by how much. That way we choose on numbers, not on gut feeling.

As with the red tomcat paradox: a good AI receptionist is not one knob you set correctly. It is measuring, changing one thing and measuring again. You can hear how that sounds for yourself on our website.

of last words go wrong
8.0%
more errors after 2 min
5-8×
of errors from the model
72%