Fundamentals
The building blocks of a great AI voice agent
The short answer
An AI voice agent is a chain of six systems. Voice activity detection hears that someone is talking. Endpointing decides they have finished. Speech recognition turns the audio into text. A language model decides what to say. Speech synthesis turns that back into audio. Telephony carries it down the line. The whole round trip has to finish in about a second, and every component spends from the same budget. Agents fail at the seams between these parts far more often than inside any one of them.
This is the short version. The full treatment is in The Complete Guide to Managed AI Call Centers.
The budget everything is spent from
In ordinary conversation people leave about 200 milliseconds between turns. Push past roughly 800 and the caller starts to feel it. Past 1,500 and they say hello again, or assume the line dropped and talk over the answer.
That is the constraint the whole stack is built around. Six systems have to run in sequence inside about a second, which makes the interesting engineering question almost never “which model is smartest.” It is where the milliseconds are going.
A useful way to watch any vendor demo is to ignore the answers and listen to the gaps. Good answers are table stakes now. Natural turn-taking is not.
Voice activity detection: is anyone talking?
Voice activity detection, usually shortened to VAD, is a small fast classifier that runs continuously and answers one question. Is this audio speech, or is it not.
That sounds trivial and it is not, because phone lines are full of things that are almost speech. A television in the next room. Road noise. Hold music bleeding in from a transfer. Set VAD too sensitive and the agent flinches at a door closing. Set it too dull and it misses a quiet caller on a cell phone in a parking lot.
VAD is also what makes barge-in possible, which is the ability to cut across the agent mid-sentence and actually be heard rather than talking into a monologue that carries on regardless. An agent that cannot be interrupted feels like a recording. Callers treat it like one. If you test one thing in a demo, talk over it.
Endpointing: have they finished?
This is the hard one, and it is where most agents give themselves away.
Knowing that someone stopped making noise is easy. Knowing they are done is not. People pause mid-sentence to think. They pause after “my policy number is” and then go looking for the card. The naive approach waits for a fixed stretch of silence, usually somewhere between 500 and 800 milliseconds, then responds. That produces the two failures everyone has lived through: the agent cutting off a caller who was mid-thought, and the agent sitting in silence after an obviously complete sentence.
Better systems use semantic endpointing, which means looking at what was actually said and judging whether it is finished. “I need to reschedule my” is clearly incomplete. “I need to reschedule my appointment” is not. The same half second of silence should mean different things in those two cases, and on a good agent it does.
When an agent feels rude, endpointing is usually why. When it feels sluggish, endpointing is usually why as well, because somebody turned the threshold up to stop it interrupting.
Speech recognition: what did they say?
Automatic speech recognition, or ASR, converts audio into text. Modern ASR is remarkable on clean audio and noticeably worse on a phone call, for a reason worth understanding.
Telephone audio is narrowband. A traditional call runs at 8 kHz through a codec like G.711, which discards most of the frequency range the model was trained on. The crispness that separates an F from an S lives up in the part of the spectrum the phone network throws away. This is why recognition that is flawless on a podcast makes mistakes on a landline, and why agents stumble on exactly the content that matters most.
That content is names, numbers, addresses and email. A model that handles a whole sentence perfectly will still render one street name three different ways across a single call. The fix is not a better model. It is design. Spell critical values back to the caller and confirm them, bias the recognizer toward the vocabulary it should expect, and never let an unconfirmed name or number land in a record somebody will later rely on.
The language model: what should we say?
This is the part everyone talks about and it is rarely the bottleneck. What matters operationally is not raw capability but two other things.
First is time to first token. A model that streams its opening words in 300 milliseconds and finishes in two seconds feels faster than one that thinks silently for a second and then delivers the whole answer at once. The agent can begin speaking before the sentence has finished generating, so the caller hears someone starting to answer rather than a machine computing.
Second is tool use, sometimes called function calling. An agent that can only talk is a brochure. An agent that can check a slot, look up an order, verify coverage and write a record is doing the job. Every one of those lookups also spends from the latency budget, which is why good agents narrate while they work. “Let me pull that up” is not filler. It buys 800 milliseconds and sounds like what a person would say.
Speech synthesis, and the line itself
Text to speech has stopped being the weak link. Two things still separate a good voice from a bad one, and neither is fidelity. The first is time to first audio byte, which matters for the same reason streaming matters on the model: the caller needs to hear something start. The second is prosody, meaning rhythm and emphasis. It is the difference between a person reading a phone number and a sequence of digits.
Then there is the telephony leg, which teams forget until it bites them. The audio crosses a carrier network with its own jitter, packet loss and codec. A stack that performs beautifully in a browser test can shed 200 milliseconds and a good deal of audio quality the moment a real call is involved. Test on the phone. Test from a cell phone. Test from a car.
The parts that are not in the stack diagram
Everything above is infrastructure, and infrastructure is increasingly something you buy rather than build. The difference between an agent that works for a week and one that works for a year lives in the layer nobody draws.
Conversation design is the map of every path a caller can take, including the hostile and confused ones. Guardrails are the written list of things the agent will not do and what it does instead, enforced in the flow rather than hoped for in a prompt. Knowledge engineering turns a company's policies, exceptions and objections into something the agent can answer from. Evaluation is a scored test suite that runs before any change ships, so improving one path does not quietly break three others.
A team can buy the strongest component for every box in that diagram and still put out an agent people hate talking to. The components are not the product. The conversation is.
Common questions
Questions people ask about this
What is VAD in a voice agent?
VAD stands for voice activity detection. It is a lightweight classifier running continuously on the incoming audio, deciding whether what it hears is speech. It is how an agent knows a caller has started talking, and it is what makes interruption possible, because the agent has to detect the caller's voice while it is still speaking in order to stop.
Why do AI phone agents interrupt people?
Almost always endpointing, which is the decision about whether a caller has finished rather than merely paused. A system waiting for a fixed stretch of silence will cut off anyone who stops mid-sentence to think or to read something out. Agents that handle this well judge the content of what was said, not only the silence after it, so an incomplete sentence earns more patience than a complete one.
Why does speech recognition get names and numbers wrong on phone calls?
Phone audio is narrowband, typically 8 kHz, and the codec discards much of the frequency range that distinguishes similar sounds. Recognition models trained on wideband audio lose accuracy on exactly the consonants that separate one letter or digit from another. The practical answer is design rather than a better model: confirm critical values back to the caller, and never write an unconfirmed name or number into a record.
How fast does an AI voice agent need to respond?
Human conversational turn-taking sits around 200 milliseconds. An agent answering inside about 800 feels natural to most callers. Past roughly 1,500 milliseconds people assume the line has dropped and start talking again. That budget has to cover endpointing, recognition, any data lookups, the language model and speech synthesis combined.
Keep reading
Want this run against your numbers?
We measure your baseline first, so the answer is arithmetic.