Operations
How to put a human in the loop in your AI agent
The short answer
Human in the loop means three separate things. During the call, it is a live transfer to a person, triggered on safety, anger, repeated failure or a direct request rather than left to the model to decide in the moment. Before an action, it is an approval gate on anything irreversible, such as a refund above a set amount. After the call, it is a scored review of recordings whose findings actually change the agent. Most operations need all three. A deployment with none of them is unsupervised, whatever the sales page says.
This is the short version. The full treatment is in The Complete Guide to Managed AI Call Centers.
The phrase covers three different jobs
Most confusion about human in the loop comes from using one phrase for three unrelated controls, built at different times, by different people, for different reasons.
During the call, a person can take over while the caller is still on the line. Before an action, the agent proposes and a human commits, so nothing irreversible happens unsupervised. After the call, somebody reviews what happened and changes the agent because of it. Same phrase, three different controls.
They solve different problems. Live transfer protects the caller in the moment. Approval gates protect the business from irreversible mistakes. Review is the only one that makes the agent better over time, and it is the one most often skipped, because it is the only one that costs the same every week forever.
In the call: designing the transfer
The decision to hand a call to a person should be made by rules, not by the model's judgment in the moment. Some triggers belong in every deployment.
A caller explicitly asking for a person. Any safety situation, which in practice means a described emergency, a medical symptom that needs triage, or anything that sounds like harm. Any request that would require advice the agent is not allowed to give. A caller who has asked the same thing twice without getting anywhere. And audible frustration, which is worth escalating on well before the caller asks, because by the time they ask it is already a bad call.
How the transfer happens matters as much as whether it happens. A cold transfer dumps the caller into a queue to start again, which is often worse than no transfer. A warm transfer carries the context across, so the person picking up already knows the name, the reason for the call and what has been tried. If the receiving human has to ask “so what is this about,” the handoff failed even though the call connected.
And there has to be an answer for when nobody picks up. An escalation path that rings an empty desk is a design that works in testing and fails at seven in the evening.
Before an action: approval gates
Not everything an agent does is reversible. Booking an appointment is easy to undo. Issuing a refund, cancelling a subscription, releasing a hold, promising a credit or dispatching a technician at a premium rate are not.
The pattern that works is a threshold rather than a blanket rule. Let the agent complete the small, common, verifiable cases on its own, because that is where the value is, and route the rest to a person. A refund under a set amount on an order the agent can see clears automatically. Above it, the agent takes the details, tells the caller when they will hear back, and queues it for approval.
Pick the threshold from your own data rather than from instinct. Look at what your team currently approves without thinking, because that is the band the agent can safely own on day one. The threshold can move later, and it usually does, in the direction of more autonomy.
After the call: the review loop
This is the one that compounds, and the one that quietly gets dropped three months in when everyone is busy.
A working review loop samples calls rather than reading all of them, weighted toward the interesting ones: anything escalated, anything abandoned, anything where the agent hit a path nobody designed, and a random sample of ordinary calls so the normal case does not drift unnoticed. Each gets scored against the same short rubric every week, so the numbers mean something across months.
The part that makes it worth doing is the loop closing. A reviewer finding that the agent mishandles a particular objection is an observation. That objection being added to the test suite, fixed in the flow, and verified on the next pass is a system improving. Without the last step, review is a weekly meeting about problems nobody fixes.
The version that is not really human in the loop
The phrase has become a reassurance, which means it now appears on products that do not have one. Three patterns are worth recognizing.
The escalation path that exists in the flow diagram and rings a number nobody answers. The review dashboard nobody has opened since onboarding. And the approval queue with a two-day response time on a decision the caller needed in the moment, which in practice means the answer is no and the customer has already gone elsewhere.
The test is simple and worth applying to any vendor. Ask who the human is, by name or by role. Ask what happens at eight in the evening. Ask to see last week's review and what changed because of it. A real loop produces evidence. A marketing loop produces a diagram.
What to measure
Containment rate, meaning the proportion of calls the agent finished without a human, is the number everyone quotes and the easiest to game. On its own it says nothing about whether those calls went well.
Read it against escalation rate and the reasons behind each escalation, resolution rate on the contained calls, and what happened to the callers who were transferred. A rising containment rate alongside a rising callback rate is not progress. It is the agent closing calls that were not finished.
A containment rate that looks too good usually is. Treat it the way you would treat a test suite where every case passes on the first run: the likely explanation is that the escalation is not firing, not that nothing needed escalating.
Common questions
Questions people ask about this
What does human in the loop mean for an AI phone agent?
It covers three different controls. A live transfer to a person during the call, an approval gate before the agent commits an irreversible action such as a refund, and a review pass after the call where recordings are scored and the findings change the agent. Most operations need all three, and they are usually built at different times.
When should an AI agent transfer to a human?
On rules rather than judgment. A direct request for a person, any safety situation, any request for advice the agent is not permitted to give, a caller who has asked the same thing twice without progress, and audible frustration. Escalating on frustration before the caller asks is what separates a recovered call from a complaint.
What is the difference between a warm and a cold transfer?
A cold transfer passes the caller to a queue or a person with nothing attached, so they start the conversation again. A warm transfer carries the context across, meaning the person picking up already has the name, the reason for the call and what has been tried. If the receiving human has to ask what the call is about, the handoff failed.
Is a high containment rate a good thing?
Not on its own. Containment measures how many calls finished without a human, not how many finished well. Read it against escalation reasons, resolution rate and callback rate. A containment rate near 100% usually means the escalation path is not firing when it should, rather than that nothing needed escalating.
Keep reading
Want this run against your numbers?
We measure your baseline first, so the answer is arithmetic.