ScaileAI  /  Blog  /  Operations

Operations

What to know before you launch your AI agent

The short answer

Before an AI phone agent takes a real call, four things should already exist. A recorded baseline of how calls are handled today, because without it no improvement can be proven. A written list of what the agent is not allowed to do, enforced in the flow rather than hoped for in a prompt. An escalation path to a named human, tested. And an adversarial test pass, because agents fail on confused and hostile callers, not on the happy path everyone demos. Launching narrow, on one call flow, is how you find the rest cheaply.

This is the short version. The full treatment is in The Complete Guide to Managed AI Call Centers.

Record the baseline, or you will argue about it later

The single most common regret is not measuring the before. Six weeks after launch somebody asks whether this is working, and the honest answer is that nobody knows, because nobody wrote down what was happening in August.

The numbers worth capturing are unglamorous and take an afternoon. How many calls come in. How many are answered, split by hours and after-hours. How long callers wait before hanging up. How many reach voicemail and how many of those ever get called back. What proportion turn into whatever counts as a win in your business.

Capture them from your phone system rather than from memory, because memory is generous about answer rates. This baseline is also the thing that protects the project internally. A number everyone agreed on in advance is very hard to argue with afterwards.

Write down what it is not allowed to do

Every agent needs a short, specific list of refusals, and it should exist on paper before anyone builds anything. Not a tone instruction in a prompt. A list.

The contents depend on the business and the obvious ones recur. No medical advice and no clinical reassurance. No legal advice and no opinion on whether someone has a case. No discounts the agent was not given authority to offer. No commitment to a delivery date the operation has not promised. No confirmation of an order that has not actually been placed.

Two things make this list real rather than decorative. It has to be enforced structurally, which usually means the agent cannot reach that state in the flow at all, rather than being asked nicely not to. And each refusal needs a destination, because “I cannot help with that” is where callers hang up and complain. Every no should be followed by a route to a yes.

Decide how it introduces itself, and check the law

Disclosure is the cheapest thing to get right and one of the more expensive to get wrong. Callers who know what they are speaking to are more direct, and the disclosure removes the moment of discovery that turns a good call into a screenshot.

Several states now regulate AI disclosure in commercial calls, and the rules are moving. Separately, call recording consent varies by state, and two-party consent jurisdictions require the notice at the top of the call rather than somewhere in a privacy policy. Both notices belong in the opening line, scripted, not left to the model to remember.

Worth knowing how this fails: an agent can disclose that the call is recorded, mention privacy, and still never say it is AI. That reads as compliant in a transcript review and is not. Test the greeting as a sentence a stranger hears, not as a checklist.

Test the calls nobody demos

Agents do not fail on the happy path. They fail on the caller who is angry, the caller who changes their mind halfway, the caller who answers a different question than the one asked, and the caller who says something the script never anticipated.

A serious pre-launch test pass covers the interruption, where someone talks over the agent mid-sentence. The correction, where a caller gives a date and then changes it. The silence, where nobody speaks for eight seconds. The second objection, because handling the first one and then folding is worse than not trying. The out-of-scope request. The emergency buried in an ordinary sentence. And the caller who simply asks to speak to a person.

Numbers and names deserve their own pass, because that is where phone audio is weakest. Have the agent spell back anything that will end up in a record, and listen to whether the spell-back sounds like a person confirming or a machine reciting.

Agree who gets called when it goes wrong

Before launch there should be a named human, a route to them, and a tested handoff. Not a plan to build one. A working path, dialed and confirmed.

Decide what triggers it. Safety situations and anything medical or legal should escalate automatically rather than on request. So should a caller who has asked twice and is not getting anywhere, and anyone who explicitly asks for a person. Decide what the human receives at the moment of transfer, because a warm transfer that arrives with no context makes the caller repeat everything and wastes the entire call.

Also decide what happens out of hours, when the named human is asleep. That answer can be a callback commitment, as long as somebody is actually going to make it.

Launch narrow

The strongest predictor of a good first ninety days is a small first scope. One call flow, one number, one measurable outcome.

There is a real temptation to go live across everything at once, usually because the business case was built on everything at once. Resist it. A narrow launch means the failures are findable, the fixes are fast, and the team builds confidence on something they can see working before volume arrives.

Expansion should follow the numbers rather than the contract. If the first flow does not beat the baseline, adding three more will not fix it, and it will make the diagnosis much harder.

Know what you are watching in week one

Daily for the first week, then weekly. Listen to calls rather than reading dashboards, at least at first, because the dashboard will tell you the containment rate was 84% and will not tell you the agent sounds impatient.

The handful worth tracking: how many calls the agent handled without a human, how many escalated and why, how often it hit a path nobody designed, how long the average call ran against the human baseline, and what proportion of records came out complete enough to act on.

One counterintuitive signal. A containment rate of 100% in week one is not good news. It usually means the escalation path is not firing when it should, and somebody is being told no who should have been passed to a person.

Common questions

Questions people ask about this

How long does it take to launch an AI phone agent?

A single, well-scoped call flow typically goes live in weeks rather than quarters. The critical path is almost never the agent itself. It is integration complexity, how fast the business can supply its own policies and exceptions, and how long script approval takes internally.

Does an AI agent have to say it is AI?

Several states now regulate AI disclosure in commercial calls and the rules are still moving, so the compliant answer depends on where the callers are. The practical answer is to disclose regardless. Callers who know what they are talking to are more direct, and nobody has ever complained about being told.

What should be tested before an AI agent goes live?

The calls nobody demos. Interruptions, mid-call corrections, long silences, repeated objections, out-of-scope requests, emergencies hidden in ordinary sentences, and callers who ask for a person. Names and numbers need a separate pass, because narrowband phone audio is weakest on exactly the content that ends up in your records.

Should we launch an AI agent on all our calls at once?

No. One call flow, one number, one measurable outcome. A narrow launch makes failures findable and fixes fast, and it gives you a clean comparison against your baseline. Expansion should follow the numbers rather than the original plan.

Keep reading

Want this run against your numbers?

We measure your baseline first, so the answer is arithmetic.