Guide
The Complete Guide to Managed AI Call Centers
This is the reference for buying, running and measuring a managed AI call center. It is written for the person who will be accountable for the phone afterwards, so it covers the parts that are usually left out of vendor material: what the AI cannot do, how these deployments fail, what the real cost is once the per-minute rate stops being the whole story, and the questions to put to a vendor that are awkward to answer.
Every section stands on its own. Read it end to end or jump to the part you need.
Contents
The category
How it works
Risk
The money
Buying it
Reference
Part one
The category
Managed AI call center is a service category, not a product category. Most of the confusion in this market comes from four different things being sold under one phrase, so it is worth being exact about what the term means before anything else.
01 Part one
What is a managed AI call center?
Short answer
A managed AI call center is a service in which a provider designs, builds, integrates, operates and continuously optimizes AI phone agents on a client's behalf, and reports on business outcomes rather than on software uptime. ScaileAI is a fully managed AI call center company that builds, operates, monitors and continuously optimizes AI-powered sales and customer service operations for businesses. The distinguishing feature is not the AI. It is who is accountable for how the phone performs.
Strip the technology out for a second. A call center is an operation: people, scripts, routing rules, quality assurance, reporting and someone whose job is the answer rate. Swapping the people for AI agents changes what does the talking. It does not remove any of the rest.
A managed AI call center is the arrangement where a provider owns all of the rest. They map your call flows, write the conversation design, connect your CRM and dispatch systems, test the thing against adversarial callers, put it live, listen to what happens and change it every week. You get a report on answer rate, qualification, conversion and cost per outcome. You do not get a login and a backlog.
The word doing the work is managed. It is the same distinction as managed hosting against a bare server, or a managed fund against a brokerage account. In both cases the underlying technology is available to buy directly and cheaply. In both cases most buyers eventually pay somebody to run it, because the running is the hard part and the failure mode of not running it is expensive.
Three things follow from that definition, and they are worth stating plainly because they are what a buyer is actually choosing between:
- The provider is accountable for outcomes, not features. If the answer rate does not move, that is the provider's problem to solve, not a support ticket for the client to file.
- The client does not maintain anything. No prompt library, no flow builder, no integration breakage to chase when a CRM field changes.
- The work never finishes. Launch is the start of the engagement. Conversations drift, offers change, callers find new ways to confuse an agent, and someone has to be watching.
That last point is the one most buyers underestimate, and it is the reason this category exists at all. An AI phone agent is not a thing you install. It is a thing you run.
02 Part one
What a managed AI call center is not
Short answer
It is not an AI voice platform, an AI receptionist product, a chatbot with a phone number, or a traditional BPO. A platform sells you tools and leaves you to build and maintain the operation. A receptionist product answers and takes a message. A BPO sells you human seats. A managed AI call center sells you the operation itself, with AI agents doing the talking and a provider accountable for the result.
Four things are currently sold under language close enough to be confusing. They are genuinely different purchases with different price points, different failure modes and different buyers.
An AI voice platform gives you a builder, a model, telephony and an API. It is the right purchase if you have engineers who want to own this, and a plausible one if your call flows are simple and stable. What you are buying is capability. What you are taking on is design, integration, testing, quality assurance, monitoring and every change forever. The per-minute rate looks cheap because it is only the metered part of the cost.
An AI receptionist answers, takes a name and a number, maybe books into a calendar, and texts you. For a small practice with modest volume this is often exactly right and costs a few hundred dollars a month. It is not built to qualify against criteria, hold an objection, read from a loan file or write structured data into a case management system, and stretching it to do so is how people conclude that AI does not work on the phone.
A chatbot with a phone number is a web assistant given voice. It usually shows: it waits for you to finish, it cannot handle interruption, and it restarts when confused. Voice is a different problem from text. Turn-taking, barge-in, silence, accents, background noise and the fact that a caller will change the subject mid-sentence are all specific to the channel.
A traditional BPO sells trained humans by the seat or by the hour. It handles nuance, empathy and genuine complexity better than any current AI, and it scales by hiring, which is slow and expensive and caps out. Most serious operations end up with both, and the interesting question is not which one wins but where the line between them sits.
A useful test: ask who fixes it at two in the morning when a call flow starts failing. On a platform, you do. With a receptionist product, nobody does, because there is nothing to fix. With a BPO, a supervisor does. With a managed AI call center, the provider does, and you find out from the report rather than from a customer.
03 Part one
Managed service or AI voice platform: how to choose
Short answer
Choose a platform if you have in-house engineers who want to own conversation design, integration and quality assurance permanently, and your call flows are simple and stable. Choose a managed service if calls carry real revenue, the flows are complex or regulated, or nobody internally owns the phone as a full-time job. The deciding question is not cost per minute. It is whether you have someone whose job is making the phone perform.
The honest version of this comparison is that both work, for different operations, and the mistake is usually buying a platform and discovering the running costs later.
A platform's advertised price is the metered component. The unadvertised component is the work: mapping the flows, writing and rewriting the conversation, building the integrations, testing against callers who behave badly, listening to recordings, finding out why containment dropped four points last Tuesday, and doing that again next week. That work does not disappear when you buy a platform. It moves to your payroll.
Some organizations should absolutely take it on. If you have an in-house team that already runs conversational systems, or a single high-volume flow that rarely changes, owning it is cheaper and gives you more control. Plenty of large operations do exactly that and are right to.
The cases where managed wins are specific:
- The calls carry real money. Where a single missed conversation is worth hundreds or thousands, the cost of the operation is a rounding error against the cost of it underperforming.
- The flows are complex or regulated. State licensing, prior authorization, consent handling and emergency escalation are not prompt-engineering problems. Getting them wrong is not a bug, it is an incident.
- Nobody owns the phone. If the honest answer to who is accountable for answer rate is a committee, a platform will be bought, half-configured and quietly abandoned.
- You need it working in weeks, not quarters. Building the first competent flow in-house typically takes a team a quarter. That is a reasonable investment; it is just not a fast one.
There is a middle path that works well: run managed to get performance quickly and learn what good looks like on your own traffic, then bring specific flows in-house later with the conversation design and the measurement baseline already proven. That sequence is far easier than the reverse.
04 Part one
Why buying AI voice software so often goes nowhere
Short answer
Most stalled AI voice projects did not fail on model quality. They failed because nobody owned the operation after purchase. The software was bought, a pilot flow was built by whoever had time, the person who built it moved on, no one was watching the calls, and the project quietly stopped. The models got dramatically better over this period. The ownership gap did not close.
There is a pattern to how these projects stall, and it is consistent enough to be worth naming.
A platform is bought on a demo. Somebody technical builds a first flow, usually well. It goes live on a narrow use case and works. Then the person who built it has other work. Nobody is listening to recordings. A CRM field gets renamed and the writeback silently breaks. Callers start hitting an edge the flow never covered. Containment drifts down and nobody notices, because no one owns the number. Six months later the conclusion in the room is that the AI was not good enough.
The AI was usually fine. What was missing was the thing every human call center has and every software purchase omits: a supervisor. Somebody whose job is to listen to calls, find the ones that went wrong, work out why, and change something.
This is also why per-minute pricing is a poor guide to what a deployment costs. A rate of nineteen cents a minute tells you nothing about the quarter of engineering time, the ongoing QA, or the cost of three months of degraded performance nobody caught. Those are the real numbers, and they do not appear on the pricing page.
None of this is an argument that software is the wrong purchase. It is an argument that the software is the cheap part, and that whoever is buying should be honest internally about who is going to run it.
Part two
How it works
What a provider is actually doing between the day you sign and the day the report lands. This part is the operational detail: scope, sequence, capability limits, the handoff, and what integration means when somebody has to maintain it.
05 Part two
What does “managed” actually include?
Short answer
Managed means the provider owns the whole operation: discovery of your call types and systems, conversation design, knowledge engineering, integration with your CRM and scheduling tools, adversarial testing, validation against your standards, launch, live monitoring, call-by-call quality assurance, weekly optimization and outcome reporting. The client supplies system access, a subject-matter expert for a few hours and sign-off on what the agent may say. Everything after that is the provider's work. If containment drops, the provider finds out first. No prompt library or flow builder changes hands.
The word covers a lot of ground and vendors use it loosely. Here is the full list of work that sits between a signature and a phone line that performs, with who does each piece.
Discovery. Not a kickoff call. It is going through your call recordings, your call mix by hour, your qualification criteria and the real state of your CRM rather than the documented one. The provider does the digging. The client supplies access and a few hours from the person who actually answers the phone today, which is rarely the person who signed the contract. Substituting one for the other is the most common reason a first build comes back wrong.
Conversation design. This is where the operation gets decided. Somebody writes how the call opens, what gets asked and in what order, what happens when a caller answers three questions at once, which objections get acknowledged and which get a straight answer, what the agent will refuse to say, and where the call is allowed to end. It should read like a floor script with every exception attached, because that is what it is. Someone who has run a phone room writes it. A prompt is the output of that work rather than a substitute for it.
Knowledge engineering. The unglamorous half of the build. Your pricing, your service territory, your hours, your policies, the product catalog, the four things your team gets asked every day that are written down nowhere. Most of it lives in somebody's head or in a PDF nobody has opened since the last price change. Turning that into something an agent can answer from, with a source behind every fact and a defined behavior when the fact is missing, is a large share of the work. It also decays. When a price moves somebody has to move it, and in a managed engagement that somebody is the provider.
Integration, testing and validation. Covered in detail further down this part. In scope terms: the provider builds the connections, writes the test scripts, runs the adversarial passes and puts real calls in front of the client before anything goes live. The client listens, and says no where it should.
Launch and monitoring. Going live is a configuration change. The monitoring that goes with it is the part worth arguing about. Automated checks watch the things that fail quietly: an API call that started returning errors, a transfer target that stopped answering, a sudden rise in calls ending in the first ten seconds. Those alert on the day they happen. Waiting for them to surface in a monthly report is how three bad months get discovered at once.
Quality assurance. Monitoring is automated. QA is a person listening. The sample has to cross hours, call types and outcomes, because the calls that went wrong are not the ones sitting at the top of the list. The reviewer codes what happened, separates a bad conversation from a broken integration and marks the calls worth changing something over. This is the activity that most distinguishes a managed engagement from a software subscription. It is also the easiest one to quietly not do, so ask how many calls get reviewed each week and by whom.
Weekly optimization. Changes ship weekly rather than quarterly. A question that confuses callers gets rewritten. An escalation rule firing too often gets tightened. A new objection turns up on three calls, an answer gets written for it, and it stops being a problem before it turns up on thirty. None of that arrives in the client's inbox as a task.
Reporting. Answer rate, abandonment, qualification, conversion, containment, repeat contacts, escalation volume with reasons and cost per outcome, measured against the baseline taken before launch. A report that shows only call counts and minutes is a usage statement rather than a performance report.
- The provider owns discovery, conversation design, knowledge engineering, the integration build, test scripting, adversarial testing, launch, monitoring, call review, the weekly changes and the report.
- The client supplies system access and credentials, a subject-matter expert for a few sessions, the escalation policy, the list of things the agent may offer and sign-off before launch.
- Nobody hands over a flow builder login, a prompt library or a maintenance backlog. A buyer who wants those should buy a platform, which is a reasonable purchase and a different one.
One test separates managed from everything else sold under the word. Ask what happens in week six, when nothing is broken and nobody has complained. If somebody is still listening to calls and changing things, the service is managed. If the answer is that support is available, you have bought software with an account manager attached.
06 Part two
How a deployment runs, stage by stage
Short answer
Eight stages: discovery, conversation design, AI build, integration, testing, validation, launch, then optimization, which does not end. A single well-scoped flow typically reaches live traffic in two to four weeks. A multi-flow operation across several systems takes longer, and integration depth is the variable that decides which. The client's part is small but sits on the critical path: system access and credentials, a subject-matter expert for a few sessions, the escalation policy and sign-off before launch. Slippage usually comes from waiting on access rather than from build time.
Eight stages, in the order they happen. The timings below describe a straightforward engagement. Anything involving an older system of record runs longer, and it is almost always the integration that decides.
01 Discovery. Three to five working days. The provider takes your call recordings, your call mix by volume and hour, your qualification criteria, your booking and dispatch rules, your escalation policy and a map of the systems involved. You supply access and a subject-matter expert. The largest single determinant of how long this stage takes is how quickly that person's calendar opens.
02 Conversation design. Roughly a week for one flow. The provider writes the conversation: openings, question order, qualification logic, objection handling, refusals, edge cases, escalation paths and how the call ends. You read it and mark it up. Say no here rather than after launch, because a design change costs an afternoon and a live change costs a week of calls.
03 AI build. A few days, and mostly invisible to you. The design becomes a configured agent: voice, pacing, knowledge base, business rules, the actions it is allowed to take and the ones it is not. Nothing for the client to do here except stay reachable for the questions that fall out of the writing.
04 Integration. The variable. A documented API with a sandbox is a few days of work. An older on-premise system, a homegrown database, or a platform where your own software vendor has to enable API access and approve the connection, can take longer than every other stage combined. What the client supplies here is credentials, a test environment if one exists and someone at their software vendor who answers email. That last one is not a joke. It is the most common cause of a slipped launch date.
05 Testing. Several days of deliberate abuse. Callers who interrupt, callers who mumble, callers on speakerphone in a moving vehicle, callers who change the subject halfway through a sentence, callers who ask for something out of scope, transfers that fail, APIs that time out. The output of this stage that matters is the list of things that broke.
06 Validation. Two to four days, and this one belongs to the client. Calls get reviewed against the standards agreed in discovery, and the client listens to a set of them. Not a highlight reel. Ask for the calls that went badly during testing, and for the ones where the agent refused, transferred or lost the thread. Approval at this gate is what makes launch a decision rather than a hope.
07 Launch. A day, often less. It rarely happens all at once. Most launches start narrow: one number, or after-hours only, or a share of overflow traffic. Forwarding is the usual mechanism rather than porting, because a forward reverses in minutes and a port does not. Monitoring, alerting and the escalation paths go live in the same moment rather than afterwards.
08 Optimize. Weekly, indefinitely. This is the stage a software purchase does not have, because a vendor's work ends at launch and the client's begins. Calls get reviewed. Friction gets found, changes get written and tested, and the next report says whether they worked. Conversations drift. Offers change. Callers keep finding new ways to confuse an agent, and someone has to be watching.
Two honest caveats about the timeline. The first is that two to four weeks describes one well-scoped flow rather than an operation. A firm replacing its whole intake across four flows and three systems should be thinking in months, while still putting the first flow live in weeks. The second is that slippage usually comes from access rather than build: a credential that needs a security review, a sandbox that has to be provisioned, a software vendor who takes eleven days to answer a partner request. A provider who quotes a launch date without asking what systems you run is quoting a build estimate and calling it something else.
07 Part two
What can an AI phone agent actually handle?
Short answer
Reliably: qualification against written criteria, booking and rescheduling, status and lookup calls read from live systems, routine service requests, order taking with full disclosure, triage against written rules, language switching mid-call, and volume that would queue a human floor. Adequately, inside limits you set: objection handling, retention offers the agent has authority to execute, light negotiation. Badly or not at all: advice requiring a license, symptom assessment, genuine distress, novel situations nobody wrote an answer for, card numbers by voice, and anything your own staff would escalate to a supervisor.
Model capability is the wrong frame for this question. What decides whether an AI agent handles a call well is whether the task has a definable right answer somebody has written down. Where it does, the agent is steadier than a human floor: the same disclosure every time, the same verification before the same lookup, at three in the morning as at eleven. Where it does not, it is guessing, and a fluent guess is worse than no answer.
Reliably. These are the call types where an AI agent does the job without a caveat, and together they cover most inbound volume in most operations:
- Qualification against written criteria. Your intake questions in your order, scored against your rules, with the ones that do not qualify declined on the call and coded with a reason.
- Booking, rescheduling and cancelling. Against a live calendar or dispatch board, with the right trade, territory and duration, and a confirmation sent before the caller hangs up.
- Status and lookup calls. Order status, technician ETA, application status, case position, account balance. Read out of the live record rather than a cached article, after verification where it is required.
- Routine service requests. Address changes, refill requests, tracking resends, appointment moves, returns and exchanges inside published policy.
- Order taking with disclosure. Including the parts a commissioned human sometimes skips. Continuity terms, cancellation rights and after-hours fees get said out loud on every call, because they sit in the flow rather than in a training memo.
- Triage against written rules. Classifying an emergency from the symptoms a caller reports, flagging a vulnerable occupant, raising the dispatch priority. Rules, not judgement.
- Language switching. A caller who opens in English and switches to Spanish gets followed on the same number, with no transfer and no callback queue.
- Volume. Concurrency is not a constraint. The four hundredth simultaneous caller gets the conversation the first one got.
Adequately, inside limits you set. This middle band is where the design work earns its fee, because the behavior is only ever as good as the authority written into it.
- Objection handling. An agent will acknowledge an objection properly and answer it, which already beats a good many human floors. It will not improvise a new argument, and it should not be asked to.
- Retention and saves. Real authority inside defined limits works: a cadence change, a skipped shipment, a size exchange, a payment arrangement. Persuasion beyond those limits does not, and attempting it is how a save call turns into a chargeback.
- Light negotiation. Holding a published price is dependable, because there is no discount authority to give away. Anything that needs a room read is a different matter.
Not at all. These belong to a person. A vendor who tells you otherwise is selling you an incident:
- Advice that requires a license. Clinical, legal, financial. The refusal has to be built into the architecture rather than trained in, and it has to survive the caller rephrasing the question.
- Symptom assessment. An agent can follow red-flag rules and escalate. Ranking how serious something is, or reassuring somebody that it is probably fine, sits outside the job.
- Genuine distress. Bereavement, a customer in real trouble, anybody who needs to be heard by a human being. Recognize it fast and transfer.
- Novel situations with no written answer. The case nobody anticipated is precisely the case where a fluent guess does damage.
- Card numbers by voice. A tokenised link keeps the recording and the archive out of scope. Reading a card number to an agent drags both back in.
- Anything your own staff escalate. A useful shortcut: if the request goes to a supervisor when a person takes it, it goes to a supervisor when the agent takes it.
There is a second boundary that has nothing to do with subject matter. Voice is a hostile channel. A caller on a job site with a compressor running, a caller on a hands-free kit at speed, a caller with a heavy accent on a bad line: accuracy degrades, and any vendor claiming otherwise has not listened to enough recordings. Names, street addresses and email addresses are the hardest things to capture by voice, which is why a well-built agent spells back, confirms and sends the detail by text. Callers also interrupt. An agent that cannot be cut off mid-sentence was built from a chat assistant and gets found out on the first difficult call.
One more thing worth saying plainly. The boundary moves. Models have improved enormously and will keep improving, so a list like this one dates. What does not date is the requirement that somebody decides where the line sits for your operation, writes it down and checks it against real calls. An agent does not know what it does not know unless the boundary was authored.
08 Part two
How the handoff to a human works
Short answer
A handoff should be a warm transfer: the agent stays with the caller, dials the target and passes a structured summary, so the person answering already knows the name, the intent, what was verified and what was promised. Escalation runs on written rules rather than model judgement: an explicit request for a person, out-of-scope intent, repetition, frustration, safety language, regulated ground. If nobody picks up, the agent falls back to a second target and then captures a callback commitment. Speed of recognition matters more than containment rate.
A warm transfer means the person picking up already knows who is on the line and why. A cold transfer means the customer starts again. They conclude, correctly, that the two halves of your business do not talk to each other. Everything here exists to stop the first quietly turning into the second.
Mechanically, the agent stays with the caller while it dials. It says what is happening and why, in plain words, without promising that the human will agree to anything. On the published billing dispute call the customer demanded a supervisor thirty-one seconds in. The agent released the duplicate charge first, so the problem was solved either way, then agreed to the transfer immediately rather than attempting one more fix. It said plainly that it would not argue. That sequence is designed rather than emergent.
What triggers an escalation. These are rules evaluated on every call, not a model deciding it feels stuck:
- An explicit request for a person. Honoured on the first ask, with no retention attempt in front of it.
- Out-of-scope intent. The request sits outside what the agent was built and authorised to do. Recognized on the intent, before an attempt.
- Repetition or confusion. The caller has said the same thing twice, or the agent has made no progress across two turns.
- Frustration. Read from the language and from the shape of the call, and treated as a transfer trigger rather than as something to work through.
- Safety language. Gas, smoke, chest pain, a threat. The commercial flow gets abandoned mid-sentence and the escalation path runs instead.
- Regulated ground. Advice requests, disputes with legal weight, anything your compliance posture says a licensed person handles.
- High value by rule. Some operations transfer every qualified case above a threshold whether or not the agent could have finished it, because the human relationship is worth more than the containment.
What crosses with the call. A transfer carrying nothing is a cold transfer with extra steps. The summary should land as structured data in the receiving system, or on the agent's screen, before they speak:
- Who the caller is and how identity was verified.
- The intent, in one line.
- What the agent already read out of your systems, so nothing gets looked up twice.
- Anything the agent committed to, including any figure it quoted.
- The escalation reason as a coded value, so it aggregates on the report.
- A link to the recording and the transcript.
The test is unglamorous. The human says hello and the customer does not have to explain.
When nobody picks up. This is the part that gets skipped in design and then happens at two in the morning. The behavior has to be defined in advance: how long the primary target rings, what the second target is, whether the agent comes back to the caller and takes a callback commitment instead, and what gets written to the record either way. A caller who has just been told a person is coming should not land in a voicemail box. Worse than that is the silent version. A transfer that fails without anybody noticing is the most expensive event in the system, so failed transfers belong on an alert rather than in a monthly summary.
Why recognition speed beats containment rate. Containment is trivially easy to inflate. An agent that refuses to transfer contains more calls, and every one of those extra containments is a customer who wanted a person and did not get one. The number that predicts whether automation helps or hurts an operation is the time between the moment a call went outside scope and the moment a human had context. On a published BPO call the agent identified an out-of-scope request in eleven seconds and transferred in six, which sits inside the abandonment window rather than testing it. A failed attempt followed by a queue is worse than no automation at all.
One design consequence follows, and it reads oddly to anyone used to buying deflection. The agent is built to stop trying. Persistence is a virtue in a salesperson and a defect in tier zero. Ask any vendor to play you a recording of a call their agent could not handle, and time the two numbers yourself: how long until it recognized, how long until the human had the context.
09 Part two
What integration and writeback actually mean
Short answer
Integration means the agent reads from your live systems during the call: the dispatch board, the order record, the matter file, the schedule. Writeback means the outcome lands back in those systems as structured fields, so reports can count it and automations can fire on it. A transcript pasted into a note field is neither. Reading from a documented API is usually quick. Writing into an older on-premise system, or one whose vendor controls API access, is the slowest part of most deployments. Whoever built it maintains it when a field changes.
Integration runs in two directions and they are not equally hard. Reading is the easier one. Writing is where deployments stall, and writing something a report can count is harder again.
Reading means the agent queries your live systems during the conversation and answers from the current state. The technician ETA comes off the dispatch board as it stands that second. The loan file shows which condition is outstanding. The order record shows whether it shipped. On one published healthcare call the agent ran an eligibility check against a patient's new plan mid-conversation and gave the exact specialist copay for that visit type before hanging up. None of that is reachable from a knowledge base. The alternative is an agent reciting what was true the last time somebody exported a spreadsheet.
Writeback means the outcome of the call lands in the system of record as structured fields. A job booked to the correct trade and territory with plan pricing applied. A matter note plus a callback task assigned to the named case manager. A new lead with the pre-qualification fields already complete before the loan officer picks up. An interaction closed with a contained disposition against the right program. Fields, with values, in the places your reports and your automations already look.
Which brings up the thing to watch for in a demo. A transcript dropped into a note field is not integration. Prose in a note cannot be filtered, counted, routed or triggered on. Nothing downstream fires. Somebody still has to open the record and read it, which is the work you were trying to remove in the first place. It proves the call happened, and you already knew that. The question to put to a vendor is concrete: show me the record this call produced, in my system, with the fields populated, and then show me the report that counts it.
What is quick. A documented REST API with a sandbox, token-based authentication and webhooks. Most modern CRMs, scheduling tools, help desks and field service platforms sit in this category, and a read-and-write connection is days of work rather than weeks.
What is slow. Older on-premise systems reached through a VPN and a change window. Homegrown databases with no API at all. Platforms where the API exists but your own software vendor has to enable it, approve the connection or admit the provider to a partner program, which turns an engineering task into a calendar problem. Anything driven by screen automation, which works and then breaks the first time the vendor moves a button. Write access behind strict validation, where every required field and every picklist value gets discovered by failing.
Two practical notes on sequencing. Read access almost always lands before write access, so it is normal to launch with the agent reading live while writing to a staging record for the first week, with the field mapping proved against real calls before it touches production data. And a partial integration beats a delayed launch. If booking writes cleanly but the disposition fields are still being argued about with your software vendor, go live on the booking and finish the rest in week three.
Who maintains it. This is where the managed part earns its fee and where the platform purchase quietly falls over. Fields change. Somebody adds a required field, renames a picklist value, tightens a validation rule, migrates to a new instance, or switches on a module that changes what a status means. The write starts failing. Worse, it keeps succeeding into a field nobody reads any more. In a managed engagement the monitoring catches it, the provider fixes the mapping, and the client finds out from the report rather than from a customer. Somebody has to own that. Where nobody does, the integration degrades over months and the conclusion in the room is that the AI stopped working.
Part three
Risk
The part vendors skip. What stops an agent saying something that creates liability, what compliance looks like in the industries where it bites, and the specific ways these deployments go wrong.
10 Part three
What are guardrails, and why do they belong in the architecture?
Short answer
A guardrail is a boundary on what a phone agent is permitted to do, enforced by the conversation architecture rather than requested of the model in a prompt. Enforced in the architecture it is a control, because the path the boundary forbids does not exist. Written into a prompt it is a hope that the model behaves under pressure. Guardrails come in five kinds: refusals, escalation triggers, required disclosures, data handling rules and authority limits. Test one by attacking it on a live line, then asking the vendor for a recording of the agent refusing.
A guardrail is a rule about what the agent is allowed to do, and the whole question is where that rule lives. Two deployments can carry the identical sentence in their documentation. In one, the boundary is a line near the bottom of a prompt asking the model politely to behave. In the other, the path the boundary forbids does not exist: no tool to call, no state the agent can enter, and a classifier that fires before the model gets its chance to be helpful. Same policy on paper. Very different system.
The difference shows up under pressure, which is where it counts. A model instructed to decline medical advice will decline on the first ask. It will usually decline on the second. On the fourth, from a caller who has rephrased the question as a hypothetical about their sister and mentioned that they are a nurse, the instruction is competing against everything else in the context and against the model's own pull toward being useful. Instructions degrade. Architecture does not negotiate.
So never ask a vendor whether they have guardrails. Everyone says yes. Ask what enforces each one, and what would have to be true for it to fail. An honest answer names a mechanism: a classifier running before generation, a tool the agent was never given, a routing rule that outranks the rest of the flow. A vague answer about careful prompting is also an answer.
They come in five kinds, enforced by different mechanisms and failing in different ways.
- Refusals. Things the agent will not do at any point in any call: give medical, legal or financial advice, quote a repair price before a technician has seen the job, invent a discount to save a sale, tell a caller what their claim might be worth. A refusal is a design fact rather than a tendency. No phrasing of the question produces the answer, because the answer is not reachable from inside the conversation.
- Escalation triggers. Conditions that stop whatever the agent was doing and move the call somewhere else. Red-flag clinical language. A safety signal. A caller who says the word lawyer. Three failed attempts at one intent. What matters is rank: the trigger outranks the commercial objective, so a booking already in progress gets abandoned rather than finished.
- Disclosures. Wording delivered regardless of how the conversation goes. The agent identifying itself as automated, the recording notice, continuity terms stated before an order is taken, the approved language on whether any relationship has been formed. Fired by the flow, in the client's words, never composed by the model in the moment.
- Data handling. Where sensitive input is permitted to go. Card capture moved to a separate path, tokenized IVR or a texted secure link, so the digits never reach the part of the call the AI can hear. Recording suppressed across those segments. Protected health information treated as such from the first second, rather than from the moment somebody says a diagnosis out loud.
- Authority limits. The size of decision the agent may make alone. Discount ceilings, goodwill limits, refund rules, which slots it can promise, which accounts it must hand straight to a person. Live capacity read from a dispatch board is an authority limit as much as an integration, because an agent reading real availability cannot offer a slot that is not there.
Every one of those five can be shown on a call, which is the whole test. So attack it. Vendors will put you on a live line; use it to break the boundary rather than to admire the voice. Ask the question the agent is supposed to refuse. Ask it again in different words. Claim an emergency. Say you only need a rough idea, off the record. Ask for a discount it has no authority to give, then get annoyed when it says no. Watch whether the refusal keeps its shape on the fourth attempt, and whether the agent stays pleasant while holding it.
Then ask for recordings, the harder request. A library of successful calls proves a vendor can demo. A published call where the agent gives up the sale is evidence of a different kind. Ours exists partly for that: an agent hearing a passing remark about a smell in a utility room, abandoning the booking mid-flow, telling the caller to leave the house without touching switches. A caller asking which product to choose and being told plainly that the agent cannot answer, with the refusal logged and a transfer to someone licensed in that caller's state. A symptom disclosure that ends a scheduling flow and then declines to accept the caller's own downgrade of the symptom. A contact recognized as out of scope eleven seconds in and transferred in six, with a summary already on the specialist's screen.
One question most buyers forget: how do you know the guardrail still holds next month? They are not static. Model versions change. Somebody edits the knowledge base on a Thursday. A new intent gets added because a client asked nicely. The boundary tested in week one is now running against a conversation that moved underneath it, so the tests have to run again, on a schedule, against real traffic.
A boundary on the boundaries, finally. All of that is engineering, and engineering is the part a provider can be held to. Whether a given configuration satisfies a given regulator is a question for your own counsel, on your own facts. The division that works: the provider describes the behavior and demonstrates it on a recorded call, and your compliance function decides whether it is enough.
11 Part three
What compliance looks like, industry by industry
Short answer
Compliance on the phone is made of behaviors you can hear, and it is a property of the deployment rather than of the model. Healthcare: the agent captures the reason for the visit and refuses to assess it. Financial services: it declines advice and transfers to someone licensed in the caller's state. Legal: it takes facts without evaluating them. Direct response and retail: card data never reaches the part of the call the AI hears. Home services: it discloses the diagnostic fee and quotes nothing else. Outsourcing: each client program is isolated. Your own counsel signs off the posture.
Buyers usually arrive with one question: is it compliant? In that shape the question has no answer. Compliance is a property of a deployment rather than of a model, which means it lives in specific behaviors on specific calls: what the agent refuses, what it escalates by rule, where sensitive data is allowed to go, what gets said whether or not the conversation invites it. Below is what that looks like in six places where the boundary has teeth.
One thing to be clear about before the specifics. None of this is legal advice, and no configuration makes anyone compliant. Rules in every one of these sectors vary by product, by state and by regulator, and they move. What a provider can do is describe the controls, show them working on a recorded call and document them for your file. What happens after that belongs to your own counsel and your compliance function, reading that description against your own facts and signing it off or sending it back. A vendor who tells you their product is compliant has told you something useful about the vendor.
Healthcare. The thing that bites is the moment an agent responds to a symptom. Characterise it, reassure about it, rank its urgency, guess at what it might be, and a phone line has practiced medicine. So the agent captures the reason for the visit in the patient's own words and stops there. Red-flag language abandons the scheduling flow ahead of any booking logic and moves the caller onto the emergency instruction and the on-call path; anything needing judgement goes to a nurse line or a clinician with the context already captured. Recordings, transcripts and structured data are treated as protected health information end to end, under a business associate agreement signed with the practice, and the agent asks for what the workflow requires and nothing past it. Clinical and compliance leads approve the scripts, the escalation triggers and the disclosures during QA, before a patient hears any of it. One warning specific to this sector: read a HIPAA badge on a vendor's website as a reason to ask more questions rather than fewer.
Financial services. The line that bites runs between information and advice, and it sits closer than people expect. A question about the current rate is on one side. A question about whether to refinance is on the other. So the agent makes no recommendations at all: not which product to choose, not whether to refinance, not when to buy or sell, not what rates will do next. Asked directly it says so, routes to a representative licensed in the caller's state, and both the request and the refusal go onto the contact record. Rates, terms and figures come only from approved sources, so there is nothing for the agent to interpolate. Disclosures, the recording notice and the AI identification are scripted and delivered every call rather than left to the flow of the conversation. Consent capture and do-not-call handling follow the client's policy and are logged on the record. Everything is recorded, transcribed and retained to the client's schedule, which is what makes supervision possible at all.
Legal. Unauthorised practice is the first thing a managing partner raises, and implying a relationship that does not exist is the second. The agent gathers what happened and does not evaluate it: no view on whether there is a case, no estimate of what it might be worth, no opinion on whether to accept an offer, no steer on which firm to call. Asked, it says an attorney will make that call and routes accordingly. The firm's approved language on whether a relationship has been formed is delivered on every call, and the agent never suggests the matter has been taken until the intake process says it has. Automated-assistant disclosure is scripted and always on. A caller who is already represented, or who matches a conflict rule the firm defines, is identified early and handled the way the firm specifies rather than screened deeper. Calls are recorded and transcribed, with consent handling configured by jurisdiction.
Direct-to-consumer and retail. Two exposures dominate. Card data is the first, and the shortest route to losing a direct-response brand is to become their payment problem, so payment moves to a separate capture path, tokenized IVR or a texted secure link, and the AI resumes afterwards without ever receiving the digits. Recording and transcription suppress those segments, which keeps the archive clean. Claims are the second. Product claims, comparisons and substantiation come out of the brand's own knowledge base, and the agent does not generate marketing language on the fly, which is the specific nightmare a DR operator has in mind when they ask what the AI might say. Continuity terms are stated in approved wording before the order is taken. The agent has no discount authority, so it cannot invent a price to rescue a call. On the retail side the same architecture carries returns, price adjustments and goodwill limits as rules the brand sets, and the agent escalates on frustration or repetition rather than looping, because looping is the failure customers remember.
Home services. Rudeness is not the exposure here. A wrong truck is, and so is a missed emergency, and so is a price quoted on the phone that a technician then has to walk back on the doorstep. Safety triggers, gas odor and flooding and no heat in freezing conditions, abandon the scheduling flow immediately and hand the call to the emergency instruction and the on-call path. The agent discloses the diagnostic or trip fee and quotes nothing else, because repair pricing happens after a diagnosis, by somebody who has seen the equipment. Address is checked against territory before anything is booked. Job type maps to the right trade, license and skill level. Availability is read live from the dispatch board, so a slot that does not exist cannot be promised.
Outsourcing, BPO and agency programs. Neither of the two exposures here is regulatory in the usual sense. The first is segregation: each client program is isolated, knowledge and recordings and transcripts and structured data, with nothing crossing between programs. The second is channel conflict, which is the real fear sitting behind every technical question a BPO owner asks. The agent answers as the client's brand, the provider contracts with the partner, and the provider does not approach, quote or contract with the partner's clients. That last commitment belongs in the agreement, so ask to read the clause rather than the web page. Alongside it, SLA metrics are produced in the shape the client contract already specifies, and escalations arrive carrying structured context, so the handoff does not burn the SLA it was meant to protect.
The pattern across all six is the same. The compliance story is made of things the agent will not do, and things it does whether or not the conversation asks for them. Both are observable. A provider who can only describe their guardrails in the abstract, with no recording and no mechanism to point at, has probably written them down rather than built them.
12 Part three
How these deployments fail
Short answer
These deployments fail operationally far more often than technically. The recurring patterns: nobody owns the engagement after launch, containment gets reported without repeat contacts so a call deflected into a callback queue counts as a win, the agent keeps trying instead of handing over, a renamed CRM field breaks writeback silently, scope creeps into calls the agent should refuse, escalation paths are configured but never tested, and launch is treated as the finish line. Every one of these is detectable in the first month by a person listening to calls.
Everything below is a failure we design against, which is another way of saying it is a list of things that have gone wrong in phone rooms for as long as there have been phone rooms. Very little of it is about AI. That is rather the point.
Nobody owns it after launch. The big one, and it does not disappear because the engagement is managed. The provider owns the operation. Somebody on the client side still owns the decisions the provider cannot make: a policy that changed, an offer that moved, who the escalation goes to now that the person it used to go to has left. When that owner is a committee, the weekly report stops being read at about week six, nobody pushes back on anything in it, and the engagement quietly becomes maintenance.
Containment reported without repeat contacts. The most common way a deployment looks healthy and is not. A call deflected into a callback queue counts as contained on most dashboards. The customer rings back tomorrow. That is a fresh contact, so containment holds, the number on the report goes up, and the customer has had a worse experience than the hold music they would otherwise have waited through. The fix is arithmetic rather than philosophy: put containment and repeat-contact rate on the same line of the same report, and count a repeat inside three days against the original contact. If a provider resists that, you have learned something.
The agent that will not stop trying. A model tuned to be helpful keeps attempting, and a containment target makes it worse, because you have now paid somebody to keep the caller away from a human. Looping is the behavior customers hate most. It is worse than a queue, because a queue at least ends. What good looks like is an attempt budget, an escalation that fires on frustration or repetition, a handoff measured in seconds, and a supervisor who reads the transcripts of the calls that escalated. One of our published calls has the agent recognizing an out-of-scope request eleven seconds in and completing the transfer in six, with a summary attached. The number to ask a vendor for is time to escalate. Almost nobody volunteers it.
Integration drift. The quiet one, and the one that costs most before anybody notices. A CRM field gets renamed. An API version is deprecated. Somebody adds a required field to the booking form. The dispatch board changes how it exposes availability. None of that makes a noise the caller can hear: the call still sounds excellent, and the record never lands. The defense is unglamorous. Treat writeback success as a monitored signal with an alert on it, and reconcile calls handled against records created every week. Two numbers, one comparison. Skip it and you find out from a customer.
Scope creep into calls it should refuse. This one starts as success. Booking is going well, so somebody adds billing questions. Then cancellations. Then the thing that genuinely needs judgement. Each addition is reasonable on its own, and the boundary that was designed and adversarially tested before launch is now carrying traffic nobody designed for. The refusals are the asset. Widening them by increments is how an operation loses them, which is why a change of scope belongs back in the design and test cycle rather than in a knowledge base edit on a Thursday afternoon.
Escalation paths nobody tested. Configured is not the same as tested. The transfer target rings a desk that moved in the spring. The after-hours number reaches an on-call rota that changed and was never updated here. The nurse line is queued at exactly the hour an escalation is most likely. Test these the way a building tests a fire alarm: on a schedule, with a human actually picking up, logged. They belong in the weekly review, not in a launch checklist where they were ticked once and forgotten.
Launch treated as the finish line. The umbrella failure, and the one every other item on this list is a symptom of. The conversation was designed against the calls you had in the spring. Callers have found new ways to phrase things since. An offer changed and the people who write the conversation heard about it three weeks later. Performance does not fall off a cliff through any of that, it erodes, which is exactly why nobody notices until a quarter has gone. The counter is dull and it works: a weekly review of sampled calls, deliberately including the ones that went badly, ending in a change.
Two honest observations to close on. The first is that none of these is a model problem, so a better model fixes none of them. The second is that all of them are visible inside the first month to a person who listens to calls and reads the writeback log. That person is the product. Everything else in this guide is downstream of whether somebody is doing that job, and whose payroll they sit on is the decision the rest of the guide is really about.
Part four
The money
What it costs, how to build a business case that survives a CFO, and which numbers to report against. The arithmetic here is deliberately conservative, because the aggressive version is what gets vendors thrown out of the room.
13 Part four
What does a managed AI call center cost?
Short answer
Four pricing models compete for the same budget. Per-minute voice platforms meter connected minutes and leave the build and the running to you. Per-seat BPO contracts price supervised human capacity. Flat-rate AI receptionists answer a small volume of simple calls for a few hundred dollars a month. A managed AI call center charges a monthly fee for the whole operation plus a one-time implementation. ScaileAI quotes usage at roughly $0.32 to $0.40 a minute and scopes the implementation and monthly minimum to the operation. Growth is built for 2,000 to 2,500 calls a month, Pro for 5,000 to 6,000, Enterprise above that.
Four prices are quoted in this market and none of them measure the same thing. A minute is an input. A seat is a person. A flat monthly fee on a receptionist product is a cap on volume. A managed fee is the price of an operation. Lining those four up in a column and taking the smallest number is how most bad AI voice decisions get made.
| Pricing model | Headline price | One-time cost | Who runs it |
|---|---|---|---|
| Per-minute voice platform | $0.08 to $0.35 a minute | None | Your team |
| AI receptionist product | $100 to $600 a month | None | Nobody |
| Per-seat BPO | Per seat or per hour | Recruitment and training | The BPO |
| Managed AI call center | Monthly fee plus usage | Implementation | The provider |
The two price ranges above are indicative rather than a quote from anyone, and they move; they are the same ranges published on the pricing explainer. The column that decides your real cost is the last one.
Per-minute platforms. You pay for connected minutes, typically somewhere between eight and thirty-five cents depending on voice quality, telephony and concurrency. The number is small because it prices the cheapest input. Everything that makes a phone agent work on your traffic sits outside it. Flow mapping, conversation design, knowledge structuring, integration, adversarial testing, and a person listening to calls every week after launch. Costed at a loaded internal rate, that work routinely runs higher in year one than the minutes do. The rate card never says so.
AI receptionist products. A flat fee, usually a hundred to six hundred dollars a month, for a capped number of calls or minutes. For a two-person office replacing voicemail this is often exactly the right purchase and nothing more needs saying about it. It stops being the right purchase the moment a call has to be qualified against criteria, or an outcome has to be written back into a system, or volume arrives in bursts. Stretching one into that job accounts for a good share of the disappointment with AI voice.
Per-seat BPO. You buy supervised human capacity, billed by seat or by hour, with a utilization assumption underneath the rate. The price is legible and the quality is contractable, which finance teams like. What it will not do is flex sideways at nine on a Friday night, and it scales by hiring, which is slow and expensive and eventually caps out. The BPO economics published elsewhere on this site put an illustrative $6.40 loaded cost per human contact against $1.15 on an AI tier-zero, which blends to $3.51 for a program containing 55% of its contacts. Those are example figures. The shape is what holds.
Managed engagements. A monthly fee covering design, build, integration, monitoring, weekly tuning and reporting, plus usage, plus a one-time implementation. The fee is larger than a platform's headline. It is smaller than the same operation run properly in-house. What the money actually buys is that nobody on your payroll has to own it.
Inside a managed engagement, four things move the price, and only one of them is volume.
- Volume. Minutes are the metered part, and the part buyers think about first. They are rarely the part that decides the quote.
- Integration depth. Reading a lead record out of a modern CRM is an afternoon. Writing a structured outcome into an older on-premise system, through middleware somebody left behind in 2014, is a project. This is the single largest source of variance in implementation cost.
- Number of distinct flows. One call type, well scoped, is a launch. Six call types with different qualification logic and different escalation rules are six launches sharing a phone number.
- Regulated review. Where disclosures, consent capture or a compliance sign-off sit in the path, conversation design goes through a review cycle that has nothing to do with engineering and adds weeks.
Which brings up the per-minute rate, the most quoted and least useful figure in this market. It prices talk time. It says nothing about who designed the conversation, who maintains it when your offer changes, who notices that containment slipped four points last Tuesday, or what a quiet quarter of degraded performance costs. Those are the expensive items and none of them appear on a rate card.
There is a second problem with the rate. Minutes are the one component of this system that gets cheaper every year. The operating work is the one component that does not. Buying on the cheapest input optimizes the part of the bill that is already falling on its own.
ScaileAI publishes a usage rate and scopes everything else. Usage runs roughly $0.32 to $0.40 a minute depending on scope and volume. Implementation and the monthly minimum are quoted against call volume and integrations, because an intake flow with two integrations and a sales floor with nine are not the same build. Growth is built for 2,000 to 2,500 calls a month and covers one number, campaign or department. Pro is built for 5,000 to 6,000 and covers multiple workflows across departments, locations or campaigns. Enterprise is scoped to the operation. The rate is published as a sanity check, not as the basis for choosing: a minute from a commodity supplier and a minute from a managed operation cost about the same and contain completely different amounts of work, so the question worth asking is what sits inside it. The plans carry the current figures.
Every engagement starts with a paid pilot on one call flow, which is the cheapest honest way to find out what your traffic actually does. The comparison worth running against any vendor is the whole operation over twelve months: managed or platform fee, usage, implementation, integration build, plus the internal hours to run the thing. Price it that way and the gap between cheap and managed narrows a long way. On complex flows it inverts. Section 14 turns that total into a number finance can argue with.
14 Part four
How to build the business case
Short answer
Build it as a leakage model. Take the calls you already receive, the share you answer today, and the gap between them. Apply recovery only to the calls currently missed, hold your qualification and conversion rates exactly where they are, multiply by contribution value rather than revenue, then subtract the provider's fee. That net figure is the business case. Assume no headcount reduction, round calls and customers to whole units before deriving money, and state every input so a CFO can attack the assumptions rather than the arithmetic.
Most AI voice business cases fall apart in the same place. They model a benefit the buyer cannot verify against a cost the buyer can. A leakage model avoids that by starting from a number the buyer already owns: the calls that arrive and reach nobody.
The calculator on the ScaileAI homepage runs this model and prints its arithmetic on every line. Below is the same model written out, with the reason for each conservative choice, because the choices are the part finance will test.
- Start with offered calls. Every inbound call that arrived last month, answered or not. This is the denominator for everything downstream, and the number most operations cannot produce on request. Take it from the carrier or the phone system, never from the CRM.
- Apply your current answer rate. Answered equals offered times the share you answer today. The difference is your missed volume, and that volume is the entire opportunity the model is permitted to price.
- Recover only the missed calls. Provider performance is applied to missed calls alone. Calls your team already answers are left exactly as they are, with no modeled improvement in how they go. This is the choice that keeps the model defensible.
- Hold qualification at parity. Of the recovered calls, the share that turn out to be real prospects is set to your existing qualification rate. No lift is modeled.
- Hold conversion at parity. Of those, the share that become customers uses your existing conversion rate. Again, no lift.
- Multiply by contribution value. Margin on one customer, not the invoice total. This is the input vendors inflate most often and the first one finance will check.
- Subtract the fee. Recovered minutes are counted at your average call length, measured against the allowance the plan includes, and the monthly fee comes off before anything is called a return.
One detail sounds pedantic and is load-bearing. Whole calls and whole customers are settled before any money is derived. Two-thirds of a customer multiplied by a contribution value produces a total that does not tie to the line above it, and a CFO who finds one broken row stops reading.
The model books no labor saving whatsoever. Not a seat, not an hour. That is a deliberate understatement, because headcount reduction is the claim that gets an AI proposal killed in the room, and the claim least likely to survive an operations manager who knows what those people actually do. If released capacity turns into cost saved later, that is upside the business case never needed.
A worked example, using the illustrative figures already published on the ScaileAI pricing explainer. An operation takes 4,200 inbound calls a month and answers 82% of them, so 3,444 are answered and 756 reach nobody. Average call length is 4 minutes. Of answered calls, 50% are real prospects and 30% of those become customers, which is a 15% rate end to end. Contribution value is $2,400 per customer. Every figure in this example is illustrative.
Run the chain with the recovered answer rate set to 100%. All 756 missed calls are recovered, which is 3,024 minutes and sits comfortably inside the 10,000 the Growth plan includes. Of those, 378 qualify, and 113 of them convert. At $2,400 of contribution each, that is $271,200 a month of illustrative value that currently never has a conversation. Those 3,024 minutes cost about $1,210 at the top of the published range, so the net is $269,990 before implementation.
Now attack it, which is what finance will do. Recovery is rarely 100%, because some callers hang up in the first two seconds regardless of who answers. After-hours and overflow traffic is often colder than the calls a team takes at eleven in the morning, so parity can flatter the result. And the illustrative $2,400 is the right input only if it is margin. If it is the order value, the case is overstated by your entire cost of delivery.
Re-run the same illustrative example with recovery at 90%, conversion cut from 30% to 20% because the recovered traffic is colder, and contribution corrected from a $2,400 order to $840 of margin. Now 680 calls are recovered, 340 qualify and 68 convert, worth $57,120 a month. Those 2,720 minutes cost about $1,088, so the net is $56,032. That is the version to take into the room, because it is the version that survives being argued with.
The cost line has to move with volume, and on a usage rate it does by construction: double the recovered minutes and the usage cost doubles with them. That is the correct order, and it is the thing most vendor models get wrong. A spreadsheet that nets a rising volume against a fixed monthly fee flatters the result every time the volume input climbs, which is exactly when somebody is most likely to be looking at it.
The same inflations turn up in vendor models often enough to be worth naming before somebody else does.
- Recovered calls modeled at a lift. The claim that the AI will also convert better than your team. It might, on speed alone. Still not something to book in advance, and parity costs nothing to defend.
- Headcount elimination in the savings line. The quickest way to turn a finance conversation into an HR conversation, and the claim most likely to be wrong. Build the case on recovered revenue and let capacity be the bonus.
- Gross revenue standing in for contribution. An illustrative $2,400 order at 35% margin is $840 of value. The distance between those two numbers is the distance between a case that gets signed and a case that gets audited.
One more line costs nothing and buys real credibility: say where each input came from and over what period. Offered volume from the carrier, answer rate from the phone system, conversion from the CRM, contribution from finance. A model whose inputs each trace to a system of record is arguing about assumptions. A model whose inputs came from the vendor is arguing about the vendor. Section 15 covers what to report once it runs.
15 Part four
The metrics that actually matter
Short answer
Report offered calls, answer rate, average speed of answer and abandonment for capacity; containment always beside repeat-contact rate; qualification rate, conversion and cost per outcome for commercial performance. Treat average handling time as a capacity input rather than a target, because the quickest way to cut it is to resolve less. Add speed to lead for anything revenue-bearing. Break every one of those by hour and by call type, since a blended monthly average hides the peaks and the after-hours window, which is where most of the damage in an inbound operation happens.
A monthly report with one number on it is not a report. The set below is what an operator would ask for, and every line has something specific it conceals when it is quoted blended across all hours and all call types.
Offered. Every call that arrived, answered or not. It is the denominator for everything else, and it is the figure most operations quietly cannot produce. Calls that ring out, hit a busy tone or abandon inside an IVR often never reach the report at all, which means answer rate is being calculated against a denominator somebody already cleaned. Pull offered from the carrier and reconcile it against whatever the dashboard says.
Answer rate. Answered divided by offered. Quoted as a monthly blend it is close to meaningless, because the hours where it collapses are the hours with the most calls in them. Ask for it by hour of day and by day of week. An operation sitting comfortably in the high eighties across a month can be answering barely half its calls between noon and two, and the lunchtime calls are not the cheap ones.
Average speed of answer. How long a caller waits before somebody picks up. The average is the wrong statistic and everybody keeps using it anyway. Ask for the distribution, or at minimum the ninetieth percentile, because the callers who wait longest are the ones who leave and the ones who bring it up later. A mean dragged down by a thousand instant answers says nothing about the two hundred people who waited a minute and a half.
Abandonment. The share of callers who hang up before reaching anyone. This is the most honest number in the set, because it is customers voting, and it is the one most often missing from a vendor dashboard. Read it against speed of answer. If abandonment is high while speed of answer looks fine, the abandoned callers are being lost before they ever enter the queue the dashboard measures.
Containment, reported next to repeat-contact rate. Containment on its own is the easiest metric in this field to game. Push a caller into a callback queue, or close the call with a promise of a text message, and most dashboards will score it as contained. Put it beside the share of those callers who made contact again within seventy-two hours and it becomes a real number. A containment figure quoted without a repeat-contact rate next to it should be treated as unreported.
Qualification rate. Of the conversations handled, the share that turn out to be genuine prospects. This is more useful as a check on the traffic than on the agent. A qualification rate that drops while everything else holds steady is usually a marketing problem arriving down the phone line, and reading it as an agent problem sends the tuning work in the wrong direction for a month.
Conversion. The share of qualified conversations that become customers, cases, bookings or whatever your business counts. Blended, it hides the difference between the calls a human took and the calls the AI took, which is precisely the comparison the engagement exists to produce. Ask for it split by who handled the call, and by hour, so the after-hours cohort can be seen on its own.
Average handling time. Talk plus hold plus after-call work. This is a capacity number. Make it a target and the quickest route to hitting it is to resolve less and get the caller off the line, which surfaces a fortnight later as repeat contacts and a worse first-contact resolution figure. Use it to size the operation and to price minutes. Keep it off the scorecard.
Speed to lead. The gap between a prospect making contact and a business responding. For anything revenue-bearing this outranks most of the list, and it is mostly a function of who is available at the moment the call lands rather than how good anyone is at selling. Measure it end to end, from the inbound ring to somebody actually engaging, rather than from the moment a record appeared in the CRM.
Cost per outcome. The number an executive can use. Cost per qualified conversation, per booked appointment, per retained case, per converted customer. Minutes are an input and nobody has a minutes target. Building the report around cost per outcome also changes the vendor conversation, because a rate card invites a procurement negotiation while a cost per qualified opportunity invites a comparison against what that opportunity costs you today.
One rule covers most of this. Any of these numbers quoted as a single monthly figure is an average of two different operations: the one that runs while there is capacity, and the one that runs when there is none. The second is where the money leaks. Ask for every metric split by hour, by day of week and by call type, and insist the baseline is measured on your own traffic before anything changes, so the comparison afterwards is arithmetic rather than opinion.
Part five
Buying it
How to tell a real operator from a demo, how to structure a pilot that proves something either way, and where to start given what you run.
16 Part five
How to evaluate an AI call center vendor
Short answer
Evaluate on evidence rather than on a demo. Ask to hear a call that went wrong, and a call where the agent refused to do something. Ask who owns the system after launch, by name. Ask how containment is measured and whether repeat-contact rate is reported beside it. Ask what breaks when a CRM field gets renamed. Ask for the total cost rather than the per-minute rate, and for the baseline methodology in writing. Then ask what flows they will not take on. A provider with no answer to that last one has not run into the edges yet.
Every vendor can show you a good call. It is the one they have run four hundred times, on a scenario they chose. What it proves is that the voice is decent. Your traffic is a different animal, and the questions below are the ones that tell you how a provider behaves on it. They are uncomfortable on purpose. Put them to ScaileAI as well.
Ask to hear a call that went wrong. Not a hard call that ended well. A call where the agent misheard something, went round a loop, or lost a sale it should have had. Anybody who has run a phone room for six months has a folder of these. The useful part of the answer is what they changed after the one they play you. A vendor who says every recording is confidential is telling you how little of this is live.
ScaileAI publishes whole calls for that reason, including the unflattering ones. The call library has a booking the agent abandoned mid-sentence because the caller mentioned a smell in the utility room, a caller who disqualified themselves on price and was logged as lost with the reason coded, and a contact that went out of scope eleven seconds in. Those are the calls that show what happens when the script runs out. Hold every vendor to the same standard.
Ask what the agent refuses to do, and to hear a refusal. A good answer is specific: it will not give advice, it will not confirm eligibility it cannot verify, it will not quote outside the published range. Then they play you the recording. A vendor who answers with a paragraph about safety training has a prompt rather than an architecture. The advice refusal in the library is what the artefact looks like. The caller pushes twice, gets declined both times, then goes to someone licensed in the right state.
Ask who owns it after launch. You want a name and a cadence. Who listens to calls, how many, how often, and what happens to what they find. If the answer is that you get access to a dashboard, you have bought software with a service wrapper on it. Then ask what happens when that person leaves. Most stalled deployments trace back to one competent individual moving on and nobody noticing for a quarter.
Ask how containment is measured, and whether repeat contacts sit next to it. Containment is the easiest number in this industry to flatter. A call deflected into a callback queue counts as contained on most dashboards, and the customer rings again the next day. If those two figures are not on the same page of the report, the containment number is decoration. A vendor who volunteers repeat-contact rate before you ask has run a real floor.
Ask what happens when an integration field changes. Somebody renames a picklist value in your CRM in March. What breaks, how quickly does anyone know, and who fixes it. The answer you want involves monitoring on the writeback and an alert that fires at the vendor rather than at you. This is the most common quiet failure in the category, because nothing stops working loudly. The calls keep happening. The data stops arriving.
Ask what the whole thing costs. A per-minute rate prices the cheapest input in isolation. Ask for the monthly fee, the included volume, the implementation charge, what an overage minute costs, what a second call flow costs later, and how many of your own hours the vendor expects to consume. ScaileAI publishes a usage rate, roughly $0.32 to $0.40 a minute, and scopes the implementation and monthly minimum to the operation. Whoever you buy from, get the figure that lands on the invoice, not the one on the slide.
Ask for the baseline methodology in writing. Not what they will report afterwards. How they establish what your phone does today, before anything changes. Calls offered against calls answered, abandonment by hour rather than blended, the conversion rate on answered calls, and where each of those figures comes from. If the baseline gets measured after go-live, or comes out of the vendor's own platform, the comparison belongs to them. Every ScaileAI engagement starts with a paid pilot on one call flow measured against the client's own baseline, and this paragraph is why.
Ask what they will not take on. A provider who has done this at scale can name the flows they decline and say why. Clinical triage. Anything where being wrong is a regulatory event rather than a bad call. A vendor who says yes to everything in the room is either new or planning to find the edges on your line. The answer also marks the honest boundary of the technology, which is worth more than a capability slide.
- Only good calls exist. Every recording they can show you ends well. Either they are new, or sales curates the library.
- Containment quoted on its own. One number, with no repeat-contact rate anywhere near it.
- The pilot is free. Free pilots are a sales activity and get staffed like one.
- Nobody is named. A support queue and a dashboard, with no individual whose job is your answer rate.
Run the whole list at us. If the answers here are worse than a competitor's, buy the competitor. These questions have no vendor-shaped answer, which is the only reason they are worth publishing.
17 Part five
How to run a pilot that proves something
Short answer
Scope the pilot to a single call flow. Record your baseline before anything changes, from your telephony and CRM data rather than from anyone's recollection. Write the success measure down in one sentence with a figure in it, and agree the failure condition in the same document. Run it four to eight weeks, long enough to cover a full weekly cycle. Instrument the escalations and the repeat contacts, not only the conversions. A pilot with no agreed way to fail has proved nothing, whatever the results look like at the end of it.
Most AI pilots are built so that nobody ever has to say it did not work. The scope drifts, the success measure arrives after the results do, and the write-up records that the team found it promising. That is a demo with a longer run time. A pilot is a measurement, and a measurement needs its number decided before you start.
Scope it to one call flow. One. Not one department and not one location: one flow, with a defined entry point and a defined set of endings. After-hours inbound for a single trade. Order status for one product line. Reschedules for one clinic. The pull is always to widen it so the exercise feels worth the effort, and widening is what makes the result unreadable, because afterwards you cannot say which change moved which number.
After-hours is usually the cleanest start. The baseline is unarguable. You already know how many calls reach nobody between six at night and eight in the morning, and you know what happens to all of them now. Nobody is displaced, so the work carries no internal politics. The floor is low enough that even a mediocre result beats the status quo, which turns the review meeting into a conversation about how much better rather than whether it works at all. Peak overflow is the second-cleanest for the same reasons.
Record the baseline before anything changes. This is the step that gets skipped, and skipping it is what turns the review into an argument about whether things feel better. Pull calls offered against calls answered, split by hour and by day of week. Pull abandonment on the same split. Pull the booking or conversion rate on the calls that were answered, and settle on what a converted call is worth. Take all of it from your telephony platform and your CRM. Expect the offered figure to be higher than you thought, because most phone systems under-count callers who hang up during ring.
Write the success measure down before you start. One sentence, with a figure in it and a date. Frame it as a comparison against the baseline you just recorded, and name the system the number will be read out of. If a vendor resists committing to that in writing before the pilot begins, you have your finding early and cheaply.
Agree the failure condition in the same document. Almost nobody does this. Decide in advance what result means you stop: a containment figure below some level, an escalation slower than the caller's patience, any single call that creates a compliance problem. Writing it down changes the vendor's behavior during the pilot, because they now have something to lose. It changes yours too, because you can end the thing without a negotiation.
Run it long enough to see a cycle. Four to eight weeks is the usual range for one flow. You want the full weekly pattern, the Monday surge included, plus at least one period of unusual volume if your business has one. Below a few hundred calls you are reading noise, so a low-volume flow needs longer rather than a braver interpretation of a small sample. Two weeks is a demo with a start date.
Instrument the calls that went badly. Conversions are the easy half. What a pilot owes you is the other half: every escalation with its trigger and the seconds it took, every outcome coded consistently enough to count, the repeat-contact rate on contained calls, and a spot check that the writeback landed as structured fields rather than in a note nobody reads. Then listen to ten of the worst calls yourself. Not a summary of them.
The pilots that prove nothing share a shape. Scope wide enough that no single number is attributable. No baseline, or a baseline taken from the vendor's platform after launch. Success defined as stakeholder satisfaction. Reporting supplied entirely by the party being evaluated. Where several of those are true, the exercise is a procurement formality, and the conclusion will be whatever the loudest person in the room already believed.
One last thing, about free pilots. They look like the lower-risk way in and they usually are not. A free pilot is a sales activity and gets staffed like one: heavy in week one, thin by week four. A paid pilot buys you a scope document, a named owner and the standing to be difficult. Every ScaileAI engagement begins with one, on a single call flow, measured against the client's own baseline.
18 Part five
Where to start
Short answer
Start with the call flow where the loss is easiest to count. If calls reach voicemail outside business hours, start with after-hours coverage. If abandonment spikes at peak and the rest of the month looks fine, start with overflow. If licensed or senior people spend their day on status lookups, start with customer service. If leads go cold before anyone rings back, start with qualification or speed-to-lead callbacks. If cancellation calls are handled by whoever picks up, start with retention. One flow, and a baseline recorded before anybody touches a call.
Pick the flow where the loss is easiest to count. Not the most interesting one. The one where you can already name the number that is wrong, because that is the number a pilot can settle.
- Calls are going to voicemail outside business hours. Start with after-hours coverage. The baseline is sitting in your phone system already. The pattern is documented on the legal, healthcare and home services pages.
- Abandonment spikes at peak and the rest of the month looks fine. Start with overflow handling, which can be set to take only the calls waiting past a threshold and leave the rest alone. Retail, direct response and BPO operations live with this shape.
- Expensive people spend their day on lookups. Order status, technician ETA, application status, case status. That is customer service at tier zero, and it is the quickest way to give senior staff their calendar back. See financial services for the version where whoever answers has to hold a license.
- Leads are not called back fast enough. If the calls are inbound and the problem is who picks up, start with lead qualification or inbound sales. If leads arrive as forms and sit there, the flow you want is outbound speed-to-lead callbacks. Agencies buying media for clients will recognize the shape on the agencies page.
- Churn calls are handled by whoever picks up. Start with retention, where the coded reason is often worth more than the save itself. Direct response operations feel this first, because the whole model is continuity.
- The diary has holes in it. Start with appointment setting, including the reschedule flow that puts a vacated slot back in front of another patient instead of losing it.
One honest caveat on that list. Outbound calling is the single use case with no published call in the library yet. That is a gap in the evidence rather than a gap in the capability, and telling those two apart is most of the skill in reading anybody's proof.
Whichever one you pick, the shape is the same: one flow, and one baseline recorded before anybody touches a call. If you want to see the output before committing to any of it, the call library holds thirty conversations with the full transcript, the coded outcome and what was written into the system behind it, including the ones where the agent declined, escalated or lost the sale. The use case index is the other way in, if you would rather browse by the job than by the symptom.
19 Reference
Glossary: the vocabulary of a real call center.
If a vendor cannot use these terms precisely, they have not run a phone room. Every one of them appears on the reports you should be asking for.
- Abandonment rate
- The share of callers who hang up before reaching anyone. The number most operations under-report, because a caller who gives up after four rings often never registers as a call at all.
- AHT (average handling time)
- Talk time plus hold plus after-call work, averaged. Useful for capacity planning and dangerous as a target: the fastest way to cut it is to resolve less.
- ASA (average speed of answer)
- The average wait before a call is answered. Usually quoted as a blended monthly figure, which hides the peaks where the damage happens.
- Barge-in
- A caller interrupting the agent mid-sentence, and the agent stopping to listen. Its absence is the clearest tell that a voice agent was built from a chat assistant.
- Blended rate
- Any metric averaged across all hours and call types. Blended answer rate is the most commonly quoted number in this industry and the least useful, because performance collapses at peak and the average conceals it.
- Containment
- The share of contacts resolved without a human. Only meaningful when reported next to repeat-contact rate: a call deflected into a callback queue counts as contained on most dashboards and has resolved nothing.
- Disposition
- The coded outcome of a call. Consistent disposition coding is what turns a month of conversations into data you can act on, and it is the thing human floors do least consistently.
- Escalation
- Handing a contact to a human. Judged on two numbers: how long the agent took to recognize it should, and whether the human received the context.
- First-contact resolution
- The share of issues settled on the first call. The metric that most reliably predicts whether customers think your service is good.
- Offered
- Calls that arrived, whether or not they were answered. Answer rate is answered divided by offered, and the denominator is where reporting usually goes wrong.
- Speed to lead
- The delay between a prospect making contact and a business responding. The largest controllable variable in inbound conversion, and mostly a function of who is available at the moment the call lands.
- Tier zero
- The layer that handles contacts before any agent is involved. Where the cost per contact in an outsourced program is decided.
- Warm transfer
- A transfer where the receiving person gets the context. The opposite is a cold transfer, where the customer explains everything again and concludes, correctly, that the two halves of your business do not talk.
- Writeback
- Pushing the outcome of a call into the system of record as structured fields rather than as a note. The difference between a transcript somebody has to read and data you can report on.
20 Reference
The questions buyers actually ask.
Is a managed AI call center the same as an AI receptionist?
No. An AI receptionist answers and takes a message. A managed AI call center runs the conversation operation.
The receptionist product is a good purchase for a small practice with modest volume and simple needs, and it costs a few hundred dollars a month. It is not built to qualify against criteria, handle an objection, read from a live system or write structured data back, and asking it to do those things is how people end up concluding that AI does not work on the phone.
How long does it take to go live?
A single well-scoped call flow typically goes live in two to four weeks. A full operation across several flows and systems takes longer.
After-hours coverage is usually fastest, because the scope is narrow and the baseline is unambiguous: the calls that currently reach nobody. Integration depth is the main variable. Reading from a modern API is quick; writing into an older on-premise system is not.
Will callers know they are talking to AI?
Yes. The agent should identify itself as automated at the start of every call.
Treat any vendor who offers to hide it as a liability. Beyond the regulatory exposure, callers who know what they are speaking to are more direct, and the disclosure removes the moment of discovery that turns a good call into a complaint and a screenshot.
What happens when the AI cannot handle something?
It should recognize that fast and transfer with context. The speed of the handoff matters more than the containment rate.
The risk in front-line automation is not the contact it cannot handle. It is the time it burns before admitting so, because a failed attempt followed by a queue is worse than no automation at all. Ask any vendor for a recording of a call their agent could not handle.
Can it work with our existing team?
That is the normal configuration, not the exception.
AI typically sits in front of the queue, taking what it is scoped for and passing the rest through with a summary, or covers only overflow and after-hours. Replacing a team wholesale is rare, slow and usually the wrong sequence: you learn what the AI should handle by watching what it handles well.
How do we know it is actually working?
By measuring your own baseline before anything changes, then reporting against it.
Answer rate, abandonment, qualification, conversion, containment, repeat contacts and cost per outcome, recorded on your traffic before deployment. Industry benchmarks are close to useless here, because the only comparison that survives scrutiny is your operation against itself.
What does it cost?
Managed engagements are priced on the operation, not per minute. Expect a monthly fee with included volume, plus a one-time implementation.
A per-minute rate covers the metered part of the purchase and says nothing about the design, integration and quality assurance that make it work, so it is a sanity check rather than a basis for choosing. Part 13 of this guide breaks the models down. ScaileAI quotes usage at roughly $0.32 to $0.40 a minute and scopes the implementation and monthly minimum to the operation.
Is it compliant in regulated industries?
Compliance is a function of how the boundaries are built, not of the model. Ask to hear the refusals.
The behaviors that matter are architectural: what the agent will not say, what it escalates by rule, where card data goes, how consent is captured. Any vendor should be able to play you a recording of their agent declining to give advice. Your own counsel signs off the final posture regardless of what a vendor tells you.
Can it handle a sudden spike in call volume?
Yes, and that is the clearest structural advantage over human staffing.
Concurrency is not a constraint, so 400 simultaneous calls are answered the same way one is. Human staffing for a spike means either paying for capacity that may not be needed or queueing the peak, and the peak is usually when intent is highest.
What happens to our data?
It should stay in your systems. The provider operates the calls, not your customer records.
Recordings, transcripts and structured outcomes belong in your CRM or case management system under your retention policy. Ask specifically about where recordings live, who can access them, how long they are kept and what happens to them if you leave.
Keep reading
Where this goes next.
Start with one call flow.
We measure your baseline before we touch a call, so the result is arithmetic rather than opinion.