Containment and AI resolution, defined precisely, measured against a baseline we capture before go live, and reported monthly. If a supplier will not write their definitions down, that is your answer.
Most vendors quote containment without telling you how they calculate it. Here are our definitions, including the cases we exclude. If another supplier will not write theirs down, that is the answer to your question.
Contacts fully handled by the agent, divided by contacts offered to the agent.
Counts: the customer got what they came for and the contact ended without a human.
Does not count: the caller hung up, the agent read a message and transferred, the contact was deflected to a web page, or the customer called back about the same issue within 48 hours. Repeat contacts are subtracted from containment, which is why our number starts lower than the market average.
Contained contacts where the intended action actually completed in the system of record.
Counts: the payment posted, the arrangement was written to the account, the case was created and routed.
Does not count: the agent said it would do something and no downstream write is evidenced. Resolution is measured against your systems, not against the transcript.
Contacts transferred to a human, divided by contacts offered.
Split three ways, because the three mean different things: the customer asked for a human, a vulnerability or compliance rule fired, or the agent could not proceed. The third is the only one you should be trying to reduce. The middle one going up can be a good sign.
Total run cost for the period, divided by contacts handled.
Includes: model inference, telephony, transcription, orchestration compute, storage, and the support effort to keep it running.
Excludes nothing. A cost per contact that leaves out inference or support is a marketing number. We report the whole bill against the whole volume.
Measured separately for agent handled and human handled contacts.
Blending them hides the effect. Automation usually leaves the harder contacts with your people, so human AHT can rise while total cost falls. We report both so nobody is surprised at month three.
Contacts where every required disclosure and check was evidenced, divided by contacts where they were required.
Scored on the transcript and the event log, not sampled by hand. For FCA Consumer Duty work this is the number that matters most, and it is the one that has to be 100 percent rather than trending upward.
Every number above comes from an event, not from an opinion. The same instrumentation serves your monthly pack and your auditor.
Intent, tool call, tool result, guardrail decision and handover reason are emitted per turn with a correlation id that spans the whole contact, across voice, email and chat.
Events land in a store with retention set to your policy, commonly seven years for regulated work, with the transcript, the model version and the prompt version that produced each decision.
Resolution is confirmed by reading back from the system of record. If the agent claims a payment and no payment exists, that contact is not a resolution.
Distributed tracing across the contact flow, the orchestrator, the tools and the model calls, so a slow or failed contact can be explained rather than guessed at.
Want to score your own deployment against this standard? Find out what is blocking you →
This is the step most projects skip, and it is why so many AI programmes cannot say whether they worked.
Volume by intent, handle time, cost per contact, first contact resolution, repeat contact rate, and your current compliance sampling result. Taken from your platform, not from ours.
Named intents, in writing, with the volume each represents. Containment is only meaningful against a defined denominator.
The definitions on this page, signed off before build, so nobody renegotiates the metric after the result is known.
Shadow mode where the agent decides but does not act, so accuracy is known before a customer is affected.
One page your executive sponsor can take to a board, plus the detail your operations team needs to act.
Containment and AI resolution against baseline. Cost per contact against baseline. Total saving for the period and cumulative. Escalation split three ways. Compliance adherence. One line on what changed and why.
Per intent performance, top failure reasons ranked by volume, the intents worth adding next, transcripts of every escalation that was not customer requested, and any prompt or model version change with its measured effect.
Week four. The agent is live on a narrow set of intents. Containment on that scope is lower than the headline numbers vendors advertise, because the definitions above are strict and because your edge cases are not yet handled. Compliance adherence should already be at 100 percent, since that is built in rather than tuned.
Month three. Containment has climbed as failure reasons are worked through in order of volume. Human handle time has probably risen, because your people now get the harder contacts. Cost per contact is measurably down. You can attribute the change because you have the baseline.
What we will not do. Quote a containment number for your operation before we have seen your traffic. Any supplier who does is quoting somebody else's contact centre.
Every Rel8 deployment passes these before it carries live traffic. They cover measurement discipline (1 to 3), production readiness (4, 5, 8 and 9) and compliance posture (6, 7 and 10). We do not run proofs of concept. We run proofs of production.
Taken from your platform before any AI goes live. Without it, every number that follows is unfalsifiable, which is why it is gate one.
A caller who gives up, gets pushed to a webpage, or rings back two days later was not contained. Subtracting all three is what separates a real figure from a headline one.
A resolution counts when something changed in your system, not when the transcript reads as though it did.
When the agent hands over, your advisor sees the whole conversation and what was already attempted. Repeating themselves is the moment a customer decides the AI wasted their time.
A test suite that runs each deploy, not a one-time accuracy claim from before launch. Models drift and prompts get edited.
Written down, and demonstrated in a real conversation. A policy document nobody has exercised is not a control.
One contact, pulled up and replayed end to end, long after the fact. This is what a regulator asks for and it cannot be retrofitted.
Measured under real load, not in a demo. The average is not the number that makes a caller talk over the agent.
An explicit path when the model is down or returns low confidence. Every deployment meets this eventually; the question is whether it was designed or discovered.
Through your review, with the findings known and closed. Not scheduled, not accepted as risk.
Identity verification is not one of the ten. It is standard on every package, so it is not something to pass.
A pass is point in time and it decays. Models are retired, prompts get edited, traffic shifts. A deployment that passed in March is not a deployment that passes today, which is what the monthly managed option exists for.
This page is the standard. The assessment scores your deployment against it and tells you which gate is blocking you.
Find out what is blocking you15 questions. 3 minutes. Free. Your recommendations immediately.
Already know what you need? Book a discovery call or email hello@rel8.cx