AI Voice Agent Post-Pilot Go/No-Go: The Criteria UK Contact Centre Leaders Should Use to Decide Whether to Scale or Stop
Author: Arkadas Kilic, Founder & CEO, Rel8 CXYou ran the pilot. You have data. Now comes the decision that most contact centre leaders get wrong.
They either scale prematurely because the demo looked good, or they kill a viable programme because they measured the wrong things. Both outcomes cost money. This post gives you a structured go/no-go framework built from production deployments, not theory.
Why Most Pilot Evaluations Fail Before They Start
The failure usually happens at the design stage. Pilots get launched without pre-agreed success criteria, which means the post-pilot review becomes a political negotiation rather than a data review.
Before you read the criteria below, confirm one thing: did you define success thresholds before the pilot went live? If not, treat this framework as your retrospective baseline and use it to set hard thresholds before any scale decision is formalised.
The Four Domains of a Go/No-Go Decision
1. Containment and Resolution Performance
This is the headline metric, but it is also the most misread one.
Containment rate measures how many calls the AI voice agent handled without transferring to a human. A pilot containment rate of 40% sounds disappointing until you account for call mix. If your pilot routed complex complaints and billing disputes alongside simple FAQs, 40% may be excellent. If it only handled password resets, 40% is a problem. What to look for:- Containment rate by intent category, not blended average
- First-contact resolution (FCR) on contained calls: did the customer actually get their answer, or did they call back within 24 hours?
- Escalation trigger analysis: are transfers happening because the agent genuinely cannot handle the query, or because of poor prompt design or missing integrations?
2. Customer Experience Signals
Containment without satisfaction is not a business outcome. A customer who gets the wrong answer quickly is worse than a customer who waits for a human.
Metrics to pull from the pilot:- Post-call CSAT scores for AI-handled calls versus human-handled calls on the same intents
- Repeat contact rate within 48 hours for contained calls
- Abandonment rate during AI interactions compared to your IVR baseline
- Sentiment analysis on call transcripts, looking specifically for frustration markers in the first 60 seconds
3. Operational and Commercial Viability
This is where pilots often look better than they are. Cost per interaction drops when you contain calls, but that number is only meaningful if you account for the full cost of running the AI layer.
Calculate the real unit economics:- Cost per contained interaction (AI infrastructure + maintenance + ongoing prompt tuning)
- Cost per escalated interaction (AI handling time plus human handling time)
- Blended cost per contact versus your current baseline
- Time to value: at what containment volume does the programme break even?
4. Technical Readiness for Scale
A pilot that works at 500 calls per month may not work at 50,000. Technical readiness is the most underweighted domain in most go/no-go reviews.
Questions to answer before scaling:- Did the integration layer hold up under peak load, or were there latency spikes above 800ms that degraded the experience?
- How many manual interventions were required to keep the pilot running? If your AI engineer spent more than four hours per week on reactive fixes, that cost does not disappear at scale, it multiplies.
- Is the system built on production-grade infrastructure, or is it a prototype running on development credentials?
- Are your compliance controls production-ready? For UK contact centres, that means PCI DSS scope is clearly defined, call recording consent flows are compliant with UK GDPR and FCA guidance where applicable, and data residency is confirmed within the UK or EEA.
- What is the rollback plan if a production incident occurs at scale?
The Go/No-Go Scorecard in Practice
Score each domain on a simple three-point scale:
- 3: Threshold met, clear evidence of production readiness
- 2: Threshold partially met, addressable gaps identified with a clear fix plan
- 1: Threshold not met, root cause unclear or fix requires significant rework
| Domain | Weight | Your Score |
|---|---|---|
| Containment and Resolution | 30% | |
| Customer Experience | 30% | |
| Operational and Commercial Viability | 25% | |
| Technical Readiness | 15% |
The Conditional Go: What It Actually Means
Most pilots land in the conditional zone. That is not failure. It means you have a viable programme with specific gaps.
The discipline is in the remediation list. Each item needs:
- A specific outcome, not a vague improvement target
- An owner with authority to deliver it
- A deadline of no more than four weeks
- A re-test protocol
If any item cannot be scoped to that standard, it is not a conditional go. It is a no-go with optimism attached.
What Scaling Actually Looks Like
Scaling an AI voice agent is not flipping a switch. A responsible scale programme moves through three stages:
Stage 1 (weeks one to two): Increase volume by 3x on proven intents only. Monitor all four domains daily. Stage 2 (weeks three to six): Expand intent coverage to the next tier of call types. Introduce any new integrations that were out of scope in the pilot. Stage 3 (weeks seven onwards): Full production volume with automated monitoring, alerting, and a documented runbook for your operations team.This is the pattern we follow on every production deployment. It is not cautious, it is how you protect the business case while building the evidence base for further investment.
A Note on Stopping
Stopping is not failure. A no-go decision that prevents a costly failed scale is a good outcome. The organisations that struggle are those that continue past clear no-go signals because of sunk cost pressure or stakeholder momentum.
If the scorecard says stop, the right question is: what would need to be true for this to be a go in six months? Answer that question honestly, and you have a programme worth restarting.
Final Checklist Before You Make the Call
- [ ] Containment rate reviewed by intent category, not blended
- [ ] FCR validated against 48-hour repeat contact data
- [ ] AI CSAT compared to human CSAT on equivalent intents
- [ ] Unit economics calculated with full infrastructure and maintenance costs
- [ ] Break-even volume validated against realistic demand forecast
- [ ] Latency data reviewed at 95th percentile, not average
- [ ] Weekly maintenance hours logged and projected to scale
- [ ] Compliance documentation reviewed by DPO or compliance lead
- [ ] Rollback procedure documented and tested
- [ ] Scorecard completed with weighted score calculated
- [ ] Conditional items (if any) scoped with owners and deadlines
We build autonomous AI voice agents for UK contact centres and take them to production in 4 to 6 weeks. If you are at the go/no-go stage and want a structured review of your pilot data before committing scale budget, let us walk through it with you.
Book a discovery callIs your pilot going to reach production?
Fifteen questions, three minutes, no cost. You get a score against the ten checks we run every deployment through, and a straight answer on what is blocking yours.
Score your readiness