AI Voice Agent Post-Pilot Go/No-Go: The Criteria UK Contact Centre Leaders Should Use to Decide Whether to Scale or Stop

Arkadas Kilic
Author: Arkadas Kilic, Founder & CEO, Rel8 CX

You ran the pilot. You have data. Now comes the decision that most contact centre leaders get wrong.

They either scale prematurely because the demo looked good, or they kill a viable programme because they measured the wrong things. Both outcomes cost money. This post gives you a structured go/no-go framework built from production deployments, not theory.


Why Most Pilot Evaluations Fail Before They Start

The failure usually happens at the design stage. Pilots get launched without pre-agreed success criteria, which means the post-pilot review becomes a political negotiation rather than a data review.

Before you read the criteria below, confirm one thing: did you define success thresholds before the pilot went live? If not, treat this framework as your retrospective baseline and use it to set hard thresholds before any scale decision is formalised.


The Four Domains of a Go/No-Go Decision

1. Containment and Resolution Performance

This is the headline metric, but it is also the most misread one.

Containment rate measures how many calls the AI voice agent handled without transferring to a human. A pilot containment rate of 40% sounds disappointing until you account for call mix. If your pilot routed complex complaints and billing disputes alongside simple FAQs, 40% may be excellent. If it only handled password resets, 40% is a problem. What to look for: Go threshold: Containment above 55% on in-scope intents, with FCR above 70% on contained calls, and fewer than 15% of escalations flagged as avoidable. No-go signal: Containment below 35% on in-scope intents, or FCR below 50%, or more than 30% of escalations caused by system gaps rather than genuine complexity.

2. Customer Experience Signals

Containment without satisfaction is not a business outcome. A customer who gets the wrong answer quickly is worse than a customer who waits for a human.

Metrics to pull from the pilot: Go threshold: AI CSAT within 8 points of human CSAT on equivalent intents. Repeat contact rate no higher than your human-handled baseline. Abandonment rate no worse than your existing IVR. No-go signal: AI CSAT more than 15 points below human CSAT. Repeat contact rate 20% higher than baseline. Abandonment rate climbing after week two of the pilot.

3. Operational and Commercial Viability

This is where pilots often look better than they are. Cost per interaction drops when you contain calls, but that number is only meaningful if you account for the full cost of running the AI layer.

Calculate the real unit economics: UK-specific consideration: Factor in your agent wage costs accurately. At a blended fully loaded cost of roughly £22 to £28 per hour for a UK contact centre agent, the break-even point on a well-configured AI voice agent typically sits between 800 and 1,200 contained interactions per month at current AWS infrastructure pricing. Go threshold: Blended cost per contact at least 20% below current baseline at pilot volume, with a clear path to 35% reduction at scale. No-go signal: Cost per contact higher than baseline, or break-even requires more than 3x current pilot volume with no credible demand forecast to support it.

4. Technical Readiness for Scale

A pilot that works at 500 calls per month may not work at 50,000. Technical readiness is the most underweighted domain in most go/no-go reviews.

Questions to answer before scaling: Go threshold: Latency consistently below 600ms at 95th percentile. Fewer than two hours per week of reactive maintenance. Compliance controls documented and signed off by your DPO or compliance team. Rollback procedure tested. No-go signal: Latency spikes above 1,000ms on more than 5% of calls. More than six hours per week of reactive maintenance during the pilot. Compliance documentation incomplete. No tested rollback procedure.

The Go/No-Go Scorecard in Practice

Score each domain on a simple three-point scale:

DomainWeightYour Score
Containment and Resolution30%
Customer Experience30%
Operational and Commercial Viability25%
Technical Readiness15%
Weighted score of 2.4 or above: Scale with a phased rollout plan. Weighted score of 1.8 to 2.3: Conditional go. Define three to five specific remediation items with owners and deadlines before committing scale budget. Weighted score below 1.8: No-go. Return to design. A failed scale is significantly more expensive than a delayed one.

The Conditional Go: What It Actually Means

Most pilots land in the conditional zone. That is not failure. It means you have a viable programme with specific gaps.

The discipline is in the remediation list. Each item needs:

If any item cannot be scoped to that standard, it is not a conditional go. It is a no-go with optimism attached.


What Scaling Actually Looks Like

Scaling an AI voice agent is not flipping a switch. A responsible scale programme moves through three stages:

Stage 1 (weeks one to two): Increase volume by 3x on proven intents only. Monitor all four domains daily. Stage 2 (weeks three to six): Expand intent coverage to the next tier of call types. Introduce any new integrations that were out of scope in the pilot. Stage 3 (weeks seven onwards): Full production volume with automated monitoring, alerting, and a documented runbook for your operations team.

This is the pattern we follow on every production deployment. It is not cautious, it is how you protect the business case while building the evidence base for further investment.


A Note on Stopping

Stopping is not failure. A no-go decision that prevents a costly failed scale is a good outcome. The organisations that struggle are those that continue past clear no-go signals because of sunk cost pressure or stakeholder momentum.

If the scorecard says stop, the right question is: what would need to be true for this to be a go in six months? Answer that question honestly, and you have a programme worth restarting.


Final Checklist Before You Make the Call


We build autonomous AI voice agents for UK contact centres and take them to production in 4 to 6 weeks. If you are at the go/no-go stage and want a structured review of your pilot data before committing scale budget, let us walk through it with you.

Book a discovery call

Is your pilot going to reach production?

Fifteen questions, three minutes, no cost. You get a score against the ten checks we run every deployment through, and a straight answer on what is blocking yours.

Score your readiness