Agentic AI Vendor Evaluation: A Scorecard for UK Contact Centre Procurement Teams
Rel8 CX is an AWS Advanced Partner that builds autonomous AI agents for regulated UK contact centres. We've been on both sides of procurement conversations, and the same pattern repeats: procurement teams ask the wrong questions, vendors answer the wrong questions, and the contact centre ends up with a proof of concept that never reaches production.
This post gives you the scorecard we'd want you to use when evaluating us, or anyone else.
Why Most AI Vendor Evaluations Fail
Most RFPs for contact centre AI are written by people who've never deployed one. They ask about "AI capabilities" and "integration flexibility" and "scalability roadmap." Vendors answer with slide decks, demo environments, and reference customers who signed NDAs.
The result is a procurement process that selects for the best presenter, not the best builder.
The questions that actually matter are operational: How long until this is live? Who owns compliance? What happens when the agent fails? What does your AWS architecture look like in production?
Here's how to structure your evaluation so you get real answers.
The Scorecard: 7 Evaluation Dimensions
Score each vendor 1 to 5 on each dimension. Weight the dimensions according to your organisation's priorities. A regulated financial services contact centre should weight compliance and architecture heavily. A high-volume insurance operation should weight time-to-production and containment rates.
Dimension 1: Production Track Record
What you're testing: Can this vendor ship working agents into live contact centre environments, or do they build prototypes? Questions to ask:- How many autonomous AI agents do you have in production today, not pilots, not POCs?
- What was the containment rate at go-live for your last three deployments?
- Can you show us a production deployment timeline from kick-off to live?
- References are all "under NDA" with no verifiable outcomes
- The vendor talks about "deployments" that turn out to be internal pilots
- Go-live timelines are measured in quarters, not weeks
Dimension 2: Regulatory and Compliance Architecture
What you're testing: Is compliance an afterthought bolted on at the end, or is it built into the architecture from day one?For UK contact centres, the regulatory landscape is specific and demanding:
- FCA Consumer Duty (effective July 2023): Requires firms to deliver good outcomes for retail customers. AI agents that deflect complaints, obscure fees, or fail vulnerable customers create direct regulatory exposure.
- FCA Consumer Credit sourcebook (CONC): Governs debt collection contact. AI agents in collections must follow CONC rules on contact frequency, time of contact, and treatment of customers in financial difficulty.
- GDPR and UK Data Protection Act 2018: All customer data processed by AI agents must meet UK GDPR standards. This includes data residency, retention policies, and the right to human review.
- PCI DSS: If the agent handles payment card data, the architecture must be scoped and assessed.
- FCA SYSC 8 (Outsourcing): AI vendors are typically third-party service providers under SYSC 8. Your firm retains regulatory responsibility. The vendor must support your oversight obligations.
- Walk me through how your agent handles a customer who identifies as vulnerable mid-interaction.
- Where does customer data reside? Which AWS regions? Can you confirm UK data residency?
- How do you support our SYSC 8 third-party oversight obligations?
- What audit logging does your architecture produce for regulatory review?
- How does the agent handle a complaint? Does it escalate, and what's the handoff process?
- The vendor says "we're GDPR compliant" without explaining the architecture
- Vulnerability handling is a manual process, not built into the agent logic
- The vendor is unfamiliar with Consumer Duty or treats it as a checkbox
- Audit logs are not available in real time or are held by the vendor, not you
Dimension 3: AWS Architecture Depth
What you're testing: Is the vendor a genuine AWS builder or a third-party tool reseller wrapped in AWS branding?This matters for three reasons. First, AWS-native architecture means your data stays in your AWS environment, not a vendor's SaaS platform. Second, AWS-native builds integrate directly with Amazon Connect, your existing contact centre infrastructure, without middleware that creates latency and failure points. Third, AWS-native vendors can leverage services like Amazon Bedrock, Amazon Transcribe, Amazon Lex, and AWS Lambda in ways that reduce cost and increase reliability compared to third-party AI platforms.
Questions to ask:- Is your architecture deployed into our AWS account or yours?
- Which AWS services does your agent stack use? Walk me through the architecture.
- Are you an AWS Advanced Partner or AWS Select Partner? What competencies do you hold?
- How do you use Amazon Connect natively versus third-party integrations?
- How is your infrastructure defined? Do you use CDK, Terraform, or CloudFormation?
| Tier | What it means |
|---|---|
| AWS Partner (entry) | Registered, minimal requirements |
| AWS Select Partner | Some validated experience, basic requirements met |
| AWS Advanced Partner | Demonstrated delivery capability, customer references, staff certifications |
| AWS Premier Partner | Highest tier, significant delivery volume and specialisation |
An AWS Advanced Partner has met specific requirements around certified staff, customer references, and delivery capability. It is not a marketing badge.
Red flags:- The vendor's architecture relies heavily on non-AWS AI platforms and uses AWS only for hosting
- Infrastructure is deployed into the vendor's AWS account, not yours
- The vendor cannot explain their CDK or IaC approach
- AWS partnership claims are not verifiable on the AWS Partner Finder
Dimension 4: Time to Production
What you're testing: How long until this agent is handling real customer interactions?This is where most vendors lose points. The industry average for enterprise AI deployments is 6 to 18 months. Most of that time is not technical. It is scoping, stakeholder alignment, procurement cycles, and change management. A vendor who has done this before knows how to compress the timeline.
Benchmark: A production-ready agentic AI deployment for a UK contact centre should take 4 to 6 weeks from kick-off to live if the vendor has done it before and the contact centre has a functioning Amazon Connect environment. Questions to ask:- What does your deployment timeline look like from contract signature to first live call?
- What are the top three reasons deployments run late? How do you mitigate them?
- What do you need from us in week one?
- Have you deployed into a contact centre that was migrating from a legacy platform at the same time?
| Week | Activities |
|---|---|
| 1 | Discovery, call flow mapping, integration scoping, AWS environment access |
| 2 | Agent architecture design, intent taxonomy, compliance review |
| 3 | Core agent build, Amazon Connect integration, test environment |
| 4 | UAT, edge case handling, vulnerability scenario testing |
| 5 | Soft launch with live traffic at 10%, monitoring, iteration |
| 6 | Full production rollout, containment baseline established |
- The vendor's timeline starts with a "discovery phase" that takes 6 to 8 weeks before any build begins
- The vendor cannot give you a week-by-week plan
- Timeline is conditional on "further scoping"
Dimension 5: Containment and Outcome Metrics
What you're testing: Does this vendor measure the right things, and do their numbers hold up?Containment rate is the percentage of interactions the AI agent resolves without transferring to a human agent. It is the primary operational metric for contact centre AI. But containment rate alone is not enough. A 70% containment rate that generates FCA complaints is worse than a 40% containment rate with zero regulatory issues.
The metrics that matter:| Metric | What it measures | Benchmark |
|---|---|---|
| Containment rate | % of interactions resolved by AI | 35% to 55% at go-live for voice; 60%+ for digital |
| Escalation accuracy | % of escalations that genuinely needed a human | Above 90% |
| Average handle time (AI) | Time per AI-handled interaction | Depends on use case, ask for actuals |
| Customer satisfaction (CSAT) | Post-interaction satisfaction for AI-handled calls | Should be within 5 points of human-handled baseline |
| Complaint rate | FCA reportable complaints generated by AI interactions | Should be lower than human baseline |
| Vulnerable customer escalation rate | % of interactions flagged and escalated for vulnerability | Benchmark against your current human rate |
- What containment rate did your last three deployments achieve at week one, week four, and week twelve?
- How do you measure customer satisfaction for AI-handled interactions?
- What is the complaint rate for AI-handled interactions versus human-handled in your production deployments?
- How do you define a successful deployment?
- The vendor only measures containment rate and not customer outcomes
- Containment numbers are from demo environments, not production
- The vendor cannot tell you what happened to CSAT when the agent went live
Dimension 6: Ongoing Operations and Ownership Model
What you're testing: What happens after go-live? Who owns the agent's performance?This is where many contact centres get burned. The vendor delivers a working agent, hands over documentation, and disappears. The contact centre team doesn't have the skills to tune the agent, handle edge cases, or adapt it when call drivers change. Performance degrades. The project is declared a failure.
Questions to ask:- What does your post-go-live support model look like?
- Who monitors the agent's performance day to day? Us or you?
- How do you handle intent drift when customer call drivers change?
- What does retraining or tuning cost after the initial deployment?
- Do you offer a managed service, or do you hand over and exit?
- What is your SLA for production incidents?
| Model | What it means | Best for |
|---|---|---|
| Build and transfer | Vendor builds, trains your team, exits | Organisations with internal AI engineering capability |
| Managed operations | Vendor runs the agent ongoing, you own the outcomes | Organisations without internal AI ops capability |
Neither model is inherently better. But you need to know which one you're buying before you sign.
Red flags:- Post-go-live support is not defined in the contract
- The vendor's exit plan involves handing over a codebase with no documentation
- Tuning and retraining are sold as separate engagements at day rates
Dimension 7: Team Composition and Delivery Model
What you're testing: Who actually builds this? Are they practitioners or account managers?The contact centre AI market is full of consultancies that sell AI transformation and deliver PowerPoint decks. The tell is the team they put on your project. If the people who present in the sales process are not the people who build the system, you are buying a brokered service.
Questions to ask:- Who will be working on our project day to day? Can we meet them before signing?
- What is the ratio of engineers to project managers on a typical deployment?
- Do your engineers hold AWS certifications? Which ones?
- Have the engineers on our project deployed production AI agents before? Where?
- Is any of the build subcontracted?
- The sales team and the delivery team are entirely different people
- The vendor talks about "our delivery partners" (this means subcontractors)
- Engineers are not named until after contract signature
- AWS certifications are held by the sales team, not the engineers
The Scorecard Summary
| Dimension | Weight (adjust to your context) | Your Score (1 to 5) | Weighted Score |
|---|---|---|---|
| Production track record | 20% | ||
| Regulatory and compliance architecture | 25% | ||
| AWS architecture depth | 15% | ||
| Time to production | 15% | ||
| Containment and outcome metrics | 10% | ||
| Ongoing operations and ownership | 10% | ||
| Team composition and delivery model | 5% | ||
| Total | 100% |
A vendor scoring below 3.0 weighted average is a risk. A vendor scoring above 4.0 on compliance and production track record is worth serious consideration regardless of their total.
Questions That Reveal Everything
If you only have 30 minutes with a vendor, ask these five questions and listen carefully:
1. "Show me a production deployment timeline from contract to live call." Not a proposed timeline. An actual one.
2. "Walk me through how your agent handles a customer who says they're struggling financially." This tests vulnerability handling, Consumer Duty awareness, and CONC compliance simultaneously.
3. "Who specifically will be building our agent? Can we meet them today?" This separates practitioners from sales organisations.
4. "What went wrong on your last deployment and how did you fix it?" Every real deployment has problems. Vendors who say everything went smoothly are lying or selling you a demo.
5. "Where does our customer data sit and who has access to it?" The answer should be: in your AWS account, with access controlled by your IAM policies.
A Note on Proof of Concepts
Many vendors will offer a free or low-cost POC to get a foot in the door. Be cautious. A POC that runs in a vendor's demo environment, uses synthetic data, and has no compliance controls is not evidence of production capability. It is a sales tool.
If you run a POC, insist that it:
- Uses real (anonymised) customer interaction data
- Runs in your AWS environment, not the vendor's
- Includes at least one vulnerable customer scenario
- Produces the same audit logs the production system would produce
- Has a defined path from POC to production with costs and timelines agreed upfront
A vendor who can't commit to these POC conditions is not ready to build your production system.
How Rel8 CX Scores on This Scorecard
We build production agentic AI agents for regulated UK contact centres. We are an AWS Advanced Partner. We deploy in 4 to 6 weeks. Our engineers are the people who present in the sales process. Compliance is built into our architecture from day one, not added at the end.
We've deployed into debt collections, insurance, and financial services environments where FCA compliance and Consumer Duty obligations are not optional. We know what vulnerable customer handling looks like in a production agent. We know what CONC-compliant collections AI sounds like on a live call.
We don't claim to score 5 on every dimension for every contact centre. We'll tell you where we're the right fit and where we're not.
If you want to run us through this scorecard, we're ready.
Book a discovery callReady to put AI agents into production?
Book a discovery call. We will assess your use case and show you what 4 to 6 weeks to production looks like.
Book a Discovery Call