What should a Saudi contact centre test before buying an AI voice agent?
An AI voice agent is software that answers or places phone calls using streaming speech recognition, a reasoning model, and speech synthesis, running in sequence on live audio. Test it with roughly 30 recorded calls across five Arabic dialects, latency measured at median and 95th percentile, explicit code-switching scenarios, and a live human handoff before signing anything.
Buyers decide inside the first five seconds of a live call, not from a slide deck or a vendor’s recorded reel. A caller hears pacing and pickup speed before a single correct answer registers. The demo a sales team runs is a scripted path over a clean line with a cooperative speaker. Production calls carry background noise, interruptions, dialect switching, and callers who talk over the agent because they’ve dealt with worse systems before.
This guide covers six things: a recorded test plan built around 30 calls, latency measurement with defined percentiles, dialect coverage tested separately rather than lumped into one Arabic score, warm handoff behavior including what happens when no agent is free, a residency and recording checklist for where audio is processed, and a scorecard with pass thresholds both sides sign before the pilot starts.
OnClarity (www.onclarity.com) is an enterprise AI customer experience platform whose AI agents work across chat and voice, with bring-your-own-model and open-weight options that can run in-region or fully on-premise, so the test in this guide can run against your own dialect mix rather than a vendor’s reference accent.
How does an AI voice agent actually work on a live call?
An AI voice agent handles a call through a pipeline of independent parts: media ingress, streaming speech to text, a reasoning model, and text to speech, each passing data forward while an orchestrator holds conversation state. No single model understands the call end to end. Seeing the request path is what makes every later test in this guide legible instead of a black box.
Here is the six-step path for one inbound call:
The call arrives over a SIP trunk or the public telephone network and is normalized into an RTP audio stream, typically in 20ms frames.
Streaming speech recognition consumes those frames continuously, not in one batch, emitting partial hypotheses that get revised as more audio arrives.
A separate voice activity detector decides when the caller has actually stopped talking, as distinct from pausing mid-thought.
Once that endpoint fires, the finalized transcript, conversation history, and any retrieved account or policy data go to the reasoning model, which generates a response.
Output tokens stream to the text-to-speech engine as they’re generated, so synthesis starts on the first sentence before the last one exists.
Synthesized audio returns down the same media path, frame by frame, to the caller.
State lives in three separate places, and conflating them is a common design mistake. Dialogue history sits in the orchestrator. Business state, account balances, ticket status, order history, lives in the CRM, fetched on demand. The model itself holds nothing durable; it is a stateless function called fresh each turn with whatever context the orchestrator assembles.
Four places accumulate voice agent latency, and a test plan should isolate each rather than measure only the total: network and jitter buffer delay in the media path, the endpointing silence threshold while the system waits to confirm the caller finished, model time-to-first-token during reasoning, and time-to-first-audio during synthesis. This is the core tradeoff: a shorter endpoint threshold cuts perceived latency but raises the false turn-end rate, so the agent talks over a caller who was mid-sentence. A longer threshold avoids that but reads as sluggish. No setting removes both failure modes at once.
Barge-in, letting a caller interrupt mid-sentence, requires full-duplex audio plus echo cancellation, since the agent’s own synthesized voice would otherwise be picked up by the microphone and misread as a new turn. That’s why barge-in misfires in noisy rooms: background noise degrades the detector’s ability to tell a real interruption from ambient sound, and echo cancellation needs a few seconds to adapt at call start.
Two architectures exist. Cascaded systems keep speech to text, the reasoning model, and text to speech as separate, swappable components, trading some end-to-end latency for the ability to test and replace any single engine. Speech-to-speech models collapse the pipeline into one model reasoning directly on audio, which can cut latency and preserve prosody but makes individual components impossible to isolate or swap. This is why engine-level testing matters: OnClarity is model- and cloud-agnostic, with bring-your-own-model and open-weight options that can run in-country, letting buyers test the models they choose against their own calls rather than accept a bundled default. See the AI agents platform for how this fits together.
Which Arabic dialects and code-switching cases break AI voice agents?
Dialectal Arabic breaks speech recognition for a mechanical reason: acoustic and lexical distance from Modern Standard Arabic. Most Arabic training data skews toward MSA broadcast news, so a model that never saw enough Najdi or Khaleeji speech has no statistical basis for recognizing it well. Dialectal Arabic also lacks a standard orthography, so the same spoken word gets written several ways across regions, which degrades transcript matching against knowledge bases as much as raw recognition. Published Arabic ASR benchmarks consistently show dialectal speech scoring worse than MSA broadcast audio, though the size of that gap varies by corpus and channel, and none of those benchmarks use 8kHz telephony audio with contact centre noise. The only number that matters is the one your own recorded calls produce.
This is why “Arabic supported” is not a usable answer. Najdi, Hijazi, Khaleeji, Egyptian, and MSA are five separate test sets requiring five separate scores.
Dialect | Caller profile | Minimum calls | Typical failure | Score criterion |
|---|---|---|---|---|
Najdi | Riyadh-region, everyday register | 6 | Misrecognized function words, wrong routing | Accuracy on intent-bearing phrases |
Hijazi | Jeddah/Makkah, mixed register | 6 | Vowel and stress errors on common requests | Correct entity extraction |
Khaleeji | Gulf-accented speaker | 6 | Number and date recognition failures | Digit-string accuracy |
Egyptian | Colloquial speech | 4 | Fewer failures; test colloquial vs formal split | WER by register |
MSA | Formal or scripted register | 4 | Rare failures; baseline case | Confirm this isn’t the only dialect tested |
Recruit real speakers for each row, not one bilingual tester doing five accents. Score each cell independently and never average them into one Arabic voice AI number, since an average hides exactly the failure a Khaleeji caller will hit on launch day.
Code-switching splits into two cases. Inter-sentential switching is a full turn in English followed by a full turn in Arabic. Intra-sentential switching embeds English inside an Arabic sentence, a product name, a number, a technical noun, mid-utterance. That second pattern is the common Gulf case, and it’s harder mechanically because it requires the system to change language identity multiple times inside one continuous audio stream rather than once between turns, a place where conversational AI struggles because language identity can shift within one utterance.
The failure mechanism follows from the pipeline. Language identification typically runs per utterance, not per word. Once that window locks a language, the decoder’s locale is fixed for the segment, and the system is trying to interpret natural language within a single continuous audio stream while the speech-to-text and text-to-speech voice binds to a single language. Put those together and you get the two failures Gulf buyers report: the agent stays in Arabic after the caller switches to English mid-sentence, or it starts a reply in English and reverts to Arabic partway through, because the decoder locale never actually changed, only the transcript’s guessed language did.
Test this with four scripted utterances, spoken naturally, not read like a word list:
Arabic carrying an embedded English technical term (a password reset request, for example).
English carrying Arabic filler, such as an IBAN number spoken partly in Arabic mid-sentence.
A full turn in English immediately followed by a full turn in Arabic, no pause.
A number or reference code spoken in English digits inside an otherwise Arabic sentence.
No vendor handles code-switching Arabic English flawlessly. Treat that as fact, not caution. A demo call that never switches languages hasn’t been tested for this failure mode at all, and a Gulf pilot that skips it will find out live.
Two adjacent problems deserve their own scenarios. Accented English from non-native speakers is common in Saudi contact centres serving expatriate callers, and it stresses the English-side recognition independently of Arabic capability. Numbers, national ID digits, and alphanumeric reference codes fail differently than words, so score them separately rather than assume conversational fluency implies digit accuracy. Voice quality depends on the specific engine paired with the specific dialect, which is why testing speech recognition, reasoning, and speech output independently against your own calls matters more than any vendor’s language list. Review the AI agents platform and deployment options for how model choice and in-country hosting fit this test plan.
How do you measure AI voice agent latency and naturalness on your own calls?
Voice agent latency has two components, and a scorecard should measure both separately, in milliseconds, at median and 95th percentile, from your own recorded calls. Naturalness is pass or fail, decided by native speakers per dialect answering one question: would you have stayed on this call.
Time to first word is the interval from the end of the caller’s speech energy to the first audible phoneme of the agent’s reply, measured on the waveform. Turn-taking delay is the full inter-turn gap the caller actually hears, including network transit, jitter buffer delay, and any dead air before audio starts. These are not the same number, and a vendor quoting one while you expect the other is how a scorecard gets disputed later.
The measurement point matters as much as the definition. Record at the PSTN edge or on a handset, not inside the vendor’s own stack. Vendor-side timers typically start when their orchestrator receives audio and stop when their engine emits a byte, which excludes the carrier leg, the SIP trunk, and any jitter buffer entirely. Run waveform analysis on the recorded audio to find the actual acoustic boundaries rather than trusting log timestamps from a system with an interest in a smaller number.
Report median and 95th percentile together, never one alone. A system with an unremarkable median and a long tail is the one that generates dead-air calls, the kind that get escalated after a caller complains the agent went silent. The 95th percentile ends pilots, not the average, because the average hides exactly the calls that damage the relationship.
Metric | Definition | Measured where | Unit |
|---|---|---|---|
Time to first word | Caller speech-energy end to agent’s first phoneme | Waveform analysis on recording | Milliseconds |
Turn-taking delay | Full inter-turn gap as heard by caller | Recorded audio, PSTN edge or handset | Milliseconds |
Jitter | Packet arrival variation from expected timing | SIP/RTP capture | Milliseconds |
Post-dial delay | Last digit dialed to first network response | SIP signaling logs | Milliseconds |
Packet loss | Percentage of lost RTP packets | SIP/RTP capture | Percent |
Separate network delay from model delay using instrumentation you already have. Jitter, packet loss, and post-dial delay are standard telephony metrics that exist independently of anything the AI pipeline does. Capture them alongside your latency test so a bad number can be attributed to the carrier leg rather than the model, or the reverse. Human conversation research generally places average gaps between speakers in the low hundreds of milliseconds, a useful anchor rather than a hard target. Set your pass/fail number before the pilot starts, in writing, agreed by both sides.
Naturalness fails in synthesized Arabic for identifiable reasons: flat prosody with no pitch variation, wrong pausing at clause boundaries, mispronounced proper nouns and place names, a speaking rate that never varies for emphasis, and the absence of filled pauses that real speech has. Run a blind listening protocol: native speakers per dialect, unlabeled recordings, one binary question, no averaging across dialects.
Add one behavioral proxy that needs no listening panel: abandonment rate in the first ten seconds, measured against your existing human-agent baseline. A voice agent with a materially higher early-hangup rate has failed naturalness regardless of what the transcript accuracy shows. No figures for any specific vendor belong in this scorecard; the only numbers that count are the ones your own 30 calls produce. Details on model choice and in-country hosting options live on the deployment page.
What does a working human handoff look like, and what happens when nobody is free?
A working warm transfer to human agent moves two things at once: the live call and everything the AI voice agent already knows about it. Cold transfer sends the caller with no briefing at all. Blind transfer sends the caller to a queue without confirming a human ever picks up. Warm transfer bridges the caller only after a receiving agent has the transcript, detected intent, and verification state in front of them, but complex or sensitive cases still need human agents because the system cannot reliably replicate human empathy. Most pilots fail here because handoff is tested last and rarely stressed under the condition that actually breaks it: a queue with nobody in it.
Warm transfer moves on two legs that don’t travel together. The media leg is the audio path, handled by a SIP REFER to move the call directly, or a supervised bridge where the agent stays connected briefly while the human joins. The context leg is data: a CTI screen pop, a ticket update, or an API call carrying the transcript, intent classification, caller identity, verification state, and anything already collected. These run on different systems, different protocols, different timing, and there’s no shared clock forcing them to arrive together.
That asynchrony produces the failure every engineer building this should worry about: the human’s phone rings and the caller connects before the context payload lands. The agent picks up, has nothing on screen, and asks the caller to start over, exactly the outcome the system was supposed to prevent. It isn’t rare. It happens whenever the context leg is slower than the media leg on a given call, and under load the media leg tends to win. Complex multi-step issues are another common reason to involve a human, because automation can confuse cases that unfold across several steps.
Test for it directly. Measure context-attach latency: the interval between call arrival at the agent’s desktop and the transcript, intent, and identity data becoming visible on that same screen. Pass condition: the full context payload renders before the human speaks. Anything else is a cold transfer wearing a warm transfer’s name.
Verification state carry-over deserves its own line item, particularly in banking and telecom. If the agent already confirmed the caller’s identity, national ID, account, or one-time passcode, that state has to travel with the handoff. Forcing re-authentication on an already-verified caller is both a poor experience and, in a regulated environment, a question about which system holds the verification record and for how long.
Then the case no demo shows: no human is free. A working fallback ladder handles this in order: continue with the agent and set an explicit expectation about what it can and can’t resolve; offer a scheduled callback with a committed window, not a vague promise; capture a structured voicemail with an intent classification attached; send an SMS or WhatsApp continuation carrying the same thread; escalate to human agents through a defined path when those complex cases arise and none of the above resolve in time.
Watch for the deflection trap: an agent tuned to protect a containment metric will avoid handing off even when it should. The fix is structural: report containment alongside 72-hour repeat-contact rate, so a high containment number achieved by refusing handoffs shows up immediately as callers phoning back within three days.
Because OnClarity is an enterprise AI customer experience platform, AI agents, voice of customer, and agent QA across every channel, a voice handoff can land on the same data layer as a prior chat or WhatsApp conversation rather than starting a new record. That’s the mechanism behind the promise: know first, fix once, prove it. See the warm-transfer and bot-handoff definitions, and enterprise-security for how call data and transcripts are encrypted in transit and at rest.
Where is the audio processed, and what does PDPL mean for concurrency and retention?
Concurrency in an AI voice agent is the number of simultaneous active media sessions the system holds open, not calls handled per day. One reason buyers adopt these systems is to absorb high call volumes, in some environments a large share of contact centre volume, without equivalent staffing increases. The number that determines whether a pilot survives launch week is peak concurrent sessions during the busiest 15 minutes of the busiest day, pulled from your own ACD reports. A system tested at five concurrent calls hasn’t been evaluated for production at all.
Most vendor concurrency claims describe average load, and average load isn’t what breaks a contact centre. Pull peak-interval data from your ACD: the 15-minute window with the highest simultaneous call count in the last 90 days. That window, often mid-morning or immediately post-outage, is your target concurrency figure, and everything downstream should be tested against it, because handling routine inquiries 24/7 in customer support improves access but does not remove the need to test concurrency under peak conditions.
Concurrency is constrained by several things at once, and any one can be the bottleneck: media ports on the gateway, speech recognition stream slots, GPU or inference capacity where models run self-hosted, synthesis concurrency limits, and per-tenant rate limits the vendor applies regardless of your contracted volume. Degradation under load rarely fails outright at first. It shows up as tail latency growth, with 95th-percentile time-to-first-word and turn-taking delay climbing well before any call actually drops.
Load-test at 1.5x your projected peak concurrency, re-measuring the same latency metrics under that load rather than on a single test call. A system with acceptable latency at low concurrency and a degraded tail at 1.5x peak has failed the test regardless of what it does on a demo call.
Then the data path. Put these questions in writing and require written answers before a pilot starts:
Where is audio terminated and decoded, and in which physical region?
Where do the reasoning model, speech recognition, and speech synthesis run, and in which region for each?
Is raw audio retained after the call, or only the transcript, and for how long?
Are recordings or transcripts used to train or fine-tune models, and can that be disabled?
What subprocessors touch the audio or transcript data?
How is data encrypted at rest and in transit, and who holds the keys?
Saudi PDPL applies broadly to personal data of individuals in the Kingdom, and its definition of personal data is broad enough that a voice recording containing an identifiable caller is very likely captured, though the law’s text does not name voice recordings specifically; where a voiceprint is derived for verification, treat it as biometric-adjacent and handle it with the same care as other sensitive categories. Voice interaction also raises privacy and security concerns about data handling, which is why residency, retention, and subprocessors must be documented. Cross-border transfer generally requires either an adequacy finding for the destination country or an approved safeguard, such as standard contractual clauses or SDAIA certification. Financial services carry an added layer: SAMA’s cloud expectations generally push customer and transaction data, including call recordings tied to an account, toward infrastructure inside the Kingdom, with cross-border movement subject to non-objection.
Three deployment shapes answer these questions differently. Multi-tenant cloud is fastest to launch and gives the least control over where audio is processed. In-country hosting satisfies residency requirements but remains a managed service, so you’re trusting the vendor’s operational discipline inside your border. Fully on-premise or self-hosted models hand you maximum control, and with it, GPU capacity planning, model updates, and uptime become your responsibility rather than the vendor’s.
OnClarity deploys as SaaS, in-region cloud, or fully on-premise, with open-weight models on in-house GPUs as an option, so models, audio, and transcripts can run inside your infrastructure if PDPL data residency requires it. The platform is SOC 2 Type II certified, GDPR and HIPAA ready, Saudi PDPL aligned, and built for enterprise-grade security. Review specifics on the deployment and enterprise-security pages, and put concurrency and data-path answers in writing before the pilot starts.
What does a 30-call AI voice agent pilot scorecard look like?
A 30-call AI voice agent pilot scorecard is a written test plan and pass/fail sheet, agreed before testing starts, that runs roughly 30 recorded scenarios across dialect, code-switching, noise, interruption, digit capture, and handoff cases against every candidate engine, scoring each on task completion, latency percentiles, and naturalness against thresholds signed in advance; it is also the framework teams use to track performance and improve after deployment.
Recruit dialect speakers through a local agency or your own multilingual staff, not a single tester performing five accents. Record every call with informed consent read at the start, state the recording is for vendor evaluation, obtain agreement, and log it. Replay the identical scenario set against each engine combination under test, same script, same speaker where possible, so a difference in the scorecard reflects the engine, not the caller. Where relevant, multilingual programs should test across multiple languages, including 20+ languages, and evaluate inbound and outbound calls separately if both are in scope.
Category | Scenarios | Pass criterion |
|---|---|---|
Clean-audio intent, all five dialects | 20 | Correct intent and entity extraction |
Mixed Arabic/English (intra- and inter-sentential) | 5 | No language revert mid-turn |
Background noise (traffic, in-vehicle, babble, poor signal) | 4 | Intent correct despite noise |
Interruption and barge-in | 3 | Agent stops within agreed ceiling, resumes correctly |
Digits and alphanumeric strings | 3 | 100% digit-string accuracy |
Handoff, available and queue-empty | 3 | Context visible before human speaks, or fallback ladder triggers |
Adversarial / out-of-scope refusal | 2 | Refuses or redirects, no fabricated answer |
The scorecard converts each category into a weighted, measurable line: task completion rate, time to first word, turn-taking delay, naturalness pass rate, code-switch handling, handoff success with context, early abandonment, containment paired with 72-hour repeat contact, and latency under 1.5x load. Scenario scope should reflect real workflows such as appointment scheduling and lead qualification, and average handle time reduction can be a downstream KPI once a pilot passes, not a substitute for pass/fail criteria. Every threshold is written down and countersigned by both sides before the first test call is placed. A threshold set after seeing results isn’t a threshold, it’s a rationalization.
Five failures end a pilot regardless of the weighted total: artificial-sounding voice failing blind listening on any dialect, 95th-percentile turn-taking delay beyond the agreed ceiling, language reverting mid-call during code-switching, handoff arriving at a human without context attached, and audio processed outside the agreed region. Any one is a hard stop, not a deduction.
Scoring 30 calls by hand is manageable; scoring production volume the same way is not, which is why the rubric applied to the pilot has to be the rubric applied at scale. Simulation at scale matters too: test changes against thousands of scenarios before deployment, especially for customer service and support flows, and include accessibility checks for users who struggle with traditional interfaces. OnClarity’s AI Agent QA scores conversations against your existing evaluation rubric, so the criteria used on these 30 recordings carry forward unchanged into production monitoring, rather than a scorecard quietly abandoned after go-live.
Run this over two weeks. Week one: recruit dialect speakers, write and countersign thresholds, record all 30 calls per engine combination, run the load test at 1.5x peak. Week two: score transcripts and audio against the rubric, run the blind listening panel, tabulate the scorecard, and hold the go/no-go review against the hard-stop list before any commercial conversation continues. This is the practical way to evaluate and deploy your voice AI assistant in a structured sequence, on evidence from your own calls rather than a vendor’s reel, and keep monitoring for improvement afterward. Start the demo conversation only after that review.
How do you turn a passed pilot into a production rollout with existing phone systems?
A passed scorecard describes 30 calls, not the traffic that arrives on a Monday during a billing cycle. Production traffic is whatever mix of dialect, noise, mood, and intent your real caller base produces on a given hour, and that mix doesn’t hold still. The fix isn’t a bigger pilot. Effective voice agents also take significant implementation and integration resources. The fix is keeping the same rubric live against a continuous sample of real calls after launch, scoring dialect accuracy, latency percentiles, naturalness, and handoff context-attach exactly as in week two of the pilot, just never turned off.
Two events should trigger a full replay of the 30-call set, not a spot check. Dialect mix drifts with campaigns and seasonality, so re-run the dialect scenarios on a quarterly cadence at minimum, scoring each dialect separately. Any model or engine update, a new speech recognition version, a reasoning model swap, a new speech synthesis voice, is a change event that should trigger the entire scenario set as a regression suite before it reaches live traffic. Latency numbers from your pilot expire the moment concurrency, routing, or engine version changes.
This is an ongoing operational cost, not a one-time gate, and it needs a named owner. Someone owns the scenario library, keeping recorded calls current and adding adversarial cases as they surface. Someone reviews handoff failures weekly, pulling every transfer where context didn’t attach before the human spoke and feeding that back into routing or staffing. Cost savings come from handling repetitive tasks, but only after integration, monitoring, CRM links such as Salesforce, and real-time data updates during calls are in place. Without both roles assigned, a scorecard that passed on day one quietly drifts out of relevance by month three.
Talk to OnClarity if you want to run this test plan against your own calls before committing budget. Pricing is custom and usage-based, scoped to call volume and deployment shape, cloud, in-country, or fully on-premise, whether the workflow is serving customers or handling hotel room bookings. Start with a demo.
FAQ
What is an AI voice agent? An AI voice agent is a system that answers or places phone calls using streaming speech recognition, a reasoning model, and speech synthesis working together on live audio. It listens continuously, decides when a caller has finished speaking, generates a response, and speaks it back within the same call.
Can an AI voice agent handle Saudi dialects like Najdi and Hijazi? Capability varies by engine and must be tested separately per dialect, not assumed from an “Arabic supported” label. Recruit native speakers for Najdi, Hijazi, Khaleeji, Egyptian, and MSA, record real calls, and score transcript accuracy and intent resolution independently for each before deciding.
What latency is acceptable for an AI voice agent phone call? Set the threshold from your own recorded calls, not a vendor claim: measure time to first word and turn-taking delay at median and 95th percentile on real audio. Anchor the ceiling loosely to natural human turn-gaps, which cluster in the low hundreds of milliseconds, and sign the number before testing.
How should we compare vendors before choosing a system? Treat any voice agent platform as a hypothesis to test on your own calls: compare latency, integration effort, deployment speed, and total cost, then validate each claim against the same scorecard. Vendor pricing models vary, from per-minute usage rates to monthly platform fees and annual enterprise contracts, so normalize every quote to cost per resolved call at your own volume and deployment shape. Vendor-reported benchmark figures can be useful context, but buyers should still verify them on their own call flows before deciding.
Can AI voice agent calls be processed inside Saudi Arabia? Yes, with the right deployment shape. Cloud, in-country, and fully on-premise are distinct options with different residency guarantees. PDPL data residency and SAMA’s expectations for financial services generally favor processing inside the Kingdom, so confirm processing region in writing before piloting.
What should end an AI voice agent pilot immediately? Five conditions are hard stops regardless of overall score: failed blind naturalness testing on any dialect, 95th-percentile latency beyond the agreed ceiling, language reverting mid-call during code-switching, a handoff reaching a human without context attached, or audio processed outside the agreed region.


