Evaluation set
Related terms
Quality assurance
Quality assurance is the structured review of conversations against a scorecard, applied to human agents and AI agents alike. Traditional QA samples a handful of conversations per agent each month; automated QA scores every conversation, which changes the exercise from spot-checking into actual measurement.
Model drift
Model drift is the quiet degradation in AI performance that occurs as products, policies, and customer language change while the model and its knowledge stay fixed. Drift rarely announces itself; it appears as a slow rise in escalations and fallbacks on intents that previously performed well.
Hallucination
A hallucination is a fluent but unsupported answer produced when a model is not grounded in real data. In customer service the risk is specific: an invented policy, price, or delivery date reads as authoritative to the customer and creates a commitment the business never made.
Utterance
An utterance is a single thing a customer said, exactly as the model receives it, before any cleaning or interpretation. Utterances are the raw material for intent classification and evaluation sets, and preserving them verbatim is what allows AI errors to be reproduced and diagnosed later.

Support teams are no longer a cost centre.
See how AI Support Agent, Agent Assist, and VoC Analyst work on your own ticket volume.