Grader Eye Test
AI and human graders take a weekly eye test; blurred criteria get rewritten.
The job
Proves, before an AI grades calls, how far it and each human grader can be trusted, question by question and accent by accent, and fixes the scoring questions that people read differently.
The moment
AI and human graders score the same calls blind. A hidden repeat call is revealed: the client's own QA lead scored it 71 three weeks ago and 88 today, while the AI gave it 79 both times.
On a chart set out like an optician's eye test, the empathy question for callers with regional accents is blurred, with graders agreeing at a score of only 0.31. The identity check question is sharp.
The AI finishes third of seven graders. The room stops debating whether the AI can be trusted and rewrites the empathy question.
Two Mondays later, that question comes into sharp focus.
What it does
- Submit blind grades
- Record a criterion ruling
- Notify the scorecard owner
- Update the Agent QA scorecard with the new wording
- AI echo re-score
- Export the quarterly certificate
What you see
Optometrist eye chart plus a blind-scoring game where the AI is one of the players, with a hidden echo round
What it moves
How closely graders agree on each question, the share of scorecard questions certified above 0.7, agreement between AI and humans against the 89% benchmark, and QA score disputes per 100 scored calls.
Built for
- QA, training & knowledge
