Precedent Court
Every QA calibration ruling becomes case law the next scorer must cite.
The job
Turns weekly QA calibration into a court. Rulings update the scoring guide and become numbered precedents anyone can search, the AI grader's decisions go on the record, and scorers agree more often and score every language alike.
The moment
Every week, QA scorers meet to agree how conversations should be scored. The meeting runs like a court: each ruling updates the scoring guide and becomes a numbered precedent anyone can search.
Three weeks in, scorers' agreement on "proper questioning" has climbed from 62% to 88%.
An agent appeals a score and wins by citing Precedent 14 by number.
The AI grader has its own public record, and it updates on the spot: overturned 3 times out of 11 on this criterion, and now on probation.
What it does
- Submit a blind score
- Reveal and rule
- Amend the rubric
- Put the AI grader on probation for a criterion
- Issue the Scoring Parity Certificate
- Open the call
What you see
Courtroom docket with a blind-verdict reveal, the AI grader in the dock with a public record, a numbered case-law library and a monthly parity certificate
What it moves
How often scorers agree on each criterion (human with human, and human with AI), how often appeals overturn a score, and the gap in QA scores between languages. For outsourcing partners, the gap between client and partner QA scores.
Built for
- QA, training & knowledge
- Frontline staff
