Rubber Stamp Detector
Proves the human in the loop is really looking, with planted test errors.
The job
Shows a regulator that human checking of AI decisions is real, not a click-through. It fixes the teams where it isn't, and gives AI steps more freedom to act alone only where humans demonstrably check.
The moment
The app tests whether people checking AI decisions are really looking, by planting errors. It is set out like an airport security scanner, and the executive goes first.
Ten AI-drafted replies slide past on the belt for approval. The ninth has invented a 30-day refund window.
The belt stops and glows red: the executive approved it in 2.1 seconds. Team B let 9 of 10 planted errors through, at a median of 1.8 seconds.
Three weeks after a change to how the team works, the same team catches 8 of 10, and the compliance officer files both results.
What it does
- Plant a logged test round
- Switch a team to evidence-first approval
- Record a result
- Recommend autonomy for an AI step
- Open the replies behind a missed plant
What you see
Airport screening lane: approvals slide along an X-ray belt, planted errors glow when missed, and leaders screen themselves first
What it moves
The share of planted errors each team catches (target 80% or more, within a range), and the number of AI steps whose freedom to act alone is backed by measured human checking.
Built for
- Compliance, complaints & legal
- Contact-centre operations
