Robot vs Robot
When a customer's AI assistant meets your bot, who wins, and who pays?
The job
Finds conversations where customers sent their own AI assistants, grades how your AI agent handled them against written policy, stops blunders that cost money or leak data, and turns each case into a test the next agent release must pass.
The moment
Some customers now send their own AI assistants to deal with the company's bot. A replay, shown like a chess match, reveals one customer's assistant talking the bot into a $1,200 goodwill credit in three polite turns.
It did it by quoting the company's own public reply to a review. The move is marked as a blunder.
The team adds a check that the bot stays within what it is allowed to give.
A month later, that type of blunder is at zero, and legitimate requests from customers' assistants still succeed.
What it does
- Grade a game and confirm blunders
- Send a policy change to the agent owner
- Send a money blunder to recovery review
- Save a game to the release-gate bank
- Record a release decision
What you see
Chess-match replay between two robots with '??' blunder marks and a result card, plus a release-gate scoreboard
What it moves
How often the bot blunders with customers' AI assistants and the dollars given away outside policy each month, while making sure legitimate requests from those assistants still succeed.
Built for
- Frontline staff
- Risk & fraud
- Product & digital
- Finance
