Jailbreak Watch
Real customers try to trick your bot daily. See who got through.
The job
Catches real customers trying to trick any AI agent, through planted instructions, role-play or social engineering, and shows whether each attempt worked. When a screenshot goes public, finds the original chat, fixes it and has a holding reply ready before the post spreads.
The moment
A screenshot of the company's bot agreeing to a one-dollar deal starts spreading online. The app catches it at 40 reposts.
A race screen pits the post's reach against the team's progress on the fix.
The team confirms the original conversation, files a fix to the bot's safety rules and drafts a holding reply.
The reach meter stops at 212, contained before it reaches 400.
What it does
- File a guardrail case or vendor security ticket
- Route to fraud
- Confirm the public-to-private match
- Refresh the post's reach
- Draft the holding reply
What you see
A CCTV security wall with HELD / BREACHED / PATCHED stamps and a race screen where a reach meter runs against a fix-progress bar
What it moves
How often each type of attack gets through, and the hours from a public post to containment: conversation found, fix filed and holding reply ready.
Built for
- Risk & fraud
- Frontline staff
