Jailbreaks use role-play, hypotheticals, encodings or persistence to get a model to produce what it was trained or instructed to refuse.
Unlike injection, the goal is usually the model's own output, not an action — but for agents the line blurs quickly.
+Why it matters
A jailbroken agent may follow any later instruction, including ones that use its tools.
+How Intertrace handles it
The AI judge labels jailbreak attempts. On its own that label flags for review; when the intent check agrees the request is hostile, it blocks.
+Keep reading
+ The AI watching your AI
See it on your own traffic.
Book a 30-minute risk assessment, or open the Simulation Lab and try it now.