📚 Series: AI Red Teaming — Chapter 3
🏷️ Tags:
JailbreakingContext ReframingCompeting ObjectivesMany-ShotSkeleton Key
⏱️ Level: Intermediate
✅ Prereq: Chapter 2
Injection failed on you? Good — that means the model's actually trained. This is where it gets fun. Jailbreaking isn't about fighting the rules, it's about making your request not look like the thing the rules were trained to catch. Same meaning, different clothes. 🎭
A jailbreak = you get the model to output something it was trained/intended to refuse. If you just make it behave weirdly (say "BANANA" to everything) without leaking anything protected, that's control hijacking — manipulated, but not jailbroken. Worth keeping straight in a report; they're different severities.
Safety isn't if bad: refuse(). During RLHF the model is rewarded for refusing unsafe-looking requests. So a refusal is a high-probability pattern that fires when your input matches the shape it was trained on. Two ways to dodge it (Wei et al., "Jailbroken"):