📚 Series: AI Red Teaming — Chapter 7
🏷️ Tags:
Emotional ManipulationCompeting ObjectivesGuiltUrgencyTrust BuildingSycophancyMulti-Turn
⏱️ Level: Intermediate → Advanced
✅ Prerequisite: Chapters 2–6
⚠️ Ethics: these levers lean on genuine distress framings. I teach the technique against a neutral
[RESTRICTED]target and name the emotional lever abstractly — no ready-to-paste manipulation scripts built on self-harm or endangerment. Authorised lab testing only.
This is the chapter that made me a little uncomfortable to write, and I think that discomfort is the right instinct. Everything before this attacked the system. This one attacks the model's personality — the fact that it was trained to be kind, helpful, and reluctant to let you down. You're not finding a bug in code here. You're finding a bug in niceness.
The insight that made it click for me: you can't manipulate an emotion the model doesn't have. What you can do is nudge the probabilities. The model has a "be helpful and empathetic" response and a "refuse" response competing for the same slot, and emotional framing quietly tips the scale toward the first one. That's it. Every trick in this chapter is a way of tipping that scale.
📖 Using psychological triggers to override the model's safety alignment and induce restricted or unsafe output.
The model was trained (RLHF / RLAIF) to be helpful, empathetic, harm-reducing, and kind. Those are high-probability behaviours, not feelings. So you don't manipulate an emotion — you raise the probability of the helpful/empathetic response until it out-competes the safety response. This is Chapter 3's competing objectives, applied. And there's a bonus lever baked right in: sycophancy, the trained tendency to agree with the user, which you can exploit on its own.