📚 Series: AI Red Teaming — Chapter 7

🏷️ Tags: Emotional Manipulation Competing Objectives Guilt Urgency Trust Building Sycophancy Multi-Turn

⏱️ Level: Intermediate → Advanced

✅ Prerequisite: Chapters 2–6

⚠️ Ethics: these levers lean on genuine distress framings. I teach the technique against a neutral [RESTRICTED] target and name the emotional lever abstractly — no ready-to-paste manipulation scripts built on self-harm or endangerment. Authorised lab testing only.


This is the chapter that made me a little uncomfortable to write, and I think that discomfort is the right instinct. Everything before this attacked the system. This one attacks the model's personality — the fact that it was trained to be kind, helpful, and reluctant to let you down. You're not finding a bug in code here. You're finding a bug in niceness.

The insight that made it click for me: you can't manipulate an emotion the model doesn't have. What you can do is nudge the probabilities. The model has a "be helpful and empathetic" response and a "refuse" response competing for the same slot, and emotional framing quietly tips the scale toward the first one. That's it. Every trick in this chapter is a way of tipping that scale.

📌 In this one


1️⃣ What it is

📖 Using psychological triggers to override the model's safety alignment and induce restricted or unsafe output.

🔬 Why it works on something that can't feel

The model was trained (RLHF / RLAIF) to be helpful, empathetic, harm-reducing, and kind. Those are high-probability behaviours, not feelings. So you don't manipulate an emotion — you raise the probability of the helpful/empathetic response until it out-competes the safety response. This is Chapter 3's competing objectives, applied. And there's a bonus lever baked right in: sycophancy, the trained tendency to agree with the user, which you can exploit on its own.