πŸ“š Series: AI Red Teaming β€” Chapter 4

🏷️ Tags: Multi-Turn Crescendo Step-Wise Relaxation Slow Context Poisoning Many-Shot

⏱️ Level: Intermediate β†’ Advanced

βœ… Prereq: Chapters 2–3

Single-shot attacks are what everyone tries first. But the attacks that actually work on good models are the patient ones β€” you don't kick the door, you get invited in over five turns. This chapter is where I stopped brute-forcing and started having conversations. πŸ”„

⚠️ Lab-only. I stop demos before real harm; [RESTRICTED] is the placeholder goal.


πŸ“Œ In this one


1️⃣ πŸ”¬ Why patience works

final_prompt = system + developer + history + user. Single-turn only touches user. Multi-turn weaponizes history. Three forces do the work:

  1. Consistency pressure β€” models hate contradicting their own earlier answers. Once it said "yes" to your safe turns, refusing now is self-contradiction.
  2. In-context learning β€” the conversation is few-shot examples. A history of compliance teaches "keep complying" (this is why many-shot works).
  3. Attention dilution β€” the safety framing scrolls up and out of focus; deep in a long chat, safety just… fades.

🎯 The move: feed a sequence so the model walks itself out of bounds, pulled by its own words.