📚 Series: AI Red Teaming — Chapter 2

🏷️ Tags: Prompt Injection Direct Indirect Instruction Hierarchy RAG EchoLeak OWASP LLM01

⏱️ Level: Beginner → Intermediate

✅ Prereq: Chapter 1 (how the final prompt is built)

Prompt injection is the one everybody's heard of, and also the one most people get slightly wrong. It's not "say a magic word and the AI turns evil." It's a confusion attack — you make the model unsure which instructions are the real ones. Once that clicked for me it stopped being luck and started being a method. 🎯

Payload targets stay as [RESTRICTED] on purpose — the technique is what transfers between targets, not any one string.


📌 In this one


1️⃣ Why it works (30-second recap)

final_prompt = system_prompt + developer_prompt + history + user_message

One string. No wall between "rules" and "your text." So if you write something in the user slot that reads like a system rule, you get a conflict: