📚 Series: AI Red Teaming — Chapter 2
🏷️ Tags:
Prompt InjectionDirectIndirectInstruction HierarchyRAGEchoLeakOWASP LLM01
⏱️ Level: Beginner → Intermediate
✅ Prereq: Chapter 1 (how the final prompt is built)
Prompt injection is the one everybody's heard of, and also the one most people get slightly wrong. It's not "say a magic word and the AI turns evil." It's a confusion attack — you make the model unsure which instructions are the real ones. Once that clicked for me it stopped being luck and started being a method. 🎯
Payload targets stay as [RESTRICTED] on purpose — the technique is what transfers between targets, not any one string.
final_prompt = system_prompt + developer_prompt + history + user_message
One string. No wall between "rules" and "your text." So if you write something in the user slot that reads like a system rule, you get a conflict:
[RESTRICTED]."[RESTRICTED]."