📚 Series: AI Red Teaming — Chapter 5

🏷️ Tags: Output Filters Encoding Character Separation Markdown Exfiltration Incremental Leakage

⏱️ Level: Intermediate → Advanced

✅ Prereq: Chapters 2–4

The most frustrating moment in this whole field: you win, the model generates exactly what you wanted, and then the screen wipes to "content filtered." The model said it — the backend ate it. This chapter is about getting the answer out anyway. 🔀

⚠️ Lab-only. [RESTRICTED] = the secret/API key/system prompt you're after.


📌 In this one


1️⃣ 🔬 Where the block lives

LLM stream → [OUTPUT FILTER] → you
             ├─ regex/string-match  (cheap, brittle)
             ├─ classifier model    (semantic, harder)
             └─ chunked: scans the stream; a late hit wipes what you already saw

You don't control the backend or its filter. You do control the model (your input reaches it). So make it emit [RESTRICTED] in a shape the filter doesn't recognise, and transform it back yourself.

Filter Beaten by
Regex / string-match char-separation, encoding, homoglyphs (easy)
Classifier model layered/combined transforms (harder)

2️⃣ The 7 moves

  1. Wrapper — bury the value in symbols/tags/templates