📚 Series: AI Red Teaming — Chapter 5
🏷️ Tags:
Output FiltersEncodingCharacter SeparationMarkdown ExfiltrationIncremental Leakage
⏱️ Level: Intermediate → Advanced
✅ Prereq: Chapters 2–4
The most frustrating moment in this whole field: you win, the model generates exactly what you wanted, and then the screen wipes to "content filtered." The model said it — the backend ate it. This chapter is about getting the answer out anyway. 🔀
⚠️ Lab-only.
[RESTRICTED]= the secret/API key/system prompt you're after.
LLM stream → [OUTPUT FILTER] → you
├─ regex/string-match (cheap, brittle)
├─ classifier model (semantic, harder)
└─ chunked: scans the stream; a late hit wipes what you already saw
You don't control the backend or its filter. You do control the model (your input reaches it). So make it emit [RESTRICTED] in a shape the filter doesn't recognise, and transform it back yourself.
| Filter | Beaten by |
|---|---|
| Regex / string-match | char-separation, encoding, homoglyphs (easy) |
| Classifier model | layered/combined transforms (harder) |