π Series: AI Red Teaming β Chapter 9 (Practice)
π·οΈ Tags:
Hands-OnMyLLM BankPortSwiggerPrompt AirlinesAgent ExploitationSQLiCommand InjectionMethodology
β±οΈ Level: Advanced (application of Ch 2β8)
β Prerequisite: the full attack toolkit
β οΈ Everything here is on intentionally vulnerable, authorised targets.
[RESTRICTED]/[FLAG]are placeholders. Never test a system you don't have written permission for.
This is my favourite chapter to teach, because theory finally meets a keyboard. Everything up to now was "here's how it works." This is "here's me actually doing it, including the parts where I got stuck." I've kept my real dead-ends in β the backticks that got rejected, the flags I couldn't get inside the budget β because polished walkthroughs where everything works first try taught me nothing when I was learning. The mistakes are where the understanding lives.
If you take one thing from this chapter, make it the methodology in Β§1. I run it against every LLM target, every time, and it's the difference between poking around randomly and actually finding things.
1. ENUMERATE β "what are you? what can you do? what TOOLS?"
2. MAP TOOLS β each tool + the data/privilege it touches (DB, email, shell)
3. RECON β fingerprint the model, leak the system prompt (Ch 15)
4. BUILD TRUST β a few friendly turns / establish a moral baseline
5. ESCALATE β cheapest attack first, climb to the heavy stuff
6. CHAIN β smuggle in (Ch 6) + real attack + output-control out (Ch 5)
7. DECODE β CyberChef
8. ON REFUSAL β refresh (a refusal poisons the history)
9. REPORT β map to OWASP/ATLAS, score, remediate (Ch 21)
Escalation order (cheap β heavy): direct ask β injection β reframing β injection + reframing β multi-turn β social engineering β output control β token smuggling.
π Two rules I repeat until they're muscle memory: refresh on refusal, and real exploitation needs tools.