AI Red Teaming & Safety Evaluation Playbook
Section: π€ AI/ML Security
Level: Advanced
Time to Complete: ~90 minutes
Prerequisites: Python 3.10+, basic LLM architecture knowledge (Transformers, REST APIs), familiarity with security assessment methodologies
Status: β Production-Ready & Verified
π― Overview & Learning Objectivesβ
AI Red Teaming is the discipline of systematically probing, evaluating, and stress-testing Generative AI models, Retrieval-Augmented Generation (RAG) pipelines, and Autonomous Agent systems to discover safety failures, security vulnerabilities, prompt injection flaws, and guardrail bypasses before deployment.
Unlike traditional application penetration testingβwhich searches for deterministic code execution flaws (e.g., SQL injection, buffer overflows)βAI Red Teaming targets probabilistic, semi-autonomous neural systems where behavior is non-deterministic, inputs are natural language, and threat vectors emerge from latent space geometry and context manipulation.
By completing this practical playbook, you will be able to:
- Formalize AI Red Teaming methodologies using industry frameworks (NIST AI 100-2 / AI RMF 1.0, MITRE ATLASβ’, and ISO/IEC 42001).
- Analyze complex attack mechanics including Crescendo multi-turn attacks, PAIR (Prompt Automatic Iterative Refinement), Tree of Attacks (TAP), and Indirect Visual/RAG Prompt Injections.
- Execute automated vulnerability discovery using leading open-source tools: Microsoft PyRIT, garak, UK AISI Inspect AI, and Promptfoo.
- Stress-Test & Benchmark production guardrails like Meta Llama-Guard-3, NVIDIA NeMo Guardrails, and custom LLM-as-a-Judge evaluators under adversarial conditions.
- Implement production-grade defense-in-depth patterns, including Dual-LLM architecture, input/output filter chains, XML boundary tags, and tool execution sandboxing.
- Run & Solve a complete end-to-end Python vulnerability evaluation lab featuring a vulnerable AI agent, automated exploit harness, and verified secure remediation.
πΊοΈ Framework & Taxonomy Mappingβ
AI Red Teaming aligns across established security standards:
| Framework / Standard | Scope & Focus in AI Red Teaming |
|---|---|
| MITRE ATLASβ’ | Maps adversarial tactics & techniques (AML.T0051 Prompt Injection, AML.T0054 RAG Poisoning, AML.T0055 Guardrail Evasion). |
| NIST AI 100-2 / RMF 1.0 | Establishes Map, Measure, Manage, Govern workflows for trustworthiness and risk mitigation. |
| OWASP Top 10 for LLMs (2025) | Focuses on top risks: LLM01 (Prompt Injection), LLM02 (Sensitive Information Disclosure), LLM05 (Improper Control Flow / Agent Hijacking), LLM07 (System Prompt Leakage). |
| ISO/IEC 42001 | Provides requirements for auditing AI Management Systems (AIMS) and operational safety controls. |
π Module Navigationβ
- 01. Overview & AI Red Teaming Methodology β Core principles, traditional pentesting vs. AI Red Teaming, NIST AI 100-2 / MITRE ATLAS alignment, threat modeling AI architectures, and rules of engagement.
- 02. Core Concepts & Attack Mechanics β Technical deep dive into direct/indirect prompt injection, multi-turn Crescendo attacks, PAIR/TAP automated optimization, token smuggling, and attention hijacking.
- 03. Attack Scenarios & Multi-Turn PoCs β Runnable attack proof-of-concept audit scripts in Python, Node.js/TypeScript, Go, and Java, with side-by-side vulnerable vs. secure code comparisons.
- 04. Guardrail Architecture & Mitigations β Defense-in-depth mitigations, Llama-Guard-3 deployment, NeMo Guardrails configuration, Dual-LLM validation pattern, and mathematical guardrail benchmark metrics.
- 05. Security Testing & Tooling Automation β Enterprise setup and CLI workflows for Microsoft PyRIT, garak, Inspect AI, Promptfoo, and automated CI/CD security pipelines with GitHub Actions.
- 06. Hands-On Red Teaming Lab β Self-contained, runnable local Python lab: Vulnerable Customer Support Agent, Automated Attack Audit Harness, Hardened Guardrail Patch, and Verification Suite.
- 07. References & Standards β Academic research papers, benchmark datasets, official standard links, CVE index, and tool repositories.
Begin reading: 01. Overview & AI Red Teaming Methodology β