Section 038 · Chapter 6, Building Evals That Matter
Adversarial and Red-Team Sampling
Random samples estimate normal behavior. Adversarial samples reveal what happens when users push the system.
RAGprompt injectionadversarial red team sampling
What to do
- Use approved test tenants, test accounts, allowlisted IPs, written rules of engagement, and internal contacts who know the work is authorized.
- Treat red-team execution like security testing.
- Define runnable checks that exercise RAG, prompt injection, and adversarial red team sampling.
Evidence to preserve
- Preserve the inputs, versions, configurations, raw outcomes, and results for RAG, prompt injection, adversarial red team sampling needed to reproduce work on Adversarial and Red-Team Sampling.
- Report results for RAG, prompt injection, adversarial red team sampling by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.
Expert note
Expert red-team programs track attack family, severity, exploitability, reproducibility, affected surface, mitigation status, and whether the same attack reappears after a prompt, policy, model, tool, or retriever change. They also refresh attacks frequently because users and attackers adapt once a system is deployed.
Continue the conversation
Apply this to your context.
Save your product context once, then open a focused conversation that combines it with this concept.
Cite this page
Jason Arbon. "Adversarial and Red-Team Sampling." Testing AI Knowledge Edition, section 38.
https://jarbon.ai/testing-ai/knowledge/ch038-adversarial-red-team-sampling.html