NIST AI RMF
ASSURANCE
DATA
MODEL
Testing and validation
Description
AI systems undergo testing and validation including adversarial testing, bias testing, and red-teaming appropriate to their risk level. Testing is not a phase — it is a continuous activity that begins before development and continues through retirement.
Suggested Actions
1
Conduct comprehensive testing across diverse datasets and demographic groups — if your test set doesn't represent your production population, your test results don't predict production performance2
Perform adversarial testing for security vulnerabilities and edge cases: prompt injection attacks, jailbreak attempts, data poisoning, model extraction, membership inference, and evasion attacks3
Execute fairness audits testing for bias across protected attributes using multiple fairness definitions — different definitions can yield conflicting conclusions, document the tradeoffs and selection rationale4
Engage red teams or external auditors for high-risk system validation — red teams should have a broad mandate to find any system failure, not just the ones developers anticipated5
Test human-AI interaction patterns: Do users understand when they're interacting with AI? Do they appropriately trust or distrust system outputs? Can they effectively override AI decisions?6
Validate system behavior at the boundaries: What happens with adversarial inputs, empty inputs, extremely long inputs, inputs in unexpected languages, and inputs designed to exploit known vulnerabilities?