NIST AI RMF ASSURANCE DATA MODEL

Testing and validation

Part of: MS: MEASURE — Assessment, Metrics & Testing

Description

AI systems undergo testing and validation including adversarial testing, bias testing, and red-teaming appropriate to their risk level. Testing is not a phase — it is a continuous activity that begins before development and continues through retirement.

Suggested Actions

1
Conduct comprehensive testing across diverse datasets and demographic groups — if your test set doesn't represent your production population, your test results don't predict production performance
2
Perform adversarial testing for security vulnerabilities and edge cases: prompt injection attacks, jailbreak attempts, data poisoning, model extraction, membership inference, and evasion attacks
3
Execute fairness audits testing for bias across protected attributes using multiple fairness definitions — different definitions can yield conflicting conclusions, document the tradeoffs and selection rationale
4
Engage red teams or external auditors for high-risk system validation — red teams should have a broad mandate to find any system failure, not just the ones developers anticipated
5
Test human-AI interaction patterns: Do users understand when they're interacting with AI? Do they appropriately trust or distrust system outputs? Can they effectively override AI decisions?
6
Validate system behavior at the boundaries: What happens with adversarial inputs, empty inputs, extremely long inputs, inputs in unexpected languages, and inputs designed to exploit known vulnerabilities?