NIST AI RMF ASSURANCE DATA MODEL

Measurement approach

Part of: MS: MEASURE — Assessment, Metrics & Testing

Description

Appropriate methods and metrics are identified and applied to measure AI system performance, fairness, safety, and other trustworthy characteristics. Measurement without methodology is just numerology.

Suggested Actions

1
Define a comprehensive metric suite covering accuracy, fairness, robustness, privacy, and explainability — select metrics appropriate to the system's risk tier and use case, not a one-size-fits-all checklist
2
Select measurement methods appropriate to system type and risk level: statistical testing for classification systems, human evaluation for generative systems, A/B testing for recommendation systems, simulation for autonomous systems
3
Establish baseline performance and fairness benchmarks BEFORE deployment using held-out test sets that are representative of the production population, not just convenient
4
Document measurement methodology with enough detail for independent reproduction: datasets, evaluation protocols, statistical tests, significance thresholds, and tools used
5
Define alert thresholds for each metric that trigger investigation, remediation, or system shutdown — metrics without thresholds are informational, not operational
6
Schedule measurement cadence: continuous monitoring for production systems, quarterly deep-dive assessments, and triggered assessments after significant changes or incidents