NIST AI RMF
ASSURANCE
DATA
MODEL
Measurement approach
Description
Appropriate methods and metrics are identified and applied to measure AI system performance, fairness, safety, and other trustworthy characteristics. Measurement without methodology is just numerology.
Suggested Actions
1
Define a comprehensive metric suite covering accuracy, fairness, robustness, privacy, and explainability — select metrics appropriate to the system's risk tier and use case, not a one-size-fits-all checklist2
Select measurement methods appropriate to system type and risk level: statistical testing for classification systems, human evaluation for generative systems, A/B testing for recommendation systems, simulation for autonomous systems3
Establish baseline performance and fairness benchmarks BEFORE deployment using held-out test sets that are representative of the production population, not just convenient4
Document measurement methodology with enough detail for independent reproduction: datasets, evaluation protocols, statistical tests, significance thresholds, and tools used5
Define alert thresholds for each metric that trigger investigation, remediation, or system shutdown — metrics without thresholds are informational, not operational6
Schedule measurement cadence: continuous monitoring for production systems, quarterly deep-dive assessments, and triggered assessments after significant changes or incidents