MEASURE — Assessment, Metrics & Testing
Description
The MEASURE function employs quantitative and qualitative methods to assess AI system trustworthiness characteristics including performance, fairness, security, resilience, transparency, and privacy throughout the lifecycle. MEASURE converts the risks identified by MAP into concrete, testable metrics and validates that controls implemented by MANAGE are actually working.
Subcategories
| ID | Name | Description |
|---|---|---|
| MS-1 | Measurement approach | Appropriate methods and metrics are identified and applied to measure AI system performance, fairness, safety, and other... |
| MS-2 | Testing and validation | AI systems undergo testing and validation including adversarial testing, bias testing, and red-teaming appropriate to th... |
| MS-3 | Competency and expertise | AI system measurement activities are performed by individuals and teams with appropriate domain knowledge, technical exp... |
| MS-4 | External inputs and validation | Measurement processes incorporate external perspectives, independent testing, and stakeholder feedback to validate AI sy... |
Implementation Guidance
Defining the Measurement Framework
Start by selecting metrics that map to each trustworthiness characteristic identified during MAP. Avoid measuring what's easy — measure what matters:
- Performance: Accuracy, precision, recall, F1, AUC-ROC — but always disaggregated by demographic group and use case segment
- Fairness: Demographic parity, equalized odds, predictive parity, calibration across groups — select metrics appropriate to the use case and document why
- Security: Adversarial robustness scores, prompt injection resistance rates, data extraction attempt detection rates
- Resilience: Performance under distribution shift, graceful degradation under load, recovery time after failures
- Transparency: Explainability scores (SHAP/LIME consistency), documentation completeness, user understanding rates
- Privacy: Re-identification risk scores, membership inference attack resistance, differential privacy budget tracking
Testing and Validation Methodology
Implement a layered testing approach:
- Unit testing — Individual model component validation
- Integration testing — End-to-end system behavior including pre/post-processing
- Fairness testing — Bias audits across protected attributes using standardized benchmarks
- Adversarial testing — Red team exercises, prompt injection testing, data poisoning resilience
- Stress testing — Performance under extreme load, edge cases, and out-of-distribution inputs
- User testing — Real-world usability, comprehension of AI outputs, and human-AI interaction patterns
Independent Validation
For high-risk AI systems, internal testing is necessary but insufficient. Engage external auditors, academic researchers, or independent testing organizations to validate your measurements. Internal teams have inherent blind spots and incentive misalignment.
Evidence Requirements
- Measurement framework document defining metrics for each trustworthiness characteristic with thresholds and rationale
- Test plans and validation reports for each AI system including methodology, datasets used, and results
- Fairness audit reports showing performance disaggregated across protected attributes with statistical significance analysis
- Adversarial testing and red team exercise reports with findings and remediation tracking
- Independent validation reports from external auditors or researchers for high-risk systems
- Measurement tool and methodology documentation enabling reproducibility
- Trend analysis showing metric trajectories over time with investigation reports for significant changes
- Competency assessments for measurement team members documenting relevant expertise