MG
NIST AI RMF APPLICATION GOVERNANCE INFRASTRUCTURE MODEL

MANAGE — Risk Response & Communication

Description

The MANAGE function allocates resources, implements controls, deploys mitigations, and establishes ongoing monitoring to manage AI risks identified by MAP and measured by MEASURE. MANAGE is where governance translates into operational reality — policies become controls, risk assessments become mitigations, and metrics become dashboards.

Subcategories

IDNameDescription
MG-1Purpose evaluationAI systems are evaluated to determine whether their intended purpose, use cases, and deployment context are appropriate ...
MG-2Risk and benefit balancingAI system risks and benefits are balanced and managed based on expected impact, with risk tolerance aligned to organizat...
MG-3Lifecycle monitoringAI systems are monitored throughout their lifecycle to detect performance degradation, drift, emerging risks, and uninte...
MG-4TEVV (Test, Evaluation, Verification, and Validation)TEVV processes are implemented and iterated throughout the AI lifecycle to ensure systems function as intended and align...

Implementation Guidance

Risk Treatment Planning

For each risk identified during MAP and quantified during MEASURE, select a treatment strategy:

  • Mitigate: Implement controls to reduce risk to acceptable levels (most common)
  • Transfer: Shift risk through insurance, contractual provisions, or third-party services
  • Accept: Formally accept residual risk with documented rationale and executive approval
  • Avoid: Do not deploy the AI system if risk cannot be reduced to acceptable levels

Document the treatment decision, rationale, responsible party, implementation timeline, and residual risk level for every identified risk.

Deployment Controls

Implement graduated deployment strategies:

  1. Canary deployment — Route 1-5% of traffic to the new system, monitor for anomalies
  2. A/B testing — Compare AI system performance against baseline (previous system or human decisions)
  3. Shadow mode — Run AI system in parallel without affecting real decisions, compare outputs
  4. Progressive rollout — Gradually increase deployment scope with monitoring at each stage
  5. Kill switch — Pre-defined criteria and mechanism for immediate system shutdown

Human Oversight Mechanisms

Design human oversight proportionate to system risk:

  • Human-in-the-loop: Human reviews and approves every AI decision (high-risk, low-volume decisions)
  • Human-on-the-loop: Human monitors AI decisions and can intervene (medium-risk, medium-volume)
  • Human-over-the-loop: Human sets policies and reviews aggregate outcomes (lower-risk, high-volume)

Continuous Monitoring

Deploy monitoring infrastructure that tracks:

  • Model performance metrics in production (accuracy, latency, error rates)
  • Data drift detection (input distribution changes)
  • Concept drift detection (relationship between inputs and correct outputs changes)
  • Fairness metrics computed on production data
  • Security monitoring (adversarial inputs, anomalous usage patterns)
  • User feedback and complaint tracking

Incident Response

Establish AI-specific incident response procedures:

  1. Detection: Automated alerting on metric threshold breaches
  2. Triage: Severity classification (critical: immediate shutdown, high: remediation within hours, medium: next business day)
  3. Investigation: Root cause analysis with model debugging tools
  4. Remediation: Fix, retrain, rollback, or decommission
  5. Communication: Notify affected stakeholders per communication plan
  6. Post-incident: Lessons learned, control updates, metric threshold adjustments

Evidence Requirements

  • Risk treatment plans for each AI system with treatment decisions, rationale, and residual risk levels signed off by risk owners
  • Deployment plans showing graduated rollout strategy with monitoring criteria at each stage
  • Human oversight mechanism documentation showing escalation criteria, reviewer qualifications, and override procedures
  • Production monitoring dashboard configurations showing tracked metrics, alert thresholds, and escalation procedures
  • Data drift and concept drift detection system configurations with historical alert logs
  • AI incident response plan with severity definitions, communication templates, and post-incident review procedures
  • Incident reports and post-incident review documentation with lessons learned and control improvements
  • Kill switch procedures and testing records demonstrating ability to rapidly shut down AI systems
  • Regular risk reassessment reports showing updated risk profiles and treatment plan adjustments