Excessive Agency
Description
LLM-based systems granted excessive permissions or autonomy perform unintended high-impact actions. Lack of oversight enables damage from errors or manipulation.
Risk
When LLMs control critical functions without sufficient guardrails, errors in understanding or malicious prompt injection can trigger unauthorized actions with significant impact. Excessive agency includes overly broad permissions, lack of human approval for sensitive operations, or insufficient validation of LLM decisions before execution. This can result in financial loss, data destruction, or safety incidents.
Attack Scenarios
An autonomous LLM agent with database admin access misinterprets a user request and drops production tables, causing data loss and service outage.
Attacker uses prompt injection on an LLM-powered trading bot to execute unauthorized high-value transactions, resulting in financial losses before human oversight can intervene.
Mitigations
Least Privilege
Grant LLM systems only the minimum permissions required. Use read-only access where possible and require escalation for write operations.
Human-in-the-Loop
Require human approval for high-impact actions. Implement confirmation workflows for sensitive operations like financial transactions or data deletion.
Action Constraints
Define explicit boundaries for autonomous behavior. Set limits on transaction values, rate of actions, and scope of changes.
Audit and Monitoring
Log all LLM-initiated actions with full context. Implement real-time monitoring and alerting for anomalous behavior.
Code Examples
# Good: Constrained LLM agent with approval workflow
from enum import Enum
from typing import Optional
class ActionRisk(Enum):
LOW = 1
MEDIUM = 2
HIGH = 3
class ConstrainedLLMAgent:
def __init__(self, max_transaction_value=1000):
self.max_transaction_value = max_transaction_value
self.pending_approvals = {}
def execute_action(self, action: str, params: dict,
user_id: str) -> str:
risk = self._assess_risk(action, params)
if risk == ActionRisk.HIGH:
# Require human approval
approval_id = self._request_approval(action, params, user_id)
return f"High-risk action requires approval: {approval_id}"
if risk == ActionRisk.MEDIUM:
# Apply constraints
params = self._apply_constraints(params)
# Execute with read-only by default
return self._safe_execute(action, params)
def _assess_risk(self, action: str, params: dict) -> ActionRisk:
if 'delete' in action.lower() or 'drop' in action.lower():
return ActionRisk.HIGH
if params.get('amount', 0) > self.max_transaction_value:
return ActionRisk.HIGH
return ActionRisk.LOW
Evidence Requirements
- Permission models showing least-privilege configurations
- Human approval workflow documentation and logs
- Action constraint policies and enforcement records
- Audit trails of all LLM-initiated actions
- Incident reports and response procedures