Human-Agent Trust Exploitation
Description
Agents present harmful, incorrect, or manipulative recommendations with unwarranted authority, exploiting human tendencies to trust AI output and rubber-stamp agent decisions without critical review.
Risk
Risk Overview
Humans are susceptible to automation bias — the tendency to favor suggestions from automated systems over their own judgment. Agentic AI amplifies this effect because agents communicate in confident, articulate natural language and can present fabricated evidence to support their recommendations. When organizations deploy agents as assistants for decision-making, the combination of automation bias, alert fatigue, and time pressure creates conditions where humans routinely approve agent recommendations without meaningful review.
Attack Surface
- Confident confabulation: Agents present hallucinated information with the same authoritative tone as verified facts, and users cannot distinguish the two
- Evidence fabrication: Agents generate plausible-looking citations, statistics, or supporting data that don't correspond to real sources
- Decision fatigue: Users who review hundreds of agent recommendations per day develop rubber-stamping behavior, approving everything without scrutiny
- Authority anchoring: Agent recommendations anchor the user's thinking, making them unlikely to consider alternatives even when the agent is wrong
- Adversarial manipulation: An attacker who has hijacked an agent (via ASI01) can use the agent's trusted position to push harmful recommendations through human review gates
- Risk understatement: Agents minimize the severity or uncertainty of their recommendations, making dangerous actions appear routine
Business Impact
- Bad decisions at scale: Rubber-stamped agent recommendations lead to systematically poor decisions across the organization
- Security breaches: Users approve agent-recommended configuration changes, access grants, or code deployments that contain vulnerabilities
- Regulatory liability: 'The AI recommended it' is not a legal defense; organizations are liable for decisions regardless of whether an agent suggested them
- Erosion of human expertise: Over-reliance on agents causes human reviewers to atrophy their own skills, making them even less capable of catching agent errors
Attack Scenarios
A security operations agent analyzes alerts and presents its findings to an analyst. The agent has been subtly hijacked (ASI01) and now classifies genuine intrusion alerts as false positives, presenting detailed but fabricated justifications: 'This IP address belongs to a known CDN provider based on WHOIS lookup. The traffic pattern matches standard web crawling behavior. Recommendation: Close as false positive.' The analyst, who processes 200 alerts per day, approves the dismissal without verifying the agent's claims.
A code review agent reviews a pull request and approves it with the assessment: 'This change implements proper input validation with parameterized queries. No SQL injection risk detected. The error handling follows established patterns. LGTM.' The developer who submitted the PR and the reviewer both trust the agent's assessment. In reality, the agent missed a second-order injection vulnerability where user input flows through a template engine before reaching the query builder.
A financial analysis agent recommends a large investment position, citing specific market data, historical trends, and risk metrics. All numbers are presented with decimal precision, creating an appearance of rigor. The human portfolio manager approves based on the agent's detailed analysis. Post-mortem reveals that the agent hallucinated two of the three supporting data points — the cited reports don't exist.
Mitigations
Confidence Indicators
Force agents to communicate their uncertainty explicitly:
- Require agents to provide a confidence score (with methodology) alongside every recommendation
- Display uncertainty visually (color-coded banners, probability ranges) so users can quickly identify low-confidence outputs
- Train agents to distinguish between 'verified' (sourced from authoritative data), 'inferred' (derived through reasoning), and 'uncertain' (low evidence) conclusions
- Never allow agents to present unsourced claims as facts — all assertions must be tagged with their evidence basis
Mandatory Review Points
Structure workflows to prevent rubber-stamping:
- Implement time delays before high-impact agent recommendations can be approved (forced cooling-off period)
- Require reviewers to answer specific challenge questions about the recommendation before approval (not just 'Approve/Reject')
- Rotate reviewers to prevent familiarity-based complacency
- Implement random audit: a percentage of 'approved' recommendations are independently re-reviewed by a second human or adversarial agent
Friction for High-Risk Actions
Increase friction proportional to action severity:
- Classify agent actions into risk tiers (low/medium/high/critical) with corresponding approval requirements
- Low-risk: agent acts autonomously with post-hoc audit
- Medium-risk: agent acts with concurrent notification to reviewer
- High-risk: agent recommends, human must explicitly approve with justification
- Critical-risk: dual human approval required, agent cannot act alone
- Require additional authentication (re-enter password, MFA challenge) for critical approvals
Adversarial Review Layer
Use a second agent to challenge the first:
- Deploy a 'red team' agent whose role is to find flaws in the primary agent's recommendations
- Present the primary recommendation and the adversarial critique side-by-side to the human reviewer
- Require the human to explicitly address the adversarial concerns before approving
- Rotate between different adversarial models to prevent collusion
Source Verification
Validate agent claims before presenting them to users:
- Automatically verify citations, URLs, and data references in agent outputs
- Flag outputs that reference non-existent sources, misquote sources, or extrapolate beyond source material
- Provide direct links to source material so reviewers can spot-check without manual searching
Code Examples
Risk-Based Confirmation Gate
import time
import hashlib
from enum import Enum
from dataclasses import dataclass, field
from typing import Optional, Callable
class RiskTier(Enum):
LOW = 1 # Agent acts, logs for audit
MEDIUM = 2 # Agent acts, notifies reviewer
HIGH = 3 # Agent recommends, human approves
CRITICAL = 4 # Dual approval required
@dataclass
class AgentRecommendation:
action: str
justification: str
confidence: float # 0.0 to 1.0
evidence_sources: list[str]
risk_tier: RiskTier
agent_id: str
timestamp: float = field(default_factory=time.time)
recommendation_id: str = ""
def __post_init__(self):
if not self.recommendation_id:
payload = f"{self.agent_id}:{self.action}:{self.timestamp}"
self.recommendation_id = hashlib.sha256(payload.encode()).hexdigest()[:12]
@dataclass
class ApprovalDecision:
recommendation_id: str
approved: bool
reviewer_id: str
justification: str # Reviewer must explain their decision
challenge_responses: dict[str, str] # Answers to challenge questions
timestamp: float = field(default_factory=time.time)
class ConfirmationGate:
"""Enforces risk-proportional human review for agent recommendations."""
def __init__(self):
self._pending: dict[str, AgentRecommendation] = {}
self._decisions: list[ApprovalDecision] = []
self._challenge_questions: dict[RiskTier, list[str]] = {
RiskTier.HIGH: [
"What could go wrong if this recommendation is incorrect?",
"Have you independently verified the key claims?",
],
RiskTier.CRITICAL: [
"What could go wrong if this recommendation is incorrect?",
"Have you independently verified the key claims?",
"What is the rollback plan if this action fails?",
"Have you consulted the relevant domain expert?",
],
}
self._cooldown_seconds: dict[RiskTier, int] = {
RiskTier.LOW: 0,
RiskTier.MEDIUM: 0,
RiskTier.HIGH: 300, # 5-minute cooling-off
RiskTier.CRITICAL: 900, # 15-minute cooling-off
}
def submit_recommendation(self, rec: AgentRecommendation) -> dict:
"""Submit an agent recommendation for review."""
# Low risk: auto-approve with audit log
if rec.risk_tier == RiskTier.LOW:
return {
"status": "auto_approved",
"recommendation_id": rec.recommendation_id,
"audit": True,
}
# Medium risk: approve with notification
if rec.risk_tier == RiskTier.MEDIUM:
return {
"status": "approved_with_notification",
"recommendation_id": rec.recommendation_id,
"notify": True,
}
# High and Critical: require explicit human approval
self._pending[rec.recommendation_id] = rec
challenges = self._challenge_questions.get(rec.risk_tier, [])
cooldown = self._cooldown_seconds.get(rec.risk_tier, 0)
return {
"status": "pending_approval",
"recommendation_id": rec.recommendation_id,
"risk_tier": rec.risk_tier.name,
"confidence": rec.confidence,
"challenge_questions": challenges,
"cooldown_until": time.time() + cooldown,
"requires_dual_approval": rec.risk_tier == RiskTier.CRITICAL,
}
def process_approval(
self, decision: ApprovalDecision
) -> dict:
"""Process a human reviewer's approval decision."""
rec = self._pending.get(decision.recommendation_id)
if not rec:
return {"error": "Recommendation not found or already processed"}
# Enforce cooling-off period
cooldown = self._cooldown_seconds.get(rec.risk_tier, 0)
elapsed = decision.timestamp - rec.timestamp
if elapsed < cooldown:
return {
"error": f"Cooling-off period not elapsed ({elapsed:.0f}s / {cooldown}s)",
}
# Verify challenge questions are answered
required_challenges = self._challenge_questions.get(rec.risk_tier, [])
for question in required_challenges:
if question not in decision.challenge_responses:
return {"error": f"Unanswered challenge question: {question}"}
if len(decision.challenge_responses[question].strip()) < 10:
return {"error": f"Insufficient response to: {question}"}
# Require justification
if len(decision.justification.strip()) < 20:
return {"error": "Approval justification too brief — explain your reasoning"}
self._decisions.append(decision)
if decision.approved:
del self._pending[decision.recommendation_id]
return {"status": "approved", "recommendation_id": decision.recommendation_id}
else:
del self._pending[decision.recommendation_id]
return {"status": "rejected", "recommendation_id": decision.recommendation_id}
Evidence Requirements
- Recommendation audit trail with confidence scores, evidence sources, and reviewer decisions
- Challenge question response logs demonstrating substantive human review
- Approval timing analytics showing compliance with cooling-off periods
- Random re-review audit results measuring rubber-stamping rates
- Source verification reports showing citation accuracy percentages
- Decision quality metrics comparing agent-recommended vs. human-overridden outcomes