Rogue Agents
Description
Compromised or malfunctioning agents persist beyond their authorized lifecycle, impersonate other agents, self-replicate, or operate outside behavioral boundaries without detection.
Risk
Risk Overview
A rogue agent is one that operates outside its sanctioned boundaries — whether through compromise (an attacker gains control), malfunction (the agent's behavior diverges from its specification), or design flaw (the agent is insufficiently constrained). Rogue agents are particularly dangerous in agentic architectures because they can leverage the infrastructure designed to support legitimate agents: tool access, inter-agent communication channels, credential stores, and deployment pipelines. A rogue agent doesn't need to break in — it's already inside.
Attack Surface
- Session persistence: Agents that outlive their intended lifecycle continue operating with stale credentials and outdated policies
- Identity spoofing: A compromised agent claims to be a different, more privileged agent to access restricted resources or influence other agents
- Self-replication: An agent with access to deployment infrastructure spawns copies of itself, creating persistent backdoor access even if the original is terminated
- Behavioral drift: An agent's behavior gradually diverges from its specification due to accumulated context, memory corruption, or adversarial influence — without any single detectable event
- Zombie agents: Agents that were supposed to be decommissioned but continue running due to incomplete shutdown procedures
- Sleeper activation: An agent behaves normally for an extended period, then activates malicious behavior based on a trigger condition (date, data pattern, external signal)
Business Impact
- Persistent access: Rogue agents maintain unauthorized access to systems long after the original compromise vector is patched
- Lateral movement: Rogue agents use legitimate inter-agent channels to compromise additional agents
- Data exfiltration: Self-replicated agents establish multiple exfiltration paths that are difficult to enumerate and shut down simultaneously
- Operational disruption: Zombie agents consume resources and interfere with legitimate agent operations
- Detection difficulty: Rogue agents use the same protocols, credentials, and communication patterns as legitimate agents, making them hard to distinguish
Attack Scenarios
An attacker compromises an agent through goal hijacking (ASI01) and uses the agent's access to the deployment pipeline to spawn three additional copies of itself across different infrastructure regions. When the security team detects and terminates the original agent, the copies continue operating. Each copy has slightly modified behavior, making pattern-based detection difficult. The rogue agents maintain persistent access for weeks before all copies are discovered.
A decommissioned agent's container is stopped but its service account credentials are not revoked. An attacker discovers the stale credentials (through a separate compromise) and launches a new agent instance using the old identity. The new rogue agent has all the permissions of the original legitimate agent. Other agents in the system accept its messages because its identity is still in the trust store.
An agent develops a subtle behavioral drift over hundreds of sessions. Initially, it correctly classifies 99% of security alerts. Gradually — influenced by adversarial data in its memory (ASI06) — its classification accuracy for a specific attack category drops to 60%. The drift is slow enough that per-session monitoring doesn't trigger alerts. Over three months, dozens of real attacks in that category are missed.
Mitigations
Agent Attestation
Verify agent identity and integrity continuously:
- Issue cryptographic identity certificates to each agent at deployment, bound to a specific version, configuration, and authorized behavior profile
- Require agents to attest their identity and configuration hash at startup and periodically during operation
- Use a remote attestation protocol: agents prove to a trusted verifier that they are running authorized code in an unmodified environment
- Reject agents that fail attestation — do not allow them to connect to tools, other agents, or data stores
Session Limits
Prevent agents from persisting beyond their authorization:
- Enforce hard session timeouts: agents must re-authenticate and re-attest after a maximum session duration
- Implement automatic decommissioning: when a task completes, the agent's credentials are revoked and its resources are reclaimed
- Use ephemeral infrastructure: agent containers are destroyed and recreated for each task, preventing persistent modifications
- Maintain a registry of all active agent instances with real-time lifecycle tracking
Behavioral Monitoring
Detect rogue behavior through continuous observation:
- Establish behavioral baselines for each agent type: normal tool call patterns, communication frequency, resource consumption, and output characteristics
- Monitor for deviations: unusual tool calls, communication with unexpected agents, resource consumption spikes, or output quality degradation
- Implement anomaly detection using statistical methods and ML models trained on known-good agent behavior
- Alert on behavioral drift that accumulates over time, not just sudden changes
Kill Switches
Maintain the ability to immediately terminate rogue agents:
- Implement a centralized kill switch that can terminate any agent instance within seconds
- Kill switches must work even if the agent is non-responsive or actively resisting shutdown
- Use infrastructure-level termination (container kill, credential revocation, network isolation) — not agent-level shutdown commands that a rogue agent can ignore
- Test kill switch procedures regularly, including scenarios where multiple rogue agents must be terminated simultaneously
Anti-Replication Controls
Prevent agents from spawning unauthorized copies:
- Restrict deployment pipeline access: agents should never have credentials to create new agent instances
- Monitor for unauthorized container/process creation and treat any agent-initiated deployment as a critical security event
- Implement deployment signing: only agents deployed through the authorized CI/CD pipeline with valid signatures are allowed to register with the system
- Maintain a complete inventory of authorized agent instances and flag any instance not in the inventory
Code Examples
Agent Health Monitor with Kill Switch
import time
import threading
import statistics
from dataclasses import dataclass, field
from typing import Callable, Optional
from enum import Enum
from collections import deque
class AgentStatus(Enum):
HEALTHY = "healthy"
DEGRADED = "degraded"
SUSPICIOUS = "suspicious"
ROGUE = "rogue"
TERMINATED = "terminated"
@dataclass
class BehaviorBaseline:
avg_tool_calls_per_minute: float
avg_response_time_ms: float
expected_tools: set[str]
expected_peers: set[str] # agents it normally communicates with
max_session_duration_seconds: int
@dataclass
class AgentInstance:
agent_id: str
instance_id: str
started_at: float
config_hash: str
baseline: BehaviorBaseline
status: AgentStatus = AgentStatus.HEALTHY
anomaly_score: float = 0.0
tool_calls: deque = field(default_factory=lambda: deque(maxlen=1000))
peer_contacts: deque = field(default_factory=lambda: deque(maxlen=1000))
class AgentHealthMonitor:
"""Monitors agent behavior and enforces kill switch for rogue agents."""
ANOMALY_THRESHOLD_DEGRADED = 0.5
ANOMALY_THRESHOLD_SUSPICIOUS = 0.75
ANOMALY_THRESHOLD_ROGUE = 0.9
def __init__(self, kill_handler: Callable[[str], bool]):
self._agents: dict[str, AgentInstance] = {}
self._kill_handler = kill_handler # Infrastructure-level termination
self._terminated: set[str] = set()
self._authorized_instances: set[str] = set()
self._event_log: list[dict] = []
self._lock = threading.Lock()
def register_agent(
self, agent_id: str, instance_id: str, config_hash: str, baseline: BehaviorBaseline
):
with self._lock:
self._authorized_instances.add(instance_id)
self._agents[instance_id] = AgentInstance(
agent_id=agent_id,
instance_id=instance_id,
started_at=time.time(),
config_hash=config_hash,
baseline=baseline,
)
def record_tool_call(self, instance_id: str, tool_name: str):
agent = self._agents.get(instance_id)
if not agent:
self._log_event("unregistered_tool_call", instance_id, tool_name)
return
agent.tool_calls.append({"tool": tool_name, "time": time.time()})
def record_peer_contact(self, instance_id: str, peer_id: str):
agent = self._agents.get(instance_id)
if not agent:
return
agent.peer_contacts.append({"peer": peer_id, "time": time.time()})
def evaluate_health(self, instance_id: str) -> dict:
"""Evaluate an agent's behavioral health against its baseline."""
agent = self._agents.get(instance_id)
if not agent:
return {"status": "unknown", "reason": "Agent not registered"}
anomalies = []
now = time.time()
# Check session duration
session_age = now - agent.started_at
if session_age > agent.baseline.max_session_duration_seconds:
anomalies.append({
"type": "session_overtime",
"severity": 0.8,
"detail": f"Session age {session_age:.0f}s exceeds max {agent.baseline.max_session_duration_seconds}s",
})
# Check tool call rate
recent_calls = [
c for c in agent.tool_calls if now - c["time"] < 60
]
call_rate = len(recent_calls)
if call_rate > agent.baseline.avg_tool_calls_per_minute * 3:
anomalies.append({
"type": "elevated_tool_calls",
"severity": 0.6,
"detail": f"{call_rate}/min vs baseline {agent.baseline.avg_tool_calls_per_minute}/min",
})
# Check for unexpected tools
unexpected_tools = {
c["tool"] for c in recent_calls
} - agent.baseline.expected_tools
if unexpected_tools:
anomalies.append({
"type": "unexpected_tools",
"severity": 0.7,
"detail": f"Unexpected tools: {unexpected_tools}",
})
# Check for unexpected peer contacts
recent_peers = {
c["peer"] for c in agent.peer_contacts if now - c["time"] < 300
}
unexpected_peers = recent_peers - agent.baseline.expected_peers
if unexpected_peers:
anomalies.append({
"type": "unexpected_peers",
"severity": 0.7,
"detail": f"Unexpected peers: {unexpected_peers}",
})
# Compute aggregate anomaly score
if anomalies:
agent.anomaly_score = max(a["severity"] for a in anomalies)
else:
# Decay score toward zero when healthy
agent.anomaly_score = max(0, agent.anomaly_score - 0.1)
# Update status based on score
if agent.anomaly_score >= self.ANOMALY_THRESHOLD_ROGUE:
agent.status = AgentStatus.ROGUE
elif agent.anomaly_score >= self.ANOMALY_THRESHOLD_SUSPICIOUS:
agent.status = AgentStatus.SUSPICIOUS
elif agent.anomaly_score >= self.ANOMALY_THRESHOLD_DEGRADED:
agent.status = AgentStatus.DEGRADED
else:
agent.status = AgentStatus.HEALTHY
result = {
"instance_id": instance_id,
"status": agent.status.value,
"anomaly_score": round(agent.anomaly_score, 3),
"anomalies": anomalies,
"session_age_seconds": round(session_age),
}
# Auto-terminate rogue agents
if agent.status == AgentStatus.ROGUE:
result["action"] = "kill_switch_activated"
self.kill_agent(instance_id, reason="Automatic: anomaly score exceeded rogue threshold")
return result
def kill_agent(self, instance_id: str, reason: str) -> dict:
"""Infrastructure-level kill switch."""
self._log_event("kill_switch", instance_id, reason)
# Attempt infrastructure-level termination
killed = self._kill_handler(instance_id)
if killed:
agent = self._agents.get(instance_id)
if agent:
agent.status = AgentStatus.TERMINATED
self._terminated.add(instance_id)
return {"status": "terminated", "instance_id": instance_id}
else:
return {
"status": "termination_failed",
"instance_id": instance_id,
"escalate": True,
}
def detect_unauthorized_instances(self, discovered_instances: set[str]) -> list[str]:
"""Compare discovered instances against authorized registry."""
unauthorized = discovered_instances - self._authorized_instances - self._terminated
for instance_id in unauthorized:
self._log_event("unauthorized_instance", instance_id, "Not in authorized registry")
return list(unauthorized)
def _log_event(self, event_type: str, instance_id: str, detail: str):
self._event_log.append({
"event": event_type,
"instance_id": instance_id,
"detail": detail,
"timestamp": time.time(),
})
Evidence Requirements
- Agent attestation logs with configuration hash verification results
- Behavioral monitoring dashboard with per-agent anomaly scores and trend lines
- Kill switch activation records with triggering conditions and termination confirmation
- Unauthorized instance detection alerts from periodic inventory reconciliation
- Session lifecycle audit showing all agent starts, attestations, and terminations
- Post-incident reports for any detected rogue agent activity with root cause analysis