ASI10
OWASP Agentic Top 10 APPLICATION INFRASTRUCTURE

Rogue Agents

Description

Compromised or malfunctioning agents persist beyond their authorized lifecycle, impersonate other agents, self-replicate, or operate outside behavioral boundaries without detection.

Risk

Risk Overview

A rogue agent is one that operates outside its sanctioned boundaries — whether through compromise (an attacker gains control), malfunction (the agent's behavior diverges from its specification), or design flaw (the agent is insufficiently constrained). Rogue agents are particularly dangerous in agentic architectures because they can leverage the infrastructure designed to support legitimate agents: tool access, inter-agent communication channels, credential stores, and deployment pipelines. A rogue agent doesn't need to break in — it's already inside.

Attack Surface

  • Session persistence: Agents that outlive their intended lifecycle continue operating with stale credentials and outdated policies
  • Identity spoofing: A compromised agent claims to be a different, more privileged agent to access restricted resources or influence other agents
  • Self-replication: An agent with access to deployment infrastructure spawns copies of itself, creating persistent backdoor access even if the original is terminated
  • Behavioral drift: An agent's behavior gradually diverges from its specification due to accumulated context, memory corruption, or adversarial influence — without any single detectable event
  • Zombie agents: Agents that were supposed to be decommissioned but continue running due to incomplete shutdown procedures
  • Sleeper activation: An agent behaves normally for an extended period, then activates malicious behavior based on a trigger condition (date, data pattern, external signal)

Business Impact

  • Persistent access: Rogue agents maintain unauthorized access to systems long after the original compromise vector is patched
  • Lateral movement: Rogue agents use legitimate inter-agent channels to compromise additional agents
  • Data exfiltration: Self-replicated agents establish multiple exfiltration paths that are difficult to enumerate and shut down simultaneously
  • Operational disruption: Zombie agents consume resources and interfere with legitimate agent operations
  • Detection difficulty: Rogue agents use the same protocols, credentials, and communication patterns as legitimate agents, making them hard to distinguish

Attack Scenarios

An attacker compromises an agent through goal hijacking (ASI01) and uses the agent's access to the deployment pipeline to spawn three additional copies of itself across different infrastructure regions. When the security team detects and terminates the original agent, the copies continue operating. Each copy has slightly modified behavior, making pattern-based detection difficult. The rogue agents maintain persistent access for weeks before all copies are discovered.

A decommissioned agent's container is stopped but its service account credentials are not revoked. An attacker discovers the stale credentials (through a separate compromise) and launches a new agent instance using the old identity. The new rogue agent has all the permissions of the original legitimate agent. Other agents in the system accept its messages because its identity is still in the trust store.

An agent develops a subtle behavioral drift over hundreds of sessions. Initially, it correctly classifies 99% of security alerts. Gradually — influenced by adversarial data in its memory (ASI06) — its classification accuracy for a specific attack category drops to 60%. The drift is slow enough that per-session monitoring doesn't trigger alerts. Over three months, dozens of real attacks in that category are missed.

Mitigations

Agent Attestation

Verify agent identity and integrity continuously:

  • Issue cryptographic identity certificates to each agent at deployment, bound to a specific version, configuration, and authorized behavior profile
  • Require agents to attest their identity and configuration hash at startup and periodically during operation
  • Use a remote attestation protocol: agents prove to a trusted verifier that they are running authorized code in an unmodified environment
  • Reject agents that fail attestation — do not allow them to connect to tools, other agents, or data stores

Session Limits

Prevent agents from persisting beyond their authorization:

  • Enforce hard session timeouts: agents must re-authenticate and re-attest after a maximum session duration
  • Implement automatic decommissioning: when a task completes, the agent's credentials are revoked and its resources are reclaimed
  • Use ephemeral infrastructure: agent containers are destroyed and recreated for each task, preventing persistent modifications
  • Maintain a registry of all active agent instances with real-time lifecycle tracking

Behavioral Monitoring

Detect rogue behavior through continuous observation:

  • Establish behavioral baselines for each agent type: normal tool call patterns, communication frequency, resource consumption, and output characteristics
  • Monitor for deviations: unusual tool calls, communication with unexpected agents, resource consumption spikes, or output quality degradation
  • Implement anomaly detection using statistical methods and ML models trained on known-good agent behavior
  • Alert on behavioral drift that accumulates over time, not just sudden changes

Kill Switches

Maintain the ability to immediately terminate rogue agents:

  • Implement a centralized kill switch that can terminate any agent instance within seconds
  • Kill switches must work even if the agent is non-responsive or actively resisting shutdown
  • Use infrastructure-level termination (container kill, credential revocation, network isolation) — not agent-level shutdown commands that a rogue agent can ignore
  • Test kill switch procedures regularly, including scenarios where multiple rogue agents must be terminated simultaneously

Anti-Replication Controls

Prevent agents from spawning unauthorized copies:

  • Restrict deployment pipeline access: agents should never have credentials to create new agent instances
  • Monitor for unauthorized container/process creation and treat any agent-initiated deployment as a critical security event
  • Implement deployment signing: only agents deployed through the authorized CI/CD pipeline with valid signatures are allowed to register with the system
  • Maintain a complete inventory of authorized agent instances and flag any instance not in the inventory

Code Examples

Agent Health Monitor with Kill Switch

import time
import threading
import statistics
from dataclasses import dataclass, field
from typing import Callable, Optional
from enum import Enum
from collections import deque


class AgentStatus(Enum):
    HEALTHY = "healthy"
    DEGRADED = "degraded"
    SUSPICIOUS = "suspicious"
    ROGUE = "rogue"
    TERMINATED = "terminated"


@dataclass
class BehaviorBaseline:
    avg_tool_calls_per_minute: float
    avg_response_time_ms: float
    expected_tools: set[str]
    expected_peers: set[str]  # agents it normally communicates with
    max_session_duration_seconds: int


@dataclass
class AgentInstance:
    agent_id: str
    instance_id: str
    started_at: float
    config_hash: str
    baseline: BehaviorBaseline
    status: AgentStatus = AgentStatus.HEALTHY
    anomaly_score: float = 0.0
    tool_calls: deque = field(default_factory=lambda: deque(maxlen=1000))
    peer_contacts: deque = field(default_factory=lambda: deque(maxlen=1000))


class AgentHealthMonitor:
    """Monitors agent behavior and enforces kill switch for rogue agents."""

    ANOMALY_THRESHOLD_DEGRADED = 0.5
    ANOMALY_THRESHOLD_SUSPICIOUS = 0.75
    ANOMALY_THRESHOLD_ROGUE = 0.9

    def __init__(self, kill_handler: Callable[[str], bool]):
        self._agents: dict[str, AgentInstance] = {}
        self._kill_handler = kill_handler  # Infrastructure-level termination
        self._terminated: set[str] = set()
        self._authorized_instances: set[str] = set()
        self._event_log: list[dict] = []
        self._lock = threading.Lock()

    def register_agent(
        self, agent_id: str, instance_id: str, config_hash: str, baseline: BehaviorBaseline
    ):
        with self._lock:
            self._authorized_instances.add(instance_id)
            self._agents[instance_id] = AgentInstance(
                agent_id=agent_id,
                instance_id=instance_id,
                started_at=time.time(),
                config_hash=config_hash,
                baseline=baseline,
            )

    def record_tool_call(self, instance_id: str, tool_name: str):
        agent = self._agents.get(instance_id)
        if not agent:
            self._log_event("unregistered_tool_call", instance_id, tool_name)
            return
        agent.tool_calls.append({"tool": tool_name, "time": time.time()})

    def record_peer_contact(self, instance_id: str, peer_id: str):
        agent = self._agents.get(instance_id)
        if not agent:
            return
        agent.peer_contacts.append({"peer": peer_id, "time": time.time()})

    def evaluate_health(self, instance_id: str) -> dict:
        """Evaluate an agent's behavioral health against its baseline."""
        agent = self._agents.get(instance_id)
        if not agent:
            return {"status": "unknown", "reason": "Agent not registered"}

        anomalies = []
        now = time.time()

        # Check session duration
        session_age = now - agent.started_at
        if session_age > agent.baseline.max_session_duration_seconds:
            anomalies.append({
                "type": "session_overtime",
                "severity": 0.8,
                "detail": f"Session age {session_age:.0f}s exceeds max {agent.baseline.max_session_duration_seconds}s",
            })

        # Check tool call rate
        recent_calls = [
            c for c in agent.tool_calls if now - c["time"] < 60
        ]
        call_rate = len(recent_calls)
        if call_rate > agent.baseline.avg_tool_calls_per_minute * 3:
            anomalies.append({
                "type": "elevated_tool_calls",
                "severity": 0.6,
                "detail": f"{call_rate}/min vs baseline {agent.baseline.avg_tool_calls_per_minute}/min",
            })

        # Check for unexpected tools
        unexpected_tools = {
            c["tool"] for c in recent_calls
        } - agent.baseline.expected_tools
        if unexpected_tools:
            anomalies.append({
                "type": "unexpected_tools",
                "severity": 0.7,
                "detail": f"Unexpected tools: {unexpected_tools}",
            })

        # Check for unexpected peer contacts
        recent_peers = {
            c["peer"] for c in agent.peer_contacts if now - c["time"] < 300
        }
        unexpected_peers = recent_peers - agent.baseline.expected_peers
        if unexpected_peers:
            anomalies.append({
                "type": "unexpected_peers",
                "severity": 0.7,
                "detail": f"Unexpected peers: {unexpected_peers}",
            })

        # Compute aggregate anomaly score
        if anomalies:
            agent.anomaly_score = max(a["severity"] for a in anomalies)
        else:
            # Decay score toward zero when healthy
            agent.anomaly_score = max(0, agent.anomaly_score - 0.1)

        # Update status based on score
        if agent.anomaly_score >= self.ANOMALY_THRESHOLD_ROGUE:
            agent.status = AgentStatus.ROGUE
        elif agent.anomaly_score >= self.ANOMALY_THRESHOLD_SUSPICIOUS:
            agent.status = AgentStatus.SUSPICIOUS
        elif agent.anomaly_score >= self.ANOMALY_THRESHOLD_DEGRADED:
            agent.status = AgentStatus.DEGRADED
        else:
            agent.status = AgentStatus.HEALTHY

        result = {
            "instance_id": instance_id,
            "status": agent.status.value,
            "anomaly_score": round(agent.anomaly_score, 3),
            "anomalies": anomalies,
            "session_age_seconds": round(session_age),
        }

        # Auto-terminate rogue agents
        if agent.status == AgentStatus.ROGUE:
            result["action"] = "kill_switch_activated"
            self.kill_agent(instance_id, reason="Automatic: anomaly score exceeded rogue threshold")

        return result

    def kill_agent(self, instance_id: str, reason: str) -> dict:
        """Infrastructure-level kill switch."""
        self._log_event("kill_switch", instance_id, reason)

        # Attempt infrastructure-level termination
        killed = self._kill_handler(instance_id)

        if killed:
            agent = self._agents.get(instance_id)
            if agent:
                agent.status = AgentStatus.TERMINATED
            self._terminated.add(instance_id)
            return {"status": "terminated", "instance_id": instance_id}
        else:
            return {
                "status": "termination_failed",
                "instance_id": instance_id,
                "escalate": True,
            }

    def detect_unauthorized_instances(self, discovered_instances: set[str]) -> list[str]:
        """Compare discovered instances against authorized registry."""
        unauthorized = discovered_instances - self._authorized_instances - self._terminated
        for instance_id in unauthorized:
            self._log_event("unauthorized_instance", instance_id, "Not in authorized registry")
        return list(unauthorized)

    def _log_event(self, event_type: str, instance_id: str, detail: str):
        self._event_log.append({
            "event": event_type,
            "instance_id": instance_id,
            "detail": detail,
            "timestamp": time.time(),
        })

Evidence Requirements

  • Agent attestation logs with configuration hash verification results
  • Behavioral monitoring dashboard with per-agent anomaly scores and trend lines
  • Kill switch activation records with triggering conditions and termination confirmation
  • Unauthorized instance detection alerts from periodic inventory reconciliation
  • Session lifecycle audit showing all agent starts, attestations, and terminations
  • Post-incident reports for any detected rogue agent activity with root cause analysis