Cascading Failures
Description
Errors in multi-agent workflows propagate unchecked across dependent agents, amplifying a single point of failure into system-wide outages, data corruption, or runaway resource consumption.
Risk
Risk Overview
Multi-agent systems create dependency chains where the output of one agent feeds the input of the next. When one agent fails — whether from a model error, tool timeout, resource exhaustion, or adversarial input — that failure propagates downstream. Unlike traditional distributed systems that have decades of established fault-tolerance patterns, agentic workflows often lack circuit breakers, retry limits, or fallback strategies. A single agent failure can cascade into a system-wide outage where dozens of agents are stuck in retry loops, consuming resources, and producing corrupted outputs.
Attack Surface
- Error amplification: One agent produces malformed output; downstream agents fail trying to parse it, each generating their own error outputs that further confuse the system
- Retry storms: Multiple agents retry failed operations simultaneously, overwhelming shared resources (databases, APIs, model endpoints)
- Resource exhaustion: An agent caught in a failure loop consumes all available compute, memory, or API quota, starving other agents
- Poison pill propagation: A corrupted data artifact passes through an agent pipeline, each agent adding its own corruption layer until the final output is completely unreliable
- Deadlocks: Circular dependencies between agents cause mutual waiting that never resolves
- Thundering herd: All agents in a pool recover from a failure simultaneously, creating a load spike that triggers another failure
Business Impact
- System-wide outage: A single agent failure cascades to take down an entire multi-agent workflow
- Resource costs: Retry storms and resource exhaustion generate unexpected compute and API costs
- Data integrity: Corrupted outputs propagate through the pipeline before anyone detects the original failure
- SLA violations: Cascading delays cause end-to-end processing times to exceed service level agreements
- Difficult diagnosis: The root cause is buried under layers of cascading effects, extending mean time to resolution
Attack Scenarios
A model serving endpoint experiences elevated latency, causing the first agent in a pipeline to time out. The orchestrator retries the request 5 times with no backoff. All 5 retries arrive at the already-overloaded endpoint simultaneously. Meanwhile, 3 downstream agents waiting for the first agent's output also time out and begin their own retry loops. Within minutes, the system generates 50x normal load on the model endpoint, causing it to crash entirely and taking down all agent workflows that depend on it.
A data ingestion agent encounters a malformed CSV file and produces a partial, corrupted output instead of failing cleanly. The next agent in the pipeline — a data validation agent — processes the corrupted data and flags some rows as invalid but passes the rest through. A downstream analysis agent produces incorrect results from the corrupted data. The final reporting agent presents these results to stakeholders as authoritative analysis. The root cause (a malformed CSV) isn't discovered until the business decisions based on the report are already executed.
An attacker deliberately sends a request that causes an agent to enter an infinite reasoning loop. The looping agent consumes its entire token budget, then requests more. The orchestrator allocates additional resources. The agent's loop generates outputs that trigger similar loops in peer agents. Within an hour, the entire agent pool is consumed by infinite loops, and the API quota for the month is exhausted.
Mitigations
Circuit Breakers
Implement circuit breakers at every agent-to-agent and agent-to-service boundary:
- Track failure rates over sliding time windows (e.g., >50% failure rate in the last 60 seconds trips the breaker)
- When the circuit opens, immediately return a fallback response instead of forwarding requests to the failing service
- After a cooldown period, allow a single probe request through to test recovery (half-open state)
- Alert operators when circuits trip and require manual review before resetting if the failure was caused by data corruption
Fallback Handlers
Define degraded-mode behavior for every agent in the pipeline:
- Each agent must declare what it does when its dependencies are unavailable: return cached results, provide a default response, or gracefully decline the task
- Implement tiered degradation: first try a backup model, then cached results, then a static fallback, then a clear error message
- Never let an agent silently produce garbage output — it must either produce validated output or explicitly fail
- Test fallback paths regularly; untested fallbacks fail when you need them most
Blast Radius Limits
Contain failures to prevent system-wide impact:
- Set resource budgets per agent: maximum tokens, maximum API calls, maximum execution time, maximum memory
- Implement bulkheads: separate agent pools for critical vs. non-critical workflows so a failure in one pool doesn't starve the other
- Use priority queues to ensure high-priority agent tasks are processed even during resource contention
- Limit retry attempts with exponential backoff and jitter to prevent thundering herds
Health Checks
Detect failures early before they cascade:
- Implement liveness probes (is the agent running?) and readiness probes (is the agent ready to accept work?) for all agents
- Monitor end-to-end pipeline latency, not just individual agent latency — a slow agent in the middle causes invisible delays downstream
- Track error rates per agent and per agent-to-agent edge in the dependency graph
- Implement synthetic transactions that exercise the full pipeline on a schedule to detect degradation before users do
Output Validation Gates
Prevent corrupted outputs from propagating:
- Define output schemas for every agent and validate output against the schema before passing it downstream
- Implement checksums or content-based validation for data artifacts passed between agents
- Use a poison pill detector that identifies outputs with anomalous characteristics (unexpected nulls, truncation, nonsensical values)
Code Examples
Circuit Breaker Pattern
import time
import threading
from enum import Enum
from dataclasses import dataclass, field
from typing import Callable, Any, Optional
from collections import deque
class CircuitState(Enum):
CLOSED = "closed" # Normal operation
OPEN = "open" # Failing, reject requests
HALF_OPEN = "half_open" # Testing recovery
@dataclass
class CircuitBreakerConfig:
failure_threshold: float = 0.5 # 50% failure rate trips breaker
window_seconds: int = 60
cooldown_seconds: int = 30
min_calls_in_window: int = 5 # Minimum calls before evaluating
max_retries: int = 3
timeout_seconds: float = 30.0
class CircuitBreaker:
"""Prevents cascading failures between agents and services."""
def __init__(self, name: str, config: CircuitBreakerConfig, fallback: Optional[Callable] = None):
self.name = name
self._config = config
self._fallback = fallback
self._state = CircuitState.CLOSED
self._call_log: deque = deque() # (timestamp, success: bool)
self._last_failure_time = 0.0
self._lock = threading.Lock()
self._state_change_log: list[dict] = []
@property
def state(self) -> CircuitState:
with self._lock:
if self._state == CircuitState.OPEN:
# Check if cooldown has elapsed
if time.time() - self._last_failure_time >= self._config.cooldown_seconds:
self._transition(CircuitState.HALF_OPEN)
return self._state
def call(self, fn: Callable, *args, **kwargs) -> dict:
"""Execute function through the circuit breaker."""
current_state = self.state
if current_state == CircuitState.OPEN:
if self._fallback:
return {
"status": "fallback",
"circuit": self.name,
"result": self._fallback(*args, **kwargs),
}
return {
"status": "rejected",
"circuit": self.name,
"reason": "Circuit open — service unavailable",
}
try:
result = fn(*args, **kwargs)
self._record_success()
return {"status": "success", "circuit": self.name, "result": result}
except Exception as e:
self._record_failure()
if self._fallback:
return {
"status": "fallback",
"circuit": self.name,
"error": str(e),
"result": self._fallback(*args, **kwargs),
}
return {"status": "error", "circuit": self.name, "error": str(e)}
def _record_success(self):
with self._lock:
self._call_log.append((time.time(), True))
self._trim_window()
if self._state == CircuitState.HALF_OPEN:
self._transition(CircuitState.CLOSED)
def _record_failure(self):
with self._lock:
now = time.time()
self._call_log.append((now, False))
self._last_failure_time = now
self._trim_window()
failure_rate = self._current_failure_rate()
total = len(self._call_log)
if total >= self._config.min_calls_in_window and failure_rate >= self._config.failure_threshold:
self._transition(CircuitState.OPEN)
def _current_failure_rate(self) -> float:
if not self._call_log:
return 0.0
failures = sum(1 for _, success in self._call_log if not success)
return failures / len(self._call_log)
def _trim_window(self):
cutoff = time.time() - self._config.window_seconds
while self._call_log and self._call_log[0][0] < cutoff:
self._call_log.popleft()
def _transition(self, new_state: CircuitState):
old_state = self._state
self._state = new_state
self._state_change_log.append({
"circuit": self.name,
"from": old_state.value,
"to": new_state.value,
"timestamp": time.time(),
"failure_rate": self._current_failure_rate(),
})
def get_health(self) -> dict:
with self._lock:
self._trim_window()
return {
"circuit": self.name,
"state": self._state.value,
"calls_in_window": len(self._call_log),
"failure_rate": round(self._current_failure_rate(), 3),
"last_failure": self._last_failure_time,
}
Evidence Requirements
- Circuit breaker state transition logs with timestamps and failure rates
- Agent health check dashboard showing liveness, readiness, and error rates per agent
- Cascading failure post-mortem reports with root cause analysis and blast radius measurements
- Resource consumption audit showing per-agent budgets and actual usage
- Synthetic transaction monitoring results validating end-to-end pipeline health
- Retry storm detection alerts with throttling effectiveness metrics