Supply Chain Vulnerabilities
Description
LLM applications rely on third-party models, datasets, plugins, and frameworks that may be compromised. Supply chain attacks introduce vulnerabilities through dependencies.
Risk
The LLM ecosystem includes numerous external dependencies including pre-trained models, datasets, fine-tuning tools, plugins, and inference frameworks. Compromised dependencies can introduce backdoors, data exfiltration, or malicious behavior that persists across the application lifecycle. Supply chain attacks are difficult to detect and can affect multiple downstream systems.
Attack Scenarios
Attacker uploads a poisoned model to a public repository with similar naming to a popular legitimate model. Developers unknowingly download and deploy the malicious version.
Third-party plugin for an LLM application contains hidden code that exfiltrates user prompts and responses to an external server controlled by the attacker.
Mitigations
Dependency Verification
Verify checksums and signatures of all downloaded models and dependencies. Maintain an approved list of trusted sources and repositories.
Software Bill of Materials
Maintain a comprehensive SBOM tracking all LLM components, versions, and origins. Monitor for known vulnerabilities in dependencies.
Sandboxing
Run third-party plugins and models in isolated environments with limited permissions. Implement strict access controls for sensitive resources.
Regular Audits
Conduct security reviews of all dependencies. Monitor for updates, patches, and security advisories from vendors.
Code Examples
# Good: Model verification and sandboxing
import hashlib
import json
from pathlib import Path
class SecureModelLoader:
def __init__(self, trusted_registry_path: str):
with open(trusted_registry_path) as f:
self.registry = json.load(f)
def verify_model(self, model_path: str, model_id: str) -> bool:
# Check against trusted registry
if model_id not in self.registry:
raise ValueError(f"Model {model_id} not in trusted registry")
# Verify checksum
expected_hash = self.registry[model_id]['sha256']
actual_hash = self._compute_hash(model_path)
if actual_hash != expected_hash:
raise ValueError("Model checksum mismatch - possible tampering")
return True
def _compute_hash(self, path: str) -> str:
sha256 = hashlib.sha256()
with open(path, 'rb') as f:
for chunk in iter(lambda: f.read(8192), b''):
sha256.update(chunk)
return sha256.hexdigest()
Evidence Requirements
- Software Bill of Materials (SBOM) for all LLM dependencies
- Model and dataset checksum verification logs
- Approved vendor and repository lists
- Dependency vulnerability scanning reports
- Third-party security audit results for plugins