LLM10
OWASP LLM Top 10 MODEL

Model Theft

Description

Attackers extract or replicate proprietary LLMs through API abuse, model inversion, or unauthorized access. Stolen models enable competitive harm and facilitate further attacks.

Risk

Proprietary LLMs represent significant investment in training data, compute, and tuning. Theft through model extraction queries, weight exfiltration, or inference API abuse can enable competitors to replicate capabilities without investment. Stolen models also facilitate offline attack development, bypass rate limits, and may reveal sensitive training data.

Attack Scenarios

Attacker systematically queries a proprietary LLM API with carefully designed inputs to extract knowledge and behavior, training a functionally equivalent model at fraction of original cost.

Insider with access to model weights exfiltrates files to external storage, selling the proprietary model to competitors or publishing it publicly.

Mitigations

Rate Limiting

Implement strict rate limits to prevent extraction via systematic querying. Monitor for unusual access patterns indicating scraping attempts.

Watermarking

Embed fingerprints in model outputs to enable detection of unauthorized copies. Use techniques that survive model distillation.

Access Controls

Restrict access to model weights and architecture. Use encryption for stored models and secure enclaves for inference where appropriate.

Query Monitoring

Detect and block model extraction attempts. Look for patterns like high query volume, systematic coverage of input space, or queries designed to probe decision boundaries.

Code Examples

# Good: Query monitoring for extraction attempts
from collections import defaultdict
import time
from typing import Optional

class ModelTheftDetector:
    def __init__(self, max_queries=1000, window=3600):
        self.max_queries = max_queries
        self.window = window
        self.query_history = defaultdict(list)
        self.similarity_threshold = 0.8
    
    def check_query(self, user_id: str, query: str) -> Optional[str]:
        now = time.time()
        
        # Clean old queries
        self.query_history[user_id] = [
            (t, q) for t, q in self.query_history[user_id]
            if now - t < self.window
        ]
        
        # Check volume
        if len(self.query_history[user_id]) >= self.max_queries:
            return "Rate limit exceeded - possible extraction attempt"
        
        # Check for systematic probing
        if self._is_systematic_probing(user_id, query):
            return "Systematic probing detected - access blocked"
        
        self.query_history[user_id].append((now, query))
        return None
    
    def _is_systematic_probing(self, user_id: str, query: str) -> bool:
        # Detect patterns like grid search of input space
        pass

Evidence Requirements

  • Rate limiting policies and enforcement logs
  • Model access control configurations and audit trails
  • Watermarking implementation details (if applicable)
  • Query pattern analysis and anomaly detection alerts
  • Incident response records for suspected theft attempts