LLM09
OWASP LLM Top 10 APPLICATION GOVERNANCE MODEL

Overreliance

Description

Users or systems trust LLM outputs without verification, accepting hallucinations or errors as fact. Overreliance leads to flawed decisions and propagation of misinformation.

Risk

LLMs can produce convincing but incorrect outputs including hallucinated facts, flawed reasoning, or outdated information. Blind trust in LLM responses without verification can lead to poor decision-making, misinformation spread, or critical errors in sensitive domains like healthcare, legal, or financial services. Overreliance is particularly dangerous when LLM outputs drive automated processes.

Attack Scenarios

Medical diagnosis system relies solely on LLM analysis without physician review, leading to incorrect treatment recommendations based on hallucinated symptoms or drug interactions.

Legal research tool presents LLM-generated case citations that don't exist, which lawyers include in court filings without verification, resulting in sanctions for citing fake cases.

Mitigations

Output Verification

Implement verification steps for critical LLM outputs. Cross-reference facts with authoritative sources and flag uncertain responses.

Confidence Scoring

Display confidence levels with LLM outputs. Use techniques like self-consistency checking or multiple sampling to assess reliability.

Human Review

Require expert human review for high-stakes decisions. Design workflows that position LLM as assistant, not decision-maker.

User Education

Clearly communicate LLM limitations to users. Display disclaimers about potential errors and the need for verification in critical contexts.

Code Examples

# Good: Verification and confidence scoring
from typing import Tuple, List
import requests

class VerifiedLLMResponse:
    def __init__(self, llm_client, fact_checker_api):
        self.llm = llm_client
        self.fact_checker = fact_checker_api
    
    def get_verified_response(self, prompt: str) -> Tuple[str, float, List[str]]:
        # Generate multiple responses
        responses = [self.llm.complete(prompt) for _ in range(3)]
        
        # Self-consistency check
        confidence = self._compute_consistency(responses)
        primary_response = responses[0]
        
        # Fact-check critical claims
        citations = self._verify_facts(primary_response)
        
        # Add disclaimer for low confidence
        if confidence < 0.7:
            primary_response = (
                "[Low confidence - verify before use]\n\n" + 
                primary_response
            )
        
        return primary_response, confidence, citations
    
    def _compute_consistency(self, responses: List[str]) -> float:
        # Compare semantic similarity
        pass

Evidence Requirements

  • User interface disclaimers about LLM limitations
  • Fact-checking and verification workflows
  • Confidence scoring implementation and thresholds
  • Human review requirements for critical decisions
  • User training materials on LLM capabilities and limits