Skip to content

24. Hallucinations

A hallucination is when an LLM confidently generates false or nonsensical information — producing text that sounds plausible but is factually incorrect. This is not a bug; it’s a consequence of how LLMs work.

LLMs are trained to predict the most plausible next token, not the correct one. These are fundamentally different objectives. A model that always outputs the statistically most likely text will inevitably state falsehoods with complete confidence.

flowchart TD
Q["User asks a question"] --> MODEL["🧠 LLM"]
MODEL --> PATTERN["Does this pattern\nexist in training data?"]
PATTERN -->|"Yes, strongly"| CORRECT["✅ Correct: 'Paris is\ncapital of France'"]
PATTERN -->|"Somewhat"| BLEND["⚠️ Blended: 'Paris is\ncapital of Germany'"]
PATTERN -->|"Similar patterns"| HALLUCINATE["❌ Hallucination:\n'Flubber's formula\nis C6H12O6'"]
style CORRECT fill:#22c55e,color:#fff
style HALLUCINATE fill:#ef4444,color:#fff

The model is optimized to generate text that looks like the text in its training data. Training data contains true facts, false statements, fiction, opinions, and confident guesses. The model doesn’t distinguish between them — it just learns to produce similar text.

An improv actor on stage can’t say “I don’t know.” When asked to explain quantum propulsion, they improvise: “It exploits the Heisenberg uncertainty principle using a flux capacitor.” It sounds brilliant. It’s complete nonsense. But the audience is entertained.

LLMs are improv actors. Their training objective rewards plausible text, not true text.


TypeDescriptionExample
FactualStates false information confidently”The Eiffel Tower is in London”
Source confusionMixes up related facts”Einstein invented the telephone”
Made-up referencesCites non-existent sources”According to a 2023 study by Smith et al.” (doesn’t exist)
Fictional specificsCreates plausible-sounding details”Flubber’s molecular weight is 180.16 g/mol”
TemporalWrong about timing/recencyGives old information as current

system_prompt = """
Only answer if you are certain.
If unsure, say: "I don't have enough information."
Never make up facts, statistics, or citations.
Cite sources only if they're real.
"""

Use temperature 0-0.3 for factual tasks. Higher temperatures increase randomness and thus hallucinations.

Ask the same question multiple times and check for agreement:

answers = []
for _ in range(3):
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": question}],
temperature=0.3
)
answers.append(response.choices[0].message.content)
# If answers agree, likely correct. If not, likely hallucinating.

Analyze the model’s log probabilities to estimate uncertainty:

def get_confidence(response):
total_logprob = 0
token_count = 0
for choice in response.choices:
if choice.logprobs:
for token_logprob in choice.logprobs.content:
total_logprob += token_logprob.logprob
token_count += 1
avg = total_logprob / token_count
return 1 / (1 + abs(avg)) # Higher = more confident

  1. Prompt for uncertainty — Explicitly instruct the model to say “I don’t know.”
  2. Lower temperature for facts — Use temperature 0-0.3 when accuracy matters.
  3. Validate critical outputs — Have a human review or a separate validation step.
  4. Monitor confidence scores — Track log probability-based confidence to detect unreliable responses.

ConceptKey Point
Root causeModel optimized for plausibility, not truth
TypesFactual, source confusion, made-up references, temporal
DetectionConfidence scoring, self-consistency checks
PreventionBetter prompting, lower temperature, RAG
RealityCan be reduced but never eliminated

Previous: 23 — Structured Output

Next: 25 — Phase Summary

Related Topics: