01. What is a Large Language Model?
Introduction
Section titled “Introduction”A Large Language Model (LLM) is a neural network trained on massive amounts of text to predict the next word in a sequence — and that simple ability, at massive scale, produces systems that can write, reason, translate, summarize, and hold conversations.
LLMs are the technology behind ChatGPT, Claude, Gemini, GitHub Copilot, and every modern AI assistant you have heard of. They represent the culmination of decades of AI research, combining the scale of the internet’s text data with the power of Transformer neural networks and the parallel processing of modern GPU hardware.
graph TD AI["🤖 Artificial Intelligence\n(Broad field: machines acting smart)"] ML["📊 Machine Learning\n(Learn from data without explicit rules)"] DL["🧠 Deep Learning\n(Learn via multi-layer neural networks)"] TF["⚡ Transformer\n(Attention-based architecture from 2017)"] LLM["💬 Large Language Model\n(Transformer trained on internet-scale text)"]
AI --> ML ML --> DL DL --> TF TF --> LLM
style AI fill:#3b82f6,color:#fff style ML fill:#8b5cf6,color:#fff style TF fill:#f59e0b,color:#fff style LLM fill:#22c55e,color:#fffWhy This Exists
Section titled “Why This Exists”The Problem Before LLMs
Section titled “The Problem Before LLMs”Before LLMs, computers could not understand natural language. They could process text — search for keywords, match patterns, follow rules — but they could not understand meaning.
Traditional NLP (Pre-2018):
- Keyword matching: “I’m not happy” → searches for “happy” → misses the “not”
- Rule-based systems: Thousands of hand-written grammar rules that break on every edge case
- Statistical models: Naive Bayes, SVM — required manual feature engineering, could only handle simple tasks
- No concept of context, sarcasm, analogy, or reasoning
Example — Customer support before LLMs:
User: “My order hasn’t arrived and it’s been three weeks. I’m extremely frustrated.”
Traditional bot: “I found the keyword ‘order’. Did you mean to track your order? Type your order number.”
The bot could not detect frustration, understand the time reference, or empathize.
What LLMs Solved
Section titled “What LLMs Solved”LLMs brought three breakthroughs:
- Understanding context — They don’t match keywords; they understand the meaning behind words
- Generalization — One model can summarize, translate, code, write poetry, and answer questions — no separate model needed for each task
- Emergent abilities — At sufficient scale, models develop abilities that were never explicitly programmed: reasoning, analogies, chain-of-thought, even theory of mind
flowchart LR subgraph BEFORE["Before LLMs (Pre-2018)"] A1["Task 1: Sentiment\n→ separate model"] A2["Task 2: Translation\n→ separate model"] A3["Task 3: Summarization\n→ separate model"] A4["Task 4: Q&A\n→ separate model"] end
subgraph AFTER["After LLMs (2020+)"] B["One Large Language Model\n(pretrained on internet text)"] B --> C1["Sentiment"] B --> C2["Translation"] B --> C3["Summarization"] B --> C4["Q&A"] B --> C5["Code"] B --> C6["Chat"] end
style BEFORE fill:#ef4444,color:#fff style AFTER fill:#22c55e,color:#fff style B fill:#8b5cf6,color:#fffDeconstructing “Large Language Model”
Section titled “Deconstructing “Large Language Model””Every word in the name tells you something important.
| Word | Meaning | Why It Matters |
|---|---|---|
| Large | The model has billions of parameters and was trained on massive data | Small models can’t do what large models can — scale creates emergent abilities |
| Language | The model processes human language — text in, text out | It understands words, sentences, documents, and conversations |
| Model | A mathematical approximation of the patterns in its training data | It’s not a database, it’s a pattern predictor — it models the probability distribution of language |
Put together: A Large Language Model is a mathematical system, trained on billions of text documents, that predicts the next word in a sequence with such accuracy that it appears to understand and generate human language.
flowchart LR A["LARGE\n(billions of parameters)"] --> D["🧠"] B["LANGUAGE\n(trained on internet text)"] --> D C["MODEL\n(probability distribution)"] --> D D --> E["Ability to generate\nhuman-quality text"] D --> F["Ability to understand\ncontext and meaning"] D --> G["Ability to reason\nand follow instructions"]
style A fill:#ef4444,color:#fff style B fill:#3b82f6,color:#fff style C fill:#f59e0b,color:#fff style D fill:#8b5cf6,color:#fff style E fill:#22c55e,color:#fff style F fill:#22c55e,color:#fff style G fill:#22c55e,color:#fffReal-World Analogy
Section titled “Real-World Analogy”The Librarian Who Read Everything
Section titled “The Librarian Who Read Everything”Imagine a librarian who has read every book ever written — every novel, textbook, blog post, poem, research paper, Wikipedia article, and Reddit thread.
You give her the start of a sentence: “The cat sat on the…”
She doesn’t think about cats or mats. She has read “The cat sat on the mat” 10,000 times, “The cat sat on the roof” 500 times, and “The cat sat on the fence” 200 times.
Based on everything she has read, she predicts the most likely next word: “mat.”
Now expand this: She doesn’t just predict one word. She predicts the next word after that. And the next. One word at a time, she builds entire paragraphs, essays, stories — all based on the statistical patterns she absorbed from her training.
The key insight: She is not thinking. She is not conscious. She is doing what she has done a trillion times: pattern matching at an unimaginable scale.
The Difference Between NLMs, ML, DL, Transformers, and LLMs
Section titled “The Difference Between NLMs, ML, DL, Transformers, and LLMs”| Technology | What It Does | Example | Limitation |
|---|---|---|---|
| Traditional NLP | Follows hand-written rules for text | Regex, keyword matching | Cannot handle ambiguity or nuance |
| Machine Learning | Learns patterns from labeled data | Spam classifier, regression | Requires feature engineering; limited on text |
| Deep Learning | Multi-layer neural networks learn hierarchical features | CNNs for images, RNNs for sequences | RNNs are slow; can’t parallelize |
| Transformer | Architecture using self-attention to process sequences in parallel | BERT (understanding), GPT (generation) | The foundation, not the final product |
| LLM | A massive Transformer trained on internet-scale text | ChatGPT, Claude, Gemini | Expensive to run; can hallucinate |
gantt title Evolution of Language AI dateFormat YYYY axisFormat %Y section Traditional Rule-based NLP :1950, 50y Statistical NLP :1990, 20y ML-based NLP :2000, 15y section Deep Learning RNNs / LSTMs :2013, 5y Attention Mechanism :2015, 2y section Transformers Transformer Paper :2017, 1y BERT / GPT-1 :2018, 1y GPT-2 :2019, 1y GPT-3 :2020, 1y ChatGPT / GPT-4 :2022, 2y Claude / Gemini / Llama :2023, 2yHow LLMs Work (High Level)
Section titled “How LLMs Work (High Level)”flowchart TD IN["Input Text\n(User prompt or query)"] --> TK["Tokenizer\n(Text → Numbers)"] TK --> TF["Transformer\n(Billions of parameters)"] TF --> PR["Prediction\n(Probability over all tokens)"] PR --> TK2["Sample next token\n(Choose the next word)"] TK2 --> D{"Is response\ncomplete?"} D -->|"No — repeat"| IN D -->|"Yes"| OUT["Final Response\n(Generated text)"]
style IN fill:#3b82f6,color:#fff style TK fill:#8b5cf6,color:#fff style TF fill:#f59e0b,color:#fff style PR fill:#ef4444,color:#fff style TK2 fill:#8b5cf6,color:#fff style OUT fill:#22c55e,color:#fff- Input — A user types a prompt: “Explain quantum computing in simple terms”
- Tokenization — The text is split into tokens (words/subwords) and converted to numbers
- Transformer — The numbers pass through 100+ layers of neural network, each layer refining the representation using self-attention and feed-forward computation
- Prediction — The final layer outputs a probability distribution over the entire vocabulary (~50,000–200,000 possible tokens)
- Sampling — The model selects one token (the most likely, or randomly weighted by probability)
- Repeat — The new token is appended to the input, and steps 2–5 repeat until the model outputs a stop token or reaches the token limit
Major LLM Families
Section titled “Major LLM Families”Here is a comparison of the major LLM families available today:
mindmap root((LLM Families)) OpenAI GPT-3.5 GPT-4 GPT-4o o1 / o3 (reasoning) Best overall quality Anthropic Claude 3 Haiku Claude 3 Sonnet Claude 3 Opus Claude 3.5 / 4 Focus on safety Google DeepMind Gemini 1.5 Pro Gemini 1.5 Flash Gemini 2.0 1M token context Meta LLaMA 2 LLaMA 3 / 3.1 Open-weight Good for self-hosting Mistral AI Mistral 7B Mixtral 8x7B Mistral Large Efficient architectures Alibaba Qwen / Qwen 2 Strong on coding Multilingual DeepSeek DeepSeek V2 / V3 R1 (reasoning) Extremely cost-efficient Microsoft Phi 3 / 3.5 Small but capable Runs on phonesDetailed Comparison
Section titled “Detailed Comparison”| Model | Creator | Size Range | Context Window | Cost | Best For | Open? |
|---|---|---|---|---|---|---|
| GPT-4o | OpenAI | Unknown | 128K | $$$ | General, coding, reasoning | No |
| Claude 3 Opus | Anthropic | Unknown | 200K | $$$ | Long docs, analysis, safety | No |
| Gemini 1.5 Pro | Unknown | 1M | $$ | Very long context, multimodal | No | |
| LLaMA 3.1 405B | Meta | 8B, 70B, 405B | 128K | Free | Self-hosting, research | Yes |
| Mistral Large | Mistral | Unknown | 128K | $$ | Efficiency, multilingual | No |
| Mixtral 8x7B | Mistral | 46B total (MoE) | 32K | Low | Self-hosting on modest hardware | Yes |
| Qwen 2.5 72B | Alibaba | 7B, 14B, 72B | 128K | Low-Med | Coding, math, Chinese/English | Yes |
| DeepSeek V3 | DeepSeek | 671B (MoE) | 128K | Very Low | Cost-efficient, coding | Yes |
| Phi 3 | Microsoft | 3.8B, 14B | 128K | Very Low | On-device, simple tasks | Yes |
| Gemma 2 | 2B, 9B, 27B | 8K | Free | Lightweight research | Yes |
Note on “Open”: Strictly open-source models (open weights + open data + open training code) are rare. Most models listed as “Yes” here are “open-weight” — you can download and run the model weights, but the training data and training code are not public.
What Makes LLMs Feel Intelligent?
Section titled “What Makes LLMs Feel Intelligent?”flowchart LR A["Input: 'What is the\ncapital of France?'"] --> B["LLM doesn't 'know'\nParis is the capital"] B --> C["LLM has seen\n'capital of France → Paris'\nmillions of times"] C --> D["LLM predicts 'Paris'\nwith 99.9% probability"] D --> E["You: 'Wow, it knows\ngeography!'"] E --> F["Reality: It's just\na statistical pattern"]
style A fill:#3b82f6,color:#fff style D fill:#22c55e,color:#fff style F fill:#ef4444,color:#fffLLMs feel intelligent because:
-
Pattern completion at scale — When you see a friend and say “Long time, no…” you complete “see” automatically. LLMs do this with every concept, across billions of patterns, simultaneously.
-
Emergent abilities — When a model reaches a certain size (~70B+ parameters), it develops abilities that smaller versions don’t have: reasoning, step-by-step thinking, analogical reasoning. No one programmed these — they emerged from scale.
-
Massive training data — The model has been exposed to more text than any human could read in 100 lifetimes. It has seen every argument, every explanation, every analogy — and can recombine them.
-
In-context learning — LLMs can learn a new task from a few examples in the prompt, without any weight updates. This makes them appear to “understand” instructions instantly.
The Critical Distinction: Prediction vs. Understanding
Section titled “The Critical Distinction: Prediction vs. Understanding”LLMs are prediction engines — not thinking machines.
- They do not understand truth
- They do not have beliefs
- They do not have consciousness
- They do not have intentions
- They do not have memory (beyond the context window)
When an LLM says “I think the answer is X,” it is not thinking. It is generating a sequence of tokens that statistically matches the pattern of a thoughtful person giving an answer.
What LLMs Can and Cannot Do
Section titled “What LLMs Can and Cannot Do”Capabilities
Section titled “Capabilities”| Capability | Example | Why It Works |
|---|---|---|
| Summarization | ”Summarize this 100-page PDF” | Training data contained countless examples of summaries |
| Translation | ”Translate to Spanish” | Training data was multilingual |
| Code generation | ”Write a Python function to sort a list” | Training data contained GitHub repositories |
| Creative writing | ”Write a poem about AI in the style of Shakespeare” | Training data contained literature in every style |
| Reasoning | Solve a logic puzzle step by step | At scale, models learn to decompose problems |
| Role-playing | ”Act as a history professor” | Training data contained dialogues, role-play, and instruction examples |
| Tool use | ”Find today’s weather” (calls a function) | Fine-tuned for function calling |
Limitations
Section titled “Limitations”| Limitation | Example | Why |
|---|---|---|
| Hallucination | States fake facts confidently | Model prioritizes plausible text over truth |
| No true reasoning | Fails on simple math if pattern is unusual | It’s pattern matching, not mathematical reasoning |
| No memory beyond context | Forgets what you said 1000 tokens ago | Fixed context window size |
| No real-time awareness | Doesn’t know today’s news (unless using tools) | Training data has a cutoff date |
| Biases from training data | Exhibits social biases | Learns biases present in internet text |
| No causal understanding | Can’t tell correlation vs causation | Predicts text, doesn’t model the world |
| Sensitive to prompt wording | Same question rephrased → different answer | No grounded understanding of meaning |
Python Example: Using an LLM
Section titled “Python Example: Using an LLM”# The simplest way to use an LLM — through an API
# Install: pip install openaifrom openai import OpenAI
client = OpenAI(api_key="your-api-key") # Get key from platform.openai.com
response = client.chat.completions.create( model="gpt-4o", messages=[ {"role": "system", "content": "You are a helpful tutor who explains things simply."}, {"role": "user", "content": "What is an LLM in one paragraph?"} ])
print(response.choices[0].message.content)# LLM stands for Large Language Model — a neural network trained on massive text data# that predicts the next word in a sequence. By repeating this prediction millions of# times, LLMs can generate coherent text, answer questions, write code, and hold# conversations. They don't "think" or "understand" — they excel at statistical# pattern matching at an enormous scale.JavaScript Example: Using an LLM
Section titled “JavaScript Example: Using an LLM”// Install: npm install openaiimport OpenAI from 'openai';
const openai = new OpenAI({ apiKey: 'your-api-key' });
async function askLLM() { const response = await openai.chat.completions.create({ model: 'gpt-4o', messages: [ { role: 'system', content: 'You explain concepts simply.' }, { role: 'user', content: 'What is an LLM in one paragraph?' } ] });
console.log(response.choices[0].message.content); // LLM stands for Large Language Model...}
askLLM();Architecture Diagram: LLM as a System
Section titled “Architecture Diagram: LLM as a System”flowchart TD subgraph CLIENT["Client Layer"] UI["User Interface\n(Chat, API call, IDE plugin)"] end
subgraph API["API / Gateway Layer"] AUTH["Authentication\n& Rate Limiting"] RTR["Router\n(selects model endpoint)"] end
subgraph INFERENCE["Inference Layer"] TOK["Tokenizer\n(text → token IDs)"] TF["Transformer\n(autoregressive generation)"] SAM["Sampling\n(top-k, top-p, temperature)"] DETOK["Detokenizer\n(token IDs → text)"] end
subgraph MODEL["Model Storage"] W["Model Weights\n(stored on GPU VRAM or disk)"] CFG["Config\n(architecture settings)"] end
UI --> AUTH AUTH --> RTR RTR --> TOK TOK --> TF W --> TF CFG --> TF TF --> SAM SAM --> DETOK DETOK --> UI
style CLIENT fill:#3b82f6,color:#fff style API fill:#8b5cf6,color:#fff style INFERENCE fill:#f59e0b,color:#fff style MODEL fill:#22c55e,color:#fffBest Practices
Section titled “Best Practices”- Choose the right model for the task — GPT-4o is great for complex reasoning; smaller models like Claude Haiku or GPT-4o mini are faster and cheaper for simple tasks
- Match model size to hardware — Don’t try to run a 405B parameter model on a laptop; use API-based models or quantized smaller models for local use
- Understand the pricing model — LLM APIs charge per token (input + output); cost can accumulate fast with long prompts or high traffic
- Use system prompts — Set the behavior, tone, and constraints in a system message rather than relying on the user to describe what they want
- Always validate outputs — LLMs can hallucinate; verify facts, test code, review generated content before using it
- Monitor for prompt injection — Users may try to override your system instructions; implement input sanitization and output filtering
Common Misconceptions
Section titled “Common Misconceptions”| Misconception | Truth |
|---|---|
| ”LLMs understand language like humans” | LLMs have no understanding, consciousness, or awareness — they perform statistical pattern matching |
| ”LLMs have access to the internet” | No — they only know what was in their training data (cutoff date varies) |
| “LLMs will replace all jobs” | LLMs are tools that augment humans, not replace them — they lack agency, physical presence, and true reasoning |
| ”Bigger models are always better” | Larger models are more capable but also more expensive and slower; smaller models can outperform on specific tasks with fine-tuning |
| ”Open-source LLMs are as good as GPT-4” | Open models are catching up fast but typically lag 6–18 months behind the best proprietary models |
| ”LLMs can do math perfectly” | LLMs are not calculators — they can approximate but may fail on simple arithmetic; use code execution for precise math |
Interview Questions
Section titled “Interview Questions”Q: What does LLM stand for and what does it do?
LLM stands for Large Language Model. It is a neural network trained on massive amounts of text data that generates text by predicting the next word in a sequence. By repeating this prediction step-by-step, it can write essays, answer questions, summarize documents, and hold conversations.
Q: What is the difference between a traditional ML model and an LLM?
Traditional ML models are trained for specific tasks (like spam classification or sentiment analysis) and require labeled data and feature engineering for each task. LLMs are general-purpose: one single model trained on diverse internet text can perform hundreds of different tasks — translation, summarization, code generation, Q&A, creative writing — without needing separate training for each one.
Medium
Section titled “Medium”Q: Why are LLMs called “large”? What makes them large?
“Large” refers to three things: (1) the number of parameters — modern LLMs have tens to hundreds of billions of parameters, (2) the training data — they are trained on trillions of tokens from the internet, books, and code repositories, and (3) the computational resources — training requires thousands of GPUs running for weeks or months. The combination of these three factors creates emergent abilities that smaller models lack.
Q: Explain why LLMs appear intelligent but are not actually thinking.
LLMs appear intelligent because they have been trained on virtually every text ever written by humans. When asked a question, the model doesn’t “think” about the answer — it calculates the most likely sequence of tokens based on statistical patterns in its training data. A useful analogy: a LLM is like a parrot that has heard every conversation ever recorded. It can say the right thing in the right context because it has heard it so many times, not because it understands what it is saying. This is why LLMs can confidently state falsehoods (hallucinations) — they are producing statistically plausible text, not verifying truth.
Q: What are emergent abilities in LLMs, and why do they matter?
Emergent abilities are capabilities that appear only when a model reaches a certain size threshold — they are not present in smaller versions of the same architecture. Examples include: chain-of-thought reasoning, in-context learning (learning a new task from examples without weight updates), code generation, and solving arithmetic problems. These abilities are notoriously unpredictable — researchers don’t know exactly at what scale they will appear or which abilities will emerge. This is significant because it means our understanding of LLM capabilities is always incomplete: a model 2x larger may do things we didn’t expect, and a model 2x smaller may lack abilities we assumed were basic.
Q: What is the difference between a foundation model and a fine-tuned model?
A foundation model (e.g., GPT-3 base, LLaMA-3 base) is trained on raw internet text using self-supervised learning — next-token prediction. It is extremely capable but hard to use because it continues text rather than following instructions. A fine-tuned model (e.g., GPT-3.5-turbo, LLaMA-3-Chat) is a foundation model that has been further trained on instruction-response pairs and preference data (RLHF). Fine-tuning aligns the model to be helpful, follow instructions, and refuse harmful requests. The foundation model is a raw intelligence engine; the fine-tuned model is a polished assistant. Most users interact with fine-tuned models.
Summary
Section titled “Summary”| Concept | Key Point |
|---|---|
| LLM | Large Language Model — neural network trained on internet-scale text to predict the next token |
| ”Large” | Billions of parameters, trillions of training tokens, thousands of GPUs |
| ”Language” | Processes and generates human language text |
| ”Model” | A mathematical approximation of language patterns — a prediction engine, not a thinking machine |
| Key technology | Transformer architecture with self-attention |
| How it works | Tokenize → predict next token → append → repeat |
| Emergent abilities | Capabilities that appear at scale (reasoning, in-context learning) |
| Major families | GPT (OpenAI), Claude (Anthropic), Gemini (Google), LLaMA (Meta), Mistral, Qwen, DeepSeek |
| Not thinking | LLMs do NOT understand, believe, or know — they generate statistically plausible text |
| Hallucination | Confident falsehoods — the model generates plausible but incorrect information |
Navigation
Section titled “Navigation”Previous: Phase 3 — Deep Learning Cheat Sheet
Next: 02 — How Language Models Work
Related Topics:
Practice Questions:
- List 5 different LLMs and describe how they differ
- Explain why an LLM is a “prediction engine” rather than a “thinking machine”
- What emergent ability of LLMs is most surprising to you and why?
- Draw the high-level flow of how an LLM generates a response
- Compare a traditional ML spam filter with an LLM — what can the LLM do that the spam filter cannot?
Further Reading: