Module 1 Summary: LLM Foundations
Module 1 Summary: LLM Foundations
Section titled “Module 1 Summary: LLM Foundations”Quick Recap
Section titled “Quick Recap”| Concept | Key Point |
|---|---|
| LLM | Neural network trained on internet-scale text to predict the next token |
| Language Model | Assigns probabilities to sequences of words |
| Tokenization | Converts text into numbers (tokens) the model can process |
| Context Window | Maximum number of tokens an LLM can process at once |
Key Equations
Section titled “Key Equations”- Parameters vs Data ratio: Modern LLMs have ~2 tokens of training data per parameter
- Context window growth: 512 (GPT-1) → 2048 (GPT-3) → 128K (GPT-4) → 1M (Gemini 1.5)
- Token ratio: ~0.75 words per token (English)
Practice Questions
Section titled “Practice Questions”- Explain why an LLM is considered a “prediction engine” rather than a “thinking machine.”
- What happens to the probability of the entire sequence as you multiply conditional probabilities?
- Why can’t you use raw word IDs as tokens? What’s the vocabulary size problem?
- If a model has a context window of 128K tokens, and each token is ~4 characters, how many pages of text can it process? (Assume 3000 characters per page)
-
What does LLM stand for?
- a) Large Logic Machine
- b) Large Language Model
- c) Linear Language Module
- d) Latent Learning Model
- Answer: b
-
Which of the following is NOT an emergent ability of large language models?
- a) Chain-of-thought reasoning
- b) In-context learning
- c) Database querying
- d) Code generation
- Answer: c
-
What is the primary purpose of tokenization?
- a) To encrypt the input text
- b) To convert text into numerical IDs
- c) To compress the input
- d) To translate between languages
- Answer: b
-
What happens when you send a prompt longer than the context window?
- a) The model expands its context window
- b) The prompt is truncated (oldest tokens dropped)
- c) The model crashes
- d) The output doubles in length
- Answer: b
Interview Questions
Section titled “Interview Questions”- Q: What is the difference between a foundation model and a fine-tuned model?
- Q: Why do larger models show emergent abilities that smaller models don’t?
- Q: Explain autoregressive generation in one sentence.