Module 4: Inference
Module 4: Inference
Section titled “Module 4: Inference”Everything that happens when you send a prompt to an LLM — from decoding strategies to streaming, function calling, and managing hallucinations.
Overview
Section titled “Overview”Module 4 is about the inference pipeline — what happens after training is complete and you send a prompt. You’ll learn how text is generated token by token, how different decoding strategies affect output quality, how streaming works, how function calling enables tool use, and why hallucinations happen.
Learning Objectives
Section titled “Learning Objectives”After completing this module, you will be able to:
- ✅ Describe the inference pipeline step by step
- ✅ Compare greedy decoding, beam search, and sampling
- ✅ Explain how temperature, top-k, and top-p affect output
- ✅ Understand how streaming delivers tokens in real-time
- ✅ Explain function calling architecture
- ✅ Describe structured output methods (JSON mode, grammar-based)
- ✅ Understand why hallucinations occur and mitigation strategies
Prerequisites
Section titled “Prerequisites”| Requirement | Level |
|---|---|
| Module 3: Training | ✅ Required |
| Understanding of probability | ⭐ Recommended |
| API usage experience | 🔄 Helpful |
Estimated Time
Section titled “Estimated Time”| Activity | Time |
|---|---|
| Reading lessons | 3 hours |
| Practice exercises | 45 minutes |
| Mini quiz | 15 minutes |
| Total | ~4 hours |
Lessons
Section titled “Lessons”| # | Lesson | 🔥 | Description |
|---|---|---|---|
| 18 | Inference | 🔥 Must Know | The complete inference pipeline |
| 19 | Decoding Strategies | 🧠 Core Concept | Greedy, beam search, top-k, top-p |
| 20 | Temperature, Top-K & Top-P | 🧠 Core Concept | Controlling randomness in generation |
| 21 | Streaming | 💼 Production | Real-time token delivery |
| 22 | Function Calling | 💼 Production | Connecting LLMs to external tools |
| 23 | Structured Output | 💼 Production | JSON mode, constrained generation |
| 24 | Hallucinations | 🧠 Core Concept | Why models fabricate and how to reduce it |
Inference Pipeline
Section titled “Inference Pipeline”flowchart TD IN["User Prompt"] --> TOK["Tokenizer\n(Text → Token IDs)"] TOK --> TF["Transformer\n(Multiple decoder blocks)"] TF --> LOGITS["Logits\n(Unnormalized scores)"] LOGITS --> SAMPLING["Decoding Strategy\n(Greedy / Sampling / Beam)"] SAMPLING --> NEXT["Next Token ID"] NEXT --> DETOK["Detokenizer\n(Token ID → Text)"] DETOK --> OUT["Output Text"] NEXT --> APPEND["Append to Input"] APPEND --> TOK OUT --> STREAM{"Streaming?"} STREAM -->|"Yes"| EMIT["Emit token to client"] STREAM -->|"No"| WAIT["Wait for complete response"]
style IN fill:#3b82f6,color:#fff style TOK fill:#8b5cf6,color:#fff style TF fill:#f59e0b,color:#fff style LOGITS fill:#ef4444,color:#fff style SAMPLING fill:#22c55e,color:#fff style OUT fill:#3b82f6,color:#fffKey Concepts
Section titled “Key Concepts”- Autoregressive generation: Each token depends on all previous tokens
- Decoding strategies: Greedy (deterministic), beam search (multiple paths), sampling (stochastic)
- Temperature: Controls the “sharpness” of the probability distribution
- Top-k: Only sample from the k most likely tokens
- Top-p (nucleus): Only sample from tokens whose cumulative probability exceeds p
- Streaming: Server-Sent Events (SSE) deliver tokens as they’re generated
- Function calling: LLM outputs structured arguments for predefined functions
- Constrained decoding: Grammar-based generation that guarantees valid output format
- Hallucinations: Model produces plausible-sounding but incorrect information
Module Summary
Section titled “Module Summary”In this module, you learned:
- Inference is autoregressive — one token at a time, each depending on previous tokens
- Decoding strategies balance creativity vs coherence
- Temperature, top-k, top-p control output randomness
- Streaming improves user experience with real-time output
- Function calling enables LLMs to interact with external systems
- Structured output ensures machine-parseable responses
- Hallucinations are inherent to LLMs and require systematic mitigation
Practice Questions
Section titled “Practice Questions”- Why is greedy decoding not always the best strategy for creative tasks?
- Explain the difference: what happens when temperature = 0 vs temperature = 1 vs temperature = 2?
- How would you design a function calling schema for a weather API?
- What are three strategies to reduce hallucinations?
Interview Questions
Section titled “Interview Questions”-
Q: Why can’t LLMs generate text in parallel? Why is it sequential?
- A: Because each token depends on all previous tokens. You can’t know token #100 without first generating tokens 1-99. This is the fundamental constraint of autoregressive generation.
-
Q: How does temperature 0 differ from greedy decoding?
- A: At temperature → 0, the softmax approaches argmax, which is equivalent to greedy decoding. Both select the most likely token deterministically.
-
Q: What are the security implications of function calling?
- A: Unrestricted function calling can lead to prompt injection where users trick the model into calling sensitive functions. Mitigations include input sanitization, least-privilege function design, and human-in-the-loop for destructive operations.
Next Steps
Section titled “Next Steps”➡️ Continue to Module 5: Revision & Project →