02. Embeddings Deep Dive
Introduction
Section titled “Introduction”An embedding is a list of numbers that captures the meaning of a piece of text. Words with similar meanings have similar numbers. That’s how computers understand semantics.
This is one of the most important concepts in AI. Every search engine, recommendation system, and RAG pipeline relies on embeddings. They are the bridge between human language and mathematical computation.
Why This Concept Exists
Section titled “Why This Concept Exists”The Problem
Section titled “The Problem”Computers don’t understand words. They understand numbers.
If you give a computer the word “dog”, it sees a string of characters — not the concept of a furry four-legged animal. If you give it “puppy”, it sees another string — with no idea that “puppy” is just a baby dog.
Computers need numbers that carry meaning.
The Story
Section titled “The Story”Imagine a huge city. Every restaurant in the city has coordinates — a latitude and longitude.
Thai restaurants tend to be on the same block. Italian restaurants cluster together. Fast food places are across the street from each other.
If you’re standing at a Thai restaurant and want to find another Thai restaurant, you look at nearby coordinates. You don’t walk across town to the Italian district.
Embeddings work exactly like this.
Words and sentences with similar meanings are placed close together in a high-dimensional space. “Dog” and “Puppy” are neighbors. “Dog” and “Car” are far apart.
flowchart TD subgraph INPUT["Input"] A["'The cat sat on the mat'"] B["'A dog played in the park'"] C["'The stock market rallied today'"] end
subgraph MODEL["Embedding Model"] D["🤖 Neural Network\nconverts text → numbers"] end
subgraph OUTPUT["Output Vectors"] E["📊 [0.23, 0.87, -0.12, 0.45...]\n(768 numbers)"] F["📊 [0.21, 0.85, -0.10, 0.42...]\n(768 numbers — similar to E)"] G["📊 [-0.67, 0.12, 0.89, -0.34...]\n(768 numbers — different from E & F)"] end
A --> D --> E B --> D --> F C --> D --> G
E --> H["'cat' and 'dog' are\nclose together\n(both are pets/animals)"] G --> I["'stock market' is\nfar away\n(different topic entirely)"]
style INPUT fill:#3b82f6,color:#fff style MODEL fill:#8b5cf6,color:#fff style OUTPUT fill:#22c55e,color:#fffReal-World Analogy
Section titled “Real-World Analogy”The Coordinate System
Section titled “The Coordinate System”Think of a 2D map, but instead of latitude and longitude, every word has a position based on meaning.
On this map:
- “Dog” is at position (10, 20)
- “Puppy” is at position (11, 21) — very close to “Dog”
- “Cat” is at position (12, 19) — close to both
- “Car” is at position (80, 90) — far away
- “Stock” is at position (90, 10) — far from everything
The distance between points represents semantic difference.
In reality, embeddings don’t use 2D. They use 384, 768, 1024, 1536, or even 3072 dimensions. A 2D map captures some meaning but misses nuance. A 1536-dimensional space captures extremely fine-grained semantic relationships.
graph TD subgraph SEMANTIC["Semantic Space (Simplified)"] DOG["🐕 Dog"] --- PUPPY["🐶 Puppy"] DOG --- CAT["🐱 Cat"] DOG --- ANIMAL["🐾 Animal"] CAT --- KITTEN["😺 Kitten"]
CAR["🚗 Car"] --- TRUCK["🚛 Truck"] CAR --- VEHICLE["🚙 Vehicle"] CAR --- BIKE["🚲 Bike"]
JS["⚡ JavaScript"] --- REACT["⚛️ React"] JS --- NODE["🟢 Node.js"] JS --- CODE["💻 Programming"] end
SEMANTIC --- DISTANCE["✂️ Distance = Semantic Difference"]
style DOG fill:#f59e0b,color:#fff style CAR fill:#3b82f6,color:#fff style JS fill:#22c55e,color:#fffHow Embeddings Are Created
Section titled “How Embeddings Are Created”Text → Numbers Pipeline
Section titled “Text → Numbers Pipeline”flowchart LR TEXT["Text Input\n'The capital of France is Paris'"] --> TOKEN["Tokenizer\n(split into tokens)"] TOKEN --> MODEL["Embedding Model\n(neural network)"] MODEL --> VECTOR["Embedding Vector\n[0.45, -0.12, 0.78, ...]"] VECTOR --> USE["Used for:\n• Similarity Search\n• Clustering\n• Classification\n• Retrieval"]
style TEXT fill:#3b82f6,color:#fff style MODEL fill:#8b5cf6,color:#fff style VECTOR fill:#22c55e,color:#fff- Input text — Any text: a word, a sentence, a paragraph, or an entire document
- Tokenize — Convert text into tokens (subword units)
- Embedding model — A neural network processes the tokens and outputs a vector
- Vector — A fixed-size array of floating-point numbers representing the text’s meaning
What the Numbers Mean
Section titled “What the Numbers Mean”Every dimension in an embedding vector captures a latent feature — something the model learned to track during training. In practice, these features are not human-interpretable. You can’t say “dimension 5 means happiness.” But the pattern of all dimensions together encodes meaning.
| Dimension | Represents (hypothetically) | High Value | Low Value |
|---|---|---|---|
| 1 | How “animal-like” a word is | ”dog” = 0.9 | ”table” = 0.1 |
| 2 | How “technology-related" | "computer” = 0.8 | ”tree” = 0.2 |
| 3 | Sentiment (positive vs negative) | “happy” = 0.7 | ”sad” = -0.6 |
| … | Hundreds more subtle features | … | … |
Popular Embedding Models
Section titled “Popular Embedding Models”| Model | Provider | Dimensions | Best For | Cost |
|---|---|---|---|---|
| text-embedding-3-small | OpenAI | 512-1536 | General purpose | $ |
| text-embedding-3-large | OpenAI | 256-3072 | High accuracy | $$ |
| voyage-3 | Voyage AI | 1024-1536 | Code + multilingual | $$ |
| BGE (BAAI General Embedding) | BAAI | 384-1024 | Open-source | Free |
| E5 (Embedding from Sentences) | Microsoft | 384-768 | Academic benchmarks | Free |
| all-MiniLM-L6-v2 | Sentence Transformers | 384 | Lightweight, fast | Free |
| Cohere Embed | Cohere | 4096 | Enterprise | $$ |
| GTE | Alibaba | 768-1024 | Multilingual | Free |
flowchart TD subgraph COMPARISON["Model Size vs Quality"] L1["384d: MiniLM, BGE-small\n🚀 Fastest, 70% accuracy"] L2["768d: BGE-base, E5\n⚡ Good balance, 80% accuracy"] L3["1024d: Voyage-2, GTE\n🎯 Strong, 85% accuracy"] L4["1536d: OpenAI 3-small\n💪 Default choice, 88% accuracy"] L5["3072d: OpenAI 3-large\n🏆 Best quality, 92% accuracy"] end
L1 --> L2 --> L3 --> L4 --> L5
style L1 fill:#22c55e,color:#fff style L3 fill:#f59e0b,color:#fff style L5 fill:#ef4444,color:#fffWhy Higher Dimensions Help
Section titled “Why Higher Dimensions Help”Think of dimensions like resolution on a screen:
- 384 dimensions — 480p: You see the general shape, but details are blurry
- 768 dimensions — 720p: Clear enough for most tasks
- 1536 dimensions — 1080p: Sharp, captures nuanced differences
- 3072 dimensions — 4K: Maximum detail, but slower and more expensive
Higher dimensions allow the model to distinguish between more subtle differences. “Excited” and “eager” might overlap in 384 dimensions but be separated in 1536 dimensions.
The trade-off: More dimensions = better accuracy but:
- More storage space
- Slower search
- Higher cost (for API-based models)
Real Production Example
Section titled “Real Production Example”How Perplexity Uses Embeddings
Section titled “How Perplexity Uses Embeddings”When you ask Perplexity a question:
- Your question is converted to an embedding vector
- Perplexity searches millions of web pages for vectors close to yours
- The closest matches are retrieved
- The LLM reads those pages and generates a response with citations
sequenceDiagram participant User participant App as Perplexity App participant Embed as Embedding API participant Index as Vector Index participant LLM as LLM
User->>App: "What is retrieval-augmented generation?" App->>Embed: Convert question to embedding Embed-->>App: [0.23, 0.87, -0.12, ...] App->>Index: Find nearest vectors Index-->>App: Top 5 relevant web pages App->>LLM: Question + retrieved pages LLM-->>App: Answer with citations App-->>User: "RAG is a technique that... [Source 1] [Source 2]"Common Mistakes
Section titled “Common Mistakes”| Mistake | Why It’s Wrong |
|---|---|
| ❌ “I can use any embedding model interchangeably” | Different models have different vector spaces. Vectors from OpenAI cannot be compared directly with vectors from Cohere |
| ❌ “Higher dimensions are always better” | Higher dimensions mean more storage and slower search. Choose based on your accuracy vs speed needs |
| ❌ “Embeddings capture exact facts” | Embeddings capture meaning, not facts. Two sentences with opposite facts but similar wording will have similar embeddings |
| ❌ “I only need one embedding per document” | Long documents need multiple embeddings (one per chunk). A single embedding for a 50-page document loses too much information |
Interview Questions
Section titled “Interview Questions”Q: What is an embedding in the context of AI?
An embedding is a numerical representation of text — a fixed-size vector of floating-point numbers that captures the semantic meaning. Similar texts have similar embeddings (vectors that are close together in space).
Intermediate
Section titled “Intermediate”Q: Why do different embedding models produce different vector dimensions? How do you choose?
Different models use different architectures and training data. Higher dimensions (1536-3072) capture more nuance but cost more to store and search. Lower dimensions (384-768) are faster and cheaper but may miss subtle differences. Choose based on your accuracy requirements and latency budget.
Senior
Section titled “Senior”Q: You’re building a multilingual retrieval system. How do embeddings handle language differences?
Most embedding models are trained primarily on English. For multilingual search, you need a model specifically trained on multiple languages (e.g., BGE-m3, GTE, Voyage-multilingual). Even then, cross-language search is less accurate than same-language search. A common pattern is to detect the language and route to language-specific embedding models, or use a translation layer before embedding.
Summary
Section titled “Summary”| Concept | Key Point |
|---|---|
| Embedding | A vector of numbers representing text meaning |
| Similarity | Similar texts have similar vectors (close together) |
| Dimensions | Higher dimensions = more nuance, more cost |
| Models | OpenAI, Voyage, BGE, E5, Cohere — each with different trade-offs |
| Usage | Search, clustering, classification, RAG |
Navigation
Section titled “Navigation”Previous: 01 — Why Retrieval Systems Exist →
Next: 03 — Vector Space →