Skip to content

05. Introduction to Vector Databases

A vector database is a database designed specifically for storing and searching vectors. It’s optimized for similarity search — finding the “nearest neighbor” vectors to a query — at massive scale.

You have a million documents. Each document is a 1536-dimensional vector. You need to find the 5 most similar documents to a query in under 100 milliseconds. MySQL can’t do that. MongoDB can’t do that. Elasticsearch can’t do that. Vector databases were created to solve exactly this problem.


Traditional databases are designed for exact matches and range queries. They can find “WHERE price = 100” or “WHERE created_at > ‘2024-01-01’” instantly. But they cannot efficiently answer “Find the 5 rows most semantically similar to this 1536-dimensional vector.”

OperationSQL DatabaseVector Database
Find exact match by ID✅ Fast✅ Fast
Find rows where price > 100✅ Fast❌ Can’t do
Sort by column✅ Fast❌ Not built for
Find semantically similar text❌ Can’t do✅ Fast
Hybrid: filter by tag + semantic search❌ Hard✅ Supported

Imagine sorting 1 million grains of sand by color. You lay them out in a line and examine each one. That’s what a traditional database does when you ask for a similarity search — it looks at every single row, one by one.

Now imagine a special sorting tray that groups sand by color family automatically. You don’t look at every grain — you go directly to the red section and find the closest red grains. That’s what a vector database does.


A library has millions of books. You want to find books about “ancient Egyptian architecture.”

Traditional Database approach: Look at every book’s metadata and check if “ancient”, “Egyptian”, and “architecture” appear in the description. Slow, and you’ll miss “The Great Monuments of Old Egypt” — which doesn’t contain any of those exact words.

Vector Database approach: Every book has already been analyzed and placed on a “meaning map.” Books about similar topics are shelved together. You find your location on the map and walk to the nearest shelf. You find relevant books in milliseconds.

flowchart TD
subgraph TRADITIONAL["Traditional Database"]
A1["Query: 'Ancient Egypt'"] --> A2["Sequential scan\nALL 1M rows"]
A2 --> A3["WHERE title LIKE '%Ancient%'\nAND title LIKE '%Egypt%'"]
A3 --> A4["❌ Misses: 'Pyramids of the Nile'\n🐌 2 seconds"]
end
subgraph VECTOR["Vector Database"]
B1["Query: 'Ancient Egypt'"] --> B2["Embed → vector"]
B2 --> B3["ANN search on\npre-built index"]
B3 --> B4["✅ Finds: 'Pyramids of the Nile'\n⚡ 50 milliseconds"]
end
style TRADITIONAL fill:#ef4444,color:#fff
style VECTOR fill:#22c55e,color:#fff

AspectSQL (PostgreSQL, MySQL)MongoDBElasticsearchVector DB (Pinecone, Qdrant)
Primary dataTables, rows, columnsJSON documentsText logsVectors (+ metadata)
Search methodB-tree indexesB-tree indexesInverted indexANN (HNSW, IVF)
Similarity search❌ Not supported❌ Not natively❌ Limited✅ Built-in
Filtering✅ Excellent✅ Excellent✅ Good✅ Supported
ScalabilityGoodGoodGoodExcellent for vectors
Best forTransactionsDocumentsLogs, full-textSemantic search

DatabaseTypeOpen SourceBest ForKey Strength
PineconeSaaS❌Production RAGZero maintenance, managed
QdrantSelf-hosted/SaaS✅PerformanceWritten in Rust, very fast
WeaviateSelf-hosted/SaaS✅Hybrid searchBuilt-in vectorizer modules
MilvusSelf-hosted✅Large scaleBillion-scale ANN
ChromaEmbedded✅PrototypingSimple API, runs locally
FAISSLibrary (not a DB)✅ResearchFastest index, no built-in persistence
pgvectorPostgreSQL extension✅Python/ML stackStore vectors in existing Postgres
flowchart TD
subgraph ECOSYSTEM["Vector Database Ecosystem"]
MANAGED["☁️ Managed Services\nPinecone, Weaviate Cloud,\nQdrant Cloud"]
SELF["🏗️ Self-Hosted\nQdrant, Weaviate,\nMilvus, Chroma"]
EMBEDDED["📦 Embedded\nChroma, FAISS,\npgvector"]
RESEARCH["🔬 Research/Library\nFAISS, ScaNN,\nHNSWlib"]
end
MANAGED --> USE["Production RAG\nChatGPT-like apps"]
SELF --> USE
EMBEDDED --> PROT["Prototyping\nLocal dev, small scale"]
RESEARCH --> BENCH["Benchmarking\nCustom solutions"]
style MANAGED fill:#3b82f6,color:#fff
style SELF fill:#8b5cf6,color:#fff
style EMBEDDED fill:#22c55e,color:#fff
style RESEARCH fill:#f59e0b,color:#fff

A collection is like a table in SQL — a logical grouping of vectors. You might have one collection for “support-docs”, another for “product-catalog”, and another for “user-notes”.

An index is the data structure that enables fast search. The most popular is HNSW (Hierarchical Navigable Small World) — a graph-based index that can search billions of vectors in milliseconds.

flowchart LR
subgraph INDEX["How HNSW Index Works"]
L1["Layer 1\n(Few nodes, long jumps)"]
L2["Layer 2\n(More nodes)"]
L3["Layer 3\n(All nodes, fine detail)"]
end
QUERY["Query"] --> L1
L1 --> L2
L2 --> L3
L3 --> RESULT["Nearest Neighbor"]
style INDEX fill:#8b5cf6,color:#fff

How HNSW works (intuitively):

  1. Start at the top layer (fewest nodes, longest connections)
  2. Find the closest node at this layer
  3. Move to the next layer (more nodes, shorter connections)
  4. Refine the search
  5. Repeat until you reach the bottom layer with the exact nearest neighbor

This is like finding a city on a map: first zoom out to see continents, then countries, then states, then streets. Each level narrows down the search area.

Metadata is additional structured data attached to each vector — like tags, category, date, author, or source URL. Vector databases can filter by metadata during similarity search, so you can say: “Find the 5 most similar documents from 2024 about machine learning.”

sequenceDiagram
participant App
participant VDB as Vector DB
App->>VDB: Search(embedding, filter={year: 2024, category: "ML"})
VDB->>VDB: 1. Filter: only vectors with year=2024 AND category="ML"
VDB->>VDB: 2. Search: find nearest neighbors in filtered set
VDB-->>App: Top 5 filtered + ranked results

Namespaces are a way to partition data within a collection — like folders. You can have separate namespaces for different customers, projects, or environments, all within the same collection and index.


flowchart TD
subgraph INGESTION["Ingestion Pipeline"]
DOC["Raw Document\n(PDF, wiki, code)"] --> CHUNK["Chunker\n(256-512 token pieces)"]
CHUNK --> EMBED["Embedding Model\n(text → vector)"]
EMBED --> STORE["Vector Database\n(store + index)"]
STORE --> META["Store Metadata\n(source, date, tags)"]
end
subgraph QUERY["Query Pipeline"]
Q["User Question"] --> Q_EMBED["Embedding Model\n(question → vector)"]
Q_EMBED --> Q_FILTER["Apply Filters\n(date, tags, access)"]
Q_FILTER --> Q_SEARCH["ANN Search\n(top-k neighbors)"]
Q_SEARCH --> Q_LLM["LLM\n(read chunks + answer)"]
Q_LLM --> Q_RESULT["Final Answer"]
end
INGESTION --> QUERY
style INGESTION fill:#3b82f6,color:#fff
style QUERY fill:#22c55e,color:#fff

How ChatGPT Answers Questions About Your PDFs

Section titled “How ChatGPT Answers Questions About Your PDFs”

When you upload a PDF to ChatGPT and ask a question:

  1. Ingestion: ChatGPT chunks the PDF (every ~500 words), embeds each chunk, and stores the vectors
  2. Query: Your question is embedded with the same model
  3. Search: The vector database finds the most relevant chunks
  4. Generate: ChatGPT reads those chunks + your question and generates an answer
flowchart TD
YOU["You upload a PDF\n(100 pages)"] --> CHUNK2["Chunked into\n200 pieces"]
CHUNK2 --> EMBED2["Each chunk → 1536d vector"]
EMBED2 --> VDB["Stored in temporary\nvector index"]
Q2["You ask:\n'What is the budget for Q3?'"] --> Q_EMBED2["Embed question"]
Q_EMBED2 --> SEARCH2["Search for\nsimilar chunks"]
SEARCH2 --> CHUNKS2["Top 3 chunks:\n'Q3 budget is $2.4M'\n'Revenue projection...'\n'Cost breakdown...'"]
CHUNKS2 --> ANSWER2["LLM reads chunks\n+ question → answer"]
style YOU fill:#3b82f6,color:#fff
style VDB fill:#f59e0b,color:#fff
style ANSWER2 fill:#22c55e,color:#fff

MistakeWhy It’s Wrong
❌ “I’ll use PostgreSQL for everything”PostgreSQL with pgvector is fine for small-scale, but dedicated vector databases have significantly better performance and features for production RAG at scale
❌ “I don’t need a vector database — I’ll just loop over all my data”Looping over 1M vectors in application code is ~2 seconds with optimized numpy. Production requires sub-100ms. Vector databases use ANN indexes to achieve this
❌ “All vector databases are the same”Different vector databases optimize for different things: Pinecone for ease of use, Qdrant for performance, Weaviate for hybrid search, Milvus for scale, Chroma for prototyping
❌ “Vector databases replace my existing database”Vector databases complement, not replace, traditional databases. You typically use both — a SQL/NoSQL database for application data and a vector database for semantic search

Q: What is a vector database and when would you use one?

A vector database is a database designed for storing and searching vectors using similarity. You use it when you need semantic search — finding content by meaning rather than exact keywords. Common use cases include RAG, recommendation systems, and semantic search.

Q: How is a vector database different from a traditional database like PostgreSQL?

Traditional databases use B-tree indexes for exact matches and range queries. Vector databases use ANN indexes (like HNSW) for similarity search. Vector databases are optimized for finding “nearest neighbors” in high-dimensional space, while traditional databases are optimized for structured queries, joins, and transactions.

Q: Your RAG system needs to search across 10 million documents with sub-100ms latency. Each document has access control tags. Design the architecture.

Solution: (1) Use HNSW index with appropriate M (16-32) and ef_construction (200-400) parameters for fast search. (2) Pre-filter using metadata — ensure the vector database supports efficient filtering during ANN search (some databases filter before search, others after — pre-filtering is faster). (3) Use partitioning/namespaces to separate data by access level. (4) Consider a two-tier approach: coarse filter first (by access level), then fine-grained semantic search. (5) Benchmark with your specific data — different vector databases perform differently depending on data distribution and query patterns.


ConceptKey Point
Vector DatabaseSpecialized for storing and searching vectors
Why It ExistsTraditional databases can’t do efficient similarity search
HNSW IndexGraph-based ANN index for fast approximate search
Metadata FilteringFilter by structured fields during similarity search
Popular OptionsPinecone, Qdrant, Weaviate, Milvus, Chroma, pgvector
Use CaseRAG, semantic search, recommendations, clustering

Previous: 04 — Similarity Search →

Next: Coming soon — Chunk 2: Building Your First RAG Pipeline →