Welcome to the end of Phase 5. You’ve built ChatPDF, a Company Knowledge Assistant, a Documentation Chatbot, and a GitHub Code Assistant. Now it’s time to review everything, compare approaches, and chart your path forward.
This summary page consolidates everything you’ve learned across 24 documents and 5 chunks. Use it as a reference, a decision guide, and a roadmap to further mastery.
subgraph DOCUMENTS["Document Sources"]
subgraph INGESTION["Ingestion Pipeline"]
EMBED["Generate Embeddings"]
STORE["Store in Vector DB"]
subgraph RETRIEVAL["Retrieval Strategies"]
SIMPLE["Basic Vector Search"]
HYBRID["Hybrid Search\n(BM25 + Vector)"]
MQ["Multi-Query\n(Query Expansion)"]
PC["Parent-Child\n(Chunk Expansion)"]
RERANK["Re-ranking\n(Cross-Encoder)"]
subgraph GEN["Generation"]
COMPRESS["Context Compression"]
CITATION["Citation Builder"]
subgraph PROD["Production"]
MONITOR["Monitoring & Evaluation"]
SCALE["Scaling & Load Balancing"]
DOCUMENTS --> INGESTION --> RETRIEVAL --> GEN --> PROD
style DOCUMENTS fill:#3b82f6,color:#fff
style INGESTION fill:#8b5cf6,color:#fff
style RETRIEVAL fill:#f59e0b,color:#fff
style GEN fill:#22c55e,color:#fff
style PROD fill:#ef4444,color:#fff
subgraph EVOLUTION["RAG Evolution"]
BASIC["Basic RAG\nEmbed → Search → Generate"]
ADV["Advanced RAG\nHybrid + Re-rank + Compression"]
ENTERPRISE["Enterprise RAG\nMulti-tenant + Security + Eval"]
AGENTIC["Agentic RAG\n(Covered in Phase 6)\nSelf-query + Tool use + Planning"]
BASIC --> ADV --> ENTERPRISE --> AGENTIC
style BASIC fill:#3b82f6,color:#fff
style ADV fill:#8b5cf6,color:#fff
style ENTERPRISE fill:#22c55e,color:#fff
style AGENTIC fill:#ef4444,color:#fff
Dimension Basic RAG Advanced RAG Enterprise RAG Search Vector only Hybrid (vector + keyword) Hybrid + multi-query + parent-child Ranking Vector similarity + Cross-encoder re-ranking + Score normalization + thresholding Context Raw chunks + Compression + deduplication + Summarization + citation building Security None None RBAC, metadata filtering, audit logs Multi-tenancy None None Tenant isolation, per-tenant indexes Evaluation Manual LLM-as-judge Automated eval pipeline + dashboards Caching None Embedding cache Embedding + response + semantic cache Monitoring None Basic metrics Full observability (traces, logs, alerts) Latency 2-5s 2-5s < 2s with caching Cost $ $$ $$$ Best for Prototypes, MVPs Production apps Enterprise systems
Q["Do you need\nkeyword matching?"]
Q -->|"Yes"| HYBRID["Hybrid Search\n(BM25 + Vector)"]
Q -->|"No"| VEC["Vector Search Only"]
HYBRID --> Q2["Is accuracy\ncritical?"]
Q2 -->|"Yes"| RERANK["Add Re-ranking\n(Cross-Encoder)"]
Q2 -->|"No"| SIMPLE["Basic Top-K"]
RERANK --> Q3["User queries\nvague?"]
Q3 -->|"Yes"| MQ["Add Multi-Query\n(Query Expansion)"]
Q3 -->|"No"| OK["Good enough"]
MQ --> Q4["Need precise\n+ context?"]
Q4 -->|"Yes"| PC["Add Parent-Child\nRetrieval"]
Q4 -->|"No"| DONE["✅ Done"]
style Q fill:#f59e0b,color:#fff
style Q2 fill:#f59e0b,color:#fff
style Q3 fill:#f59e0b,color:#fff
style Q4 fill:#f59e0b,color:#fff
style DONE fill:#22c55e,color:#fff
Scenario Better Approach You need the model to learn new facts permanently Fine-tuning You have < 20 documents and the LLM context window fits them all Just put them in the prompt You need the model to follow a specific output format Fine-tuning or structured output Your documents change every minute Streaming data pipeline rather than batch RAG You need 100% factual accuracy with no hallucination risk Constrained extraction (no generation) Your users need real-time data from APIs Function calling / tool use (Phase 6)
Feature Keyword Search Semantic Search Hybrid Search Method BM25 / TF-IDF Vector similarity BM25 + Vector fusion Match Exact words Meaning Both Handles typos No Yes Yes Handles synonyms No Yes Yes Handles code Good (exact API names) Poor Good Latency < 50ms < 100ms < 200ms Best for Code, exact terms General text, concepts Production systems
Aspect Development Production Vector DB Local (Chroma, FAISS) Managed (Qdrant, Pinecone) LLM GPT-4o-mini GPT-4o / Claude 3 Caching None Embedding + Response + Semantic Monitoring Console logs Grafana + Prometheus + Alerts Security None Auth + RBAC + Encryption + Audit Scaling Single process Horizontal, auto-scaling Uptime 99% 99.9%+ Cost < $100/mo $500-$5000/mo
Strategy Best For Chunk Size Overlap Fixed size Simple documents 256-512 tokens 10-20% Recursive General purpose 512 tokens 10-20% Semantic Well-structured text Variable None needed Section-based Documentation Per section None needed Function-level Code Per function Imports as context Parent-Child Precision + context Small child (128), large parent (1024) Hierarchical
subgraph LEVEL1["Level 1: Foundation ✓"]
L1_1["✅ Embeddings & Vector Search"]
L1_2["✅ Vector Databases"]
L1_3["✅ Basic RAG Pipeline"]
subgraph LEVEL2["Level 2: Building ✓"]
L2_1["✅ Chunking Strategies"]
L2_2["✅ Document Ingestion"]
subgraph LEVEL3["Level 3: Advanced ✓"]
L3_3["✅ Multi-Query & Parent-Child"]
subgraph LEVEL4["Level 4: Production ✓"]
L4_1["✅ Production Architecture"]
L4_2["✅ Security & Multi-tenancy"]
L4_3["✅ Evaluation & Monitoring"]
L4_4["✅ Scaling & Caching"]
subgraph LEVEL5["Level 5: Projects ✓"]
L5_2["✅ Knowledge Assistant"]
L5_3["✅ Documentation Bot"]
subgraph NEXT["Next: Phase 6"]
N4["👥 Multi-Agent Systems"]
LEVEL1 --> LEVEL2 --> LEVEL3 --> LEVEL4 --> LEVEL5 --> NEXT
style LEVEL1 fill:#3b82f6,color:#fff
style LEVEL2 fill:#8b5cf6,color:#fff
style LEVEL3 fill:#f59e0b,color:#fff
style LEVEL4 fill:#22c55e,color:#fff
style LEVEL5 fill:#ef4444,color:#fff
style NEXT fill:#6366f1,color:#fff
Project Key Lesson Architecture Pattern ChatPDF Document ingestion pipeline + Query pipeline Two-pipeline architecture Knowledge Assistant Access control via metadata filtering Pre-filter before search Documentation Bot Section-based chunking + version management Citation system + version tags Code Assistant AST-based chunking + dependency graph Symbol index + multi-search
subgraph COST["Relative Cost to Run"]
P1["📄 ChatPDF\n$200-800/mo"]
P2["🏢 Knowledge Assistant\n$1,000-5,000/mo"]
P3["📚 Documentation Bot\n$100-500/mo"]
P4["💻 Code Assistant\n$500-2,000/mo"]
subgraph COMPLEXITY["Implementation Complexity"]
C2["⭐⭐ Knowledge Assistant\nHigh"]
C3["⭐ Documentation Bot\nLow-Medium"]
C4["⭐⭐⭐ Code Assistant\nVery High"]
subgraph TIME["Time to MVP"]
T1["📄 ChatPDF\n1-2 weeks"]
T2["🏢 Knowledge Assistant\n4-8 weeks"]
T3["📚 Documentation Bot\n1-2 weeks"]
T4["💻 Code Assistant\n6-12 weeks"]
style P1 fill:#3b82f6,color:#fff
style P2 fill:#8b5cf6,color:#fff
style P3 fill:#22c55e,color:#fff
style P4 fill:#ef4444,color:#fff
style C1 fill:#3b82f6,color:#fff
style C2 fill:#8b5cf6,color:#fff
style C3 fill:#22c55e,color:#fff
style C4 fill:#ef4444,color:#fff
style T1 fill:#3b82f6,color:#fff
style T2 fill:#8b5cf6,color:#fff
style T3 fill:#22c55e,color:#fff
style T4 fill:#ef4444,color:#fff
You’ve mastered retrieval — giving LLMs access to knowledge. Now learn how to give LLMs access to tools, actions, and autonomy :
Agent Architecture — How agents think: perceive → reason → act → observe
Tool Use — How agents call APIs, run code, and interact with the world
Agent Memory — Short-term, long-term, and episodic memory for agents
Planning — How agents break down complex tasks into steps
Multi-Agent Systems — How agents collaborate, delegate, and debate
Agent Safety — Guardrails, human-in-the-loop, and failure handling
Chunk Topics Documents Chunk 1: Embeddings & Vector Search Why retrieval, embeddings, vector space, similarity search, vector DBs 01-05 Chunk 2: Building Your First RAG Pipeline What is RAG, chunking, ingestion pipeline, retrievers, complete pipeline 06-10 Chunk 3: Advanced Retrieval Techniques Hybrid search, re-ranking, context compression, parent-child, multi-query 11-15 Chunk 4: Production RAG Systems Architecture, metadata filtering, evaluation, scaling, best practices 16-20 Chunk 5: Enterprise RAG Projects ChatPDF, Knowledge Assistant, Documentation Bot, Code Assistant, Summary 21-25
Previous: 24 — GitHub Code Assistant
Next: Phase 6 — AI Agents & Agentic Systems (Coming Soon)
🎉 Congratulations on completing Phase 5: Retrieval Systems & RAG!