Skip to content

16. Production RAG Architecture

Production RAG is not a single script — it is a distributed system with microservices for ingestion, indexing, retrieval, re-ranking, prompt building, LLM inference, caching, monitoring, and security.

You have built a RAG pipeline that works on your laptop. Now imagine scaling that to millions of documents, thousands of queries per second, and enterprise security requirements. That is Production RAG.

flowchart TD
subgraph DEV["Development Environment"]
A["Single Python Script\n📝 Jupyter Notebook\n💻 Local Vector DB"]
end
subgraph PROD["Production Environment"]
B["📦 Ingestion Service\n📦 Chunking Service\n📦 Embedding Service\n📦 Vector DB Cluster\n📦 Retriever Service\n📦 Re-ranker Service\n📦 Prompt Builder\n📦 LLM Gateway\n📦 Cache Layer\n📦 Monitoring Stack"]
end
DEV -->|"10x Complexity\n100x Documents\n1000x Users"| PROD
style DEV fill:#3b82f6,color:#fff
style PROD fill:#22c55e,color:#fff

Your local RAG script works for 10 documents and 1 user. Production requires:

DimensionPrototypeProduction
Documents10–100Millions
Users1 developerThousands of concurrent users
Latency”It runs eventually”< 2 seconds
ReliabilityRun on demand99.9% uptime
SecurityNoneRBAC, encryption, audit logs
MonitoringPrint statementsDashboards, alerts, tracing
UpdatesRe-run manuallyContinuous indexing

Production RAG systems solve four core problems:

  1. Scale — Handle millions of documents and thousands of queries per second
  2. Reliability — No single point of failure, automatic retries, circuit breakers
  3. Security — Multi-tenant isolation, access control, compliance
  4. Observability — Know exactly why every answer was produced

Imagine turning a single librarian into a library factory with specialized workers:

  • Ingestion Workers — Receive new books, scan them, prepare them for storage
  • Chunking Specialists — Divide books into chapters and sections
  • Indexing Team — Create card catalog entries for every section
  • Retrieval Desk — Find the right sections when someone asks a question
  • Quality Reviewers — Re-rank results to ensure the best answers come first
  • Prompt Architects — Package retrieved sections with the question
  • LLM Experts — Generate the final answer
  • Cache Clerks — Store frequent answers for instant response
  • Monitoring Team — Track every request, measure quality, alert on issues

Each team works independently, scales independently, and can be improved independently.


flowchart TD
DOCS["📄 Source Documents\nPDF, Word, Websites, DB"] --> INGEST["📦 Ingestion Service\nExtract, Clean, Normalize"]
INGEST --> QUEUE["📬 Message Queue\n(RabbitMQ / Kafka / SQS)"]
QUEUE --> CHUNK["✂️ Chunking Service\nRecursive / Semantic / Sentence"]
CHUNK --> EMBED["🧠 Embedding Service\nBatch Embedding Generation"]
EMBED --> VDB[("🗄️ Vector Database\nPinecone / Qdrant / Weaviate")]
EMBED --> KV[("⚡ Key-Value Store\nRedis / DynamoDB\nMetadata & Cache")]
subgraph RETRIEVAL["Retrieval Pipeline"]
QUERY["🔍 User Query"] --> AUTH["🔐 Auth & Filtering\nRBAC, Tenant Isolation"]
AUTH --> RET["📡 Retriever\nVector + Keyword + Hybrid"]
RET --> RERANK["📊 Re-ranker\nCross-encoder scoring"]
RERANK --> PROMPT["📝 Prompt Builder\nContext + Instruction"]
PROMPT --> LLM["🤖 LLM Gateway\nGPT-4 / Claude / Gemini"]
LLM --> CACHE["💾 Response Cache\nRedis"]
CACHE --> RESP["✅ Final Response"]
end
RET --> VDB
RET --> KV
subgraph OPS["Operations"]
MON["📊 Monitoring\nLangSmith / Phoenix"]
LOG["📋 Logging\nOpenTelemetry"]
METRIC["📈 Metrics\nLatency, Cost, Quality"]
ALERT["🔔 Alerts\nPagerDuty / Slack"]
end
style RETRIEVAL fill:#1a1a2e,color:#fff
style OPS fill:#2d2d44,color:#fff
style DOCS fill:#3b82f6,color:#fff
style RESP fill:#22c55e,color:#fff

The ingestion service handles everything that happens before a query arrives.

ComponentResponsibilityTechnology Examples
Document ParserExtract text from PDF, Word, HTML, MarkdownUnstructured.io, LangChain, PyMuPDF
Text CleanerRemove headers, footers, noiseCustom rules, regex
NormalizerStandardize format, fix encodingutf-8, Unicode NFKC
Metadata ExtractorExtract author, date, source, tagsCustom extractors
RouterSend to appropriate processing queueKafka topics, SQS queues

A dedicated service for splitting documents into searchable pieces.

  • Input: Clean raw text + metadata
  • Process: Apply chunking strategy (recursive, semantic, sentence)
  • Output: Structured chunks with parent IDs, metadata, and position info

Generates vector embeddings for every chunk.

  • Batch Processing: Embeds chunks in batches for efficiency
  • Model Selection: Configurable embedding model (text-embedding-3-small, voyage-2, BGE)
  • Dimension Handling: Supports 384, 768, 1024, 1536 dimensions
  • Caching: Reuses embeddings for identical chunks

The storage and retrieval backbone.

  • Indexing: HNSW, IVF, or DiskANN indexes for fast approximate nearest neighbor search
  • Metadata Storage: Stores document IDs, timestamps, permissions alongside vectors
  • Filtering: Pre-filtering by metadata before vector search
  • Replication: Multi-node replication for high availability
  • Backup: Snapshot-based backups for disaster recovery

Orchestrates the retrieval strategy.

  • Multi-Strategy: Supports vector, keyword, and hybrid search
  • Fusion: Combines results from multiple strategies (RRF, weighted scoring)
  • Top-K Config: Dynamically adjustable number of results
  • Fallback: Falls back to keyword search if vector search fails

Improves retrieval quality with cross-encoder scoring.

  • Slower but Better: Cross-encoders are slower but more accurate than bi-encoders
  • Top-100 → Top-5: Re-ranks the top 100 results to find the best 5
  • Configurable: Can be skipped for latency-sensitive applications

Constructs the final prompt sent to the LLM.

  • Template System: Versioned prompt templates
  • Context Packing: Inserts retrieved chunks into context window
  • Instruction Injection: Adds system instructions and formatting rules
  • Token Budgeting: Ensures context fits within model limits

Manages LLM inference with production reliability.

  • Load Balancing: Distributes requests across multiple models/providers
  • Fallback: Switches to backup model on failure
  • Rate Limiting: Prevents API quota exhaustion
  • Cost Tracking: Logs token usage per request

In a production environment, these services are deployed as a distributed system with redundancy, auto-scaling, and observability.

flowchart TD
subgraph NETWORK["Network Layer"]
LB["⚖️ Load Balancer
(ALB / Nginx)"]
CDN["🌐 CDN
(CloudFront / Cloudflare)"]
end
subgraph APP["Application Layer
Kubernetes Cluster"]
API["📡 API Servers
Auto-scaled (2-20 pods)"]
INGEST["📦 Ingestion Workers
Auto-scaled by queue depth"]
RETRIEVE["📡 Retriever Pods
Auto-scaled by CPU"]
RERANK["📊 Re-ranker Pods
GPU-accelerated"]
end
subgraph CACHE["Cache Layer"]
REDIS1["💾 Redis Cluster
Embedding Cache
Response Cache
Session Store"]
end
subgraph STORAGE["Data Layer"]
VDB_CLUSTER[("🗄️ Vector DB Cluster
Primary + 2 Read Replicas")]
S3_BUCKET["☁️ S3 / Blob Storage
Raw documents, backups"]
POSTGRES["🐘 PostgreSQL
Metadata, audit logs,
user config"]
end
subgraph OBSERV["Observability"]
PROM["📊 Prometheus
Metrics collection"]
GRAF["📈 Grafana
Dashboards + Alerts"]
TRACE["🔍 OpenTelemetry
Distributed tracing"]
LOGS["📋 ELK / Loki
Log aggregation"]
end
CDN --> LB
LB --> API
API --> INGEST
API --> RETRIEVE
RETRIEVE --> RERANK
API --> REDIS1
RETRIEVE --> VDB_CLUSTER
INGEST --> S3_BUCKET
API --> POSTGRES
API --> PROM
API --> TRACE
API --> LOGS
PROM --> GRAF
style NETWORK fill:#3b82f6,color:#fff
style APP fill:#22c55e,color:#fff
style CACHE fill:#f59e0b,color:#fff
style STORAGE fill:#8b5cf6,color:#fff
style OBSERV fill:#ef4444,color:#fff

sequenceDiagram
participant User as 👤 User
participant API as API Gateway
participant Ingest as Ingestion Service
participant Queue as Message Queue
participant Chunk as Chunking Service
participant Embed as Embedding Service
participant VDB as Vector Database
User->>API: Upload Document (PDF)
API->>Ingest: Send to ingestion
Ingest->>Ingest: Extract text & metadata
Ingest->>Queue: Publish chunking job
Queue->>Chunk: Consume chunking job
Chunk->>Chunk: Split into chunks
Chunk->>Queue: Publish embedding job
Queue->>Embed: Consume embedding job
Embed->>Embed: Generate embeddings
Embed->>VDB: Store vectors + metadata
VDB-->>Embed: Confirm storage
Embed-->>Queue: Ack complete
Queue-->>Ingest: Pipeline complete
Ingest-->>API: Document indexed
API-->>User: ✅ Upload complete

sequenceDiagram
participant User as 👤 User
participant Gateway as API Gateway
participant Auth as Auth Service
participant Ret as Retriever
participant VDB as Vector Database
participant Rerank as Re-ranker
participant Prompt as Prompt Builder
participant LLM as LLM Gateway
participant Cache as Cache Layer
User->>Gateway: Query
Gateway->>Auth: Authenticate & authorize
Auth-->>Gateway: User context + permissions
Gateway->>Cache: Check cache
alt Cache Hit
Cache-->>Gateway: Cached response
Gateway-->>User: ✅ Response (cached)
else Cache Miss
Gateway->>Ret: Retrieve with filters
Ret->>VDB: Vector search + metadata filter
VDB-->>Ret: Top 100 results
Ret->>Rerank: Re-rank results
Rerank-->>Ret: Top 5 results
Ret->>Prompt: Build prompt
Prompt->>LLM: Generate answer
LLM-->>Prompt: Generated response
Prompt->>Cache: Store response
Prompt-->>Gateway: Final answer
Gateway-->>User: ✅ Response
end

ComponentImplementation
IngestionAutomated crawler that indexes company wikis, SharePoint, Google Drive
ChunkingSemantic chunking with overlap, optimized for different document types
EmbeddingOpenAI’s text-embedding-3-large (3072 dimensions)
Vector DBCustom vector storage with metadata filtering for tenant isolation
RetrievalHybrid search (BM25 + vector) with re-ranking
SecurityEnterprise SSO, RBAC, document-level permissions
MonitoringUsage dashboards, audit logs, cost tracking per workspace
ComponentImplementation
IngestionNative M365 connectors — SharePoint, OneDrive, Teams, Exchange
ChunkingDocument structure-aware chunking (respects headings, tables)
EmbeddingMicrosoft’s internal embedding models
Vector DBAzure AI Search with integrated vector + keyword search
SecurityInherits existing M365 permissions — users only see their authorized documents
MonitoringMicrosoft 365 admin center, compliance reports
ComponentImplementation
IngestionCloud Storage, BigQuery, web crawling connectors
ChunkingConfigurable chunk size and strategy
EmbeddingGoogle’s embedding models via Vertex AI
Vector DBVertex AI Vector Search (formerly Matching Engine) — extremely scalable
RetrievalHybrid search with re-ranking, natural language understanding
SecurityIAM roles, VPC-SC, CMEK encryption

MistakeWhy It’s WrongFix
Monolithic architectureOne service does everything → hard to scale, deploy, debugSplit into microservices with clear responsibilities
Synchronous ingestionUser waits for document to be indexedUse async queues, notify user when ready
No retry logicTransient failures cause pipeline failuresImplement exponential backoff + dead letter queues
Ignoring metadataNo permission filtering → security leaksStore metadata with every vector, filter before retrieval
Single LLM providerProvider outage = system downMulti-provider fallback strategy
No cachingEvery query recomputes embeddings and LLM responsesCache at multiple layers (embedding, response, retriever)

  1. Design for failure — Every service should assume downstream services will fail. Use circuit breakers, retries, and fallbacks.

  2. Version everything — Prompt templates, embedding models, chunking strategies, and retrieval configurations should all be versioned.

  3. Measure everything — Latency per service, cost per query, retrieval recall, answer faithfulness. If you can’t measure it, you can’t improve it.

  4. Separate ingestion from retrieval — Ingestion can be slower and batch-oriented. Retrieval must be fast and real-time.

  5. Use message queues — Decouple ingestion stages with queues (Kafka, RabbitMQ, SQS) for reliability and scalability.

  6. Implement graceful degradation — If the re-ranker fails, fall back to raw retriever results. If the LLM fails, return retrieved documents directly.

  7. Plan for index rebuilds — When embedding models change, you need to re-embed all documents. Build this into your architecture from day one.


Q: What is the difference between a prototype RAG system and a production RAG system?

A prototype runs on a single machine with a few documents. A production system is distributed across multiple services, handles millions of documents, supports concurrent users, implements security and access control, includes monitoring and alerting, and is designed for reliability with retries, fallbacks, and disaster recovery.

Q: Why does production RAG use message queues?

Message queues decouple the ingestion pipeline stages. Documents can be uploaded immediately, but chunking and embedding happen asynchronously. This prevents slow ingestion from blocking users and allows each stage to scale independently.

Q: How would you design a multi-tenant RAG system where Company A and Company B both use the same infrastructure but should never see each other’s documents?

Each document is tagged with a tenant_id metadata field. The vector database stores this as a filterable attribute. Every query includes the authenticated tenant’s ID as a pre-filter. This ensures vector search only returns the tenant’s documents. Additional isolation can be achieved with separate index partitions or separate database instances for compliance requirements.

Q: What strategies would you use to reduce P95 latency in a production RAG system?

  1. Caching — Cache embedding results for frequent queries, cache LLM responses for identical questions
  2. Pre-computation — Pre-compute embeddings during ingestion, never at query time
  3. Load balancing — Distribute retrieval across replicas
  4. Connection pooling — Reuse database and LLM connections
  5. Streaming — Stream LLM responses to users instead of waiting for the full response
  6. Reduce re-ranking — Only re-rank when quality is critical; skip for simple queries

Q: How would you handle embedding model version migration in production without downtime?

  1. Dual-writing — When a new embedding model is deployed, write vectors for both old and new models simultaneously
  2. Gradual migration — Run both embedding models in parallel, serve queries using the old model while the new one indexes
  3. Index swap — Build the new index in the background, then swap query routing atomically
  4. A/B testing — Route a percentage of queries to the new model to validate quality before full cutover
  5. Rollback plan — Keep the old index accessible for at least one week in case quality regression is detected

Q: Design a production RAG system that handles 10,000 queries per second across 1,000 tenants with document-level access control.

Architecture:

Ingestion: Documents flow through a Kafka pipeline: Parser → Chunker → Embedder → Vector DB. Each document is tagged with tenant_id and document_id metadata.

Storage: Vector database partitioned by tenant (tenant_id as partition key). Each partition has its own HNSW index. Metadata stored alongside vectors for document-level filtering.

Caching: Two-layer cache — (1) Embedding cache: LRU cache mapping query text to embedding vector, (2) Response cache: Redis cluster mapping (query_hash + tenant_id) to response.

Retrieval: Each query goes through: Auth service (extract tenant + permissions) → Cache check → Retriever (hybrid search with tenant pre-filter) → Re-ranker (cross-encoder for top 20) → LLM (with load balancing across GPT-4, Claude, and a local fallback model).

Scaling: Horizontal scaling for all stateless services. Vector DB uses read replicas for retrieval, single writer for ingestion. Auto-scaling based on queue depth and query latency.

Reliability: Circuit breakers between every service. Dead letter queues for failed ingestion. Multi-region deployment with active-passive failover.

Q: Draw the architecture of an enterprise RAG system and explain how you would ensure it meets SOC 2 compliance.

Architecture: (Refer to the Production RAG Architecture diagram above)

SOC 2 Compliance:

  • Security: Encryption at rest (AES-256) and in transit (TLS 1.3). Access control via IAM roles. Secrets stored in HashiCorp Vault.
  • Availability: Multi-AZ deployment, auto-scaling, health checks, circuit breakers. SLA target: 99.9% uptime.
  • Processing Integrity: Prompt versioning, embedding model versioning, audit trail of all changes. Data validation at every pipeline stage.
  • Confidentiality: Document-level access control. PII masking in retrieved chunks. Tenant isolation via metadata filtering.
  • Privacy: Data retention policies. Automatic deletion of documents after specified TTL. User data export/deletion APIs for GDPR compliance.

ConceptKey Point
Production RAGDistributed system with specialized microservices, not a single script
Ingestion PipelineAsync, queue-based document processing
Retrieval PipelineMulti-stage: filter → search → re-rank → prompt → LLM → cache
CachingEmbedding cache + response cache for latency reduction
SecurityTenant isolation, RBAC, encryption, audit logs
ReliabilityRetries, circuit breakers, fallbacks, multi-region
ObservabilityTracing, metrics, logging, dashboards, alerts

Previous: 15 — Multi-Query Retrieval

Next: 17 — Metadata Filtering & Security

Related Topics: