14. Production AI Agent Architecture
Introduction
Section titled “Introduction”Production AI agents are not just LLMs with tool calls. They are distributed systems with planning, memory, retrieval, execution, observability, guardrails, and scaling — all working together to deliver reliable, safe, and fast agent experiences.
This document shows how real products like Cursor, Devin, OpenAI Operator, and Claude Desktop are architected. These systems cost millions to build and operate — understanding their architecture is essential for building your own production agents.
flowchart TD USER["👤 User"] --> GATEWAY["🚪 API Gateway"] GATEWAY --> AUTH["🔐 Auth & Rate Limiting"] AUTH --> ROUTER["🔀 Router\n(Intent Classification)"]
ROUTER --> PLAN["📋 Planner\n(Task Decomposition)"] PLAN --> QUEUE["📨 Task Queue"]
QUEUE --> MEM["💾 Memory Layer"] QUEUE --> RETRIEVER["🔍 Retriever\n(Knowledge Base)"] QUEUE --> LLM["🧠 LLM Service"] QUEUE --> TOOLS["🛠️ Tool Registry"]
TOOLS --> BROWSER["🌐 Browser"] TOOLS --> CODE["💻 Code Execution"] TOOLS --> FS["📁 File System"] TOOLS --> API["🔌 External APIs"] TOOLS --> DB["🗄️ Database"]
TOOLS --> EXEC["⚡ Execution Engine"] EXEC --> OBS["👁️ Observer"] OBS --> REFLECT["🪞 Reflection"] REFLECT -->|"Continue"| PLAN REFLECT -->|"Done"| RESPONSE["✅ Response"]
OBSERV["📊 Observability\n(Tracing, Logs, Metrics)"] -.-> ALL GUARD["🛡️ Guardrails\n(Safety, Validation)"] -.-> ALL
style USER fill:#3b82f6,color:#fff style GATEWAY fill:#8b5cf6,color:#fff style PLAN fill:#f59e0b,color:#fff style LLM fill:#22c55e,color:#fff style EXEC fill:#ef4444,color:#fff style RESPONSE fill:#22c55e,color:#fff style OBSERV fill:#6366f1,color:#fff style GUARD fill:#ec4899,color:#fffHow Real AI Agent Products Work
Section titled “How Real AI Agent Products Work”flowchart LR subgraph PRODUCTS["Real AI Agent Products"] CURSOR["Cursor\nCode Editor Agent"] DEVIN["Devin\nSoftware Engineer Agent"] OPERATOR["OpenAI Operator\nWeb Task Agent"] CLAUDE["Claude Desktop\nComputer Use Agent"] end
PRODUCTS --> COMMON["Common Architecture Patterns"]
style CURSOR fill:#3b82f6,color:#fff style DEVIN fill:#8b5cf6,color:#fff style OPERATOR fill:#f59e0b,color:#fff style CLAUDE fill:#22c55e,color:#fff style COMMON fill:#6366f1,color:#fffCursor Architecture
Section titled “Cursor Architecture”flowchart TD USER_CODE["👤 User writes code"] --> AI["🤖 AI Engine"] AI --> CONTEXT["📄 Context Builder\n(Current file, imports, errors)"] CONTEXT --> INDEX["📑 Code Index\n(AST + Embeddings)"] INDEX --> LLM_CURSOR["🧠 LLM\n(Code completion)"] LLM_CURSOR --> SUGGEST["💡 Suggestion"]
USER_CODE --> DIFF["🔍 Diff Detection"] DIFF --> ERRORS["❌ Error Detection"] ERRORS --> AI
AI --> TERMINAL["💻 Terminal Execution"] TERMINAL --> OUTPUT["📊 Output Analysis"] OUTPUT --> AI
style USER_CODE fill:#3b82f6,color:#fff style AI fill:#8b5cf6,color:#fff style LLM_CURSOR fill:#f59e0b,color:#fff style SUGGEST fill:#22c55e,color:#fffHow Cursor works:
- Context gathering — Reads current file, open tabs, imports, recent edits, compiler errors
- Code indexing — AST-based code understanding + embeddings for semantic search
- AI engine — LLM generates code suggestions based on full repository context
- Terminal integration — Executes commands, captures output, detects errors
- Iterative loop — Coder → Run → Check output → Fix errors → Repeat
Devin Architecture
Section titled “Devin Architecture”flowchart TD GOAL["🎯 User Goal:\n'Build a web app'"] --> PLANNER_D["📋 Planner\n(High-level plan)"] PLANNER_D --> SHELL["💻 Shell Agent\n(Commands & scripts)"] PLANNER_D --> EDITOR["📝 Editor Agent\n(Code writing)"] PLANNER_D --> BROWSER_D["🌐 Browser Agent\n(Testing & debugging)"] PLANNER_D --> SEARCH["🔍 Search Agent\n(Research & docs)"]
SHELL --> MEM_D["🧠 Memory\n(Project state)"] EDITOR --> MEM_D BROWSER_D --> MEM_D SEARCH --> MEM_D
MEM_D --> REVIEW_D["📝 Reviewer\n(Code review)"] REVIEW_D -->|"Issues"| EDITOR REVIEW_D -->|"Approved"| DEPLOY["🚀 Deploy"]
style GOAL fill:#3b82f6,color:#fff style PLANNER_D fill:#8b5cf6,color:#fff style MEM_D fill:#f59e0b,color:#fff style DEPLOY fill:#22c55e,color:#fffHow Devin works:
- Planner — Receives a high-level goal, creates a multi-step plan
- Specialized agents — Shell, Editor, Browser, Search agents each have specific tools
- Shared memory — All agents write to a shared project state (files created, commands run, errors)
- Reviewer — Continuously reviews code quality and progress
- Iterative — Loops between coding, testing, and fixing until the task is complete
OpenAI Operator & Claude Desktop
Section titled “OpenAI Operator & Claude Desktop”flowchart LR subgraph OPERATOR_ARCH["OpenAI Operator"] O_USER["👤 User"] --> O_LLM["🧠 LLM"] O_LLM --> O_BROWSER["🌐 Browser Control"] O_BROWSER --> O_SS["📸 Screenshot Analysis"] O_SS --> O_ACT["🖱️ Click / Type / Navigate"] O_ACT --> O_CHECK["✅ Check Result"] O_CHECK -->|"Failed"| O_RETRY["🔄 Retry with different approach"] O_CHECK -->|"Success"| O_DONE["✅ Done"] O_HITL["👤 Human Approval\n(for purchases, logins)"] -.-> O_ACT end
subgraph CLAUDE_ARCH["Claude Desktop"] C_USER["👤 User"] --> C_LLM["🧠 LLM"] C_LLM --> C_SCREEN["📸 Screen Capture"] C_SCREEN --> C_ANALYZE["👁️ Visual Analysis"] C_ANALYZE --> C_MOUSE["🖱️ Mouse Control"] C_ANALYZE --> C_KEYBOARD["⌨️ Keyboard Control"] C_MOUSE --> C_RESULT["📊 Check Result"] C_KEYBOARD --> C_RESULT C_RESULT -->|"Continue"| C_SCREEN C_RESULT -->|"Done"| C_DONE["✅ Task Complete"] end
style OPERATOR_ARCH fill:#3b82f6,color:#fff style CLAUDE_ARCH fill:#22c55e,color:#fffProduction Architecture Components
Section titled “Production Architecture Components”flowchart TD subgraph LAYERS["Production Agent Layers"] L1["🔹 Layer 1: Gateway\nAPI Gateway, Auth, Rate Limiting"] L2["🔹 Layer 2: Orchestration\nPlanner, Router, Task Queue"] L3["🔹 Layer 3: Intelligence\nLLM, Memory, Retrieval"] L4["🔹 Layer 4: Execution\nTool Registry, Execution Engine"] L5["🔹 Layer 5: Observability\nTracing, Logging, Monitoring"] L6["🔹 Layer 6: Safety\nGuardrails, Validation, Human-in-Loop"] end
L1 --> L2 --> L3 --> L4 L5 -.-> ALL L6 -.-> ALL
style L1 fill:#3b82f6,color:#fff style L2 fill:#8b5cf6,color:#fff style L3 fill:#f59e0b,color:#fff style L4 fill:#22c55e,color:#fff style L5 fill:#ef4444,color:#fff style L6 fill:#ec4899,color:#fffComponent Details
Section titled “Component Details”| Layer | Component | Purpose | Example |
|---|---|---|---|
| Gateway | API Gateway | Route requests, authenticate | Kong, Envoy |
| Gateway | Auth | User identity & permissions | JWT, OAuth, RBAC |
| Gateway | Rate Limiting | Prevent abuse | Token bucket, 100 req/min |
| Orchestration | Planner | Decompose goals into steps | LLM-based planner |
| Orchestration | Router | Route to appropriate agent | Intent classifier |
| Orchestration | Task Queue | Async task processing | RabbitMQ, Redis |
| Intelligence | LLM Service | Reasoning engine | OpenAI, Anthropic, self-hosted |
| Intelligence | Memory | Cross-session context | Vector DB + Redis |
| Intelligence | Retriever | Knowledge access | RAG pipeline |
| Execution | Tool Registry | Available tools | Plugin system |
| Execution | Execution Engine | Run tools safely | Sandboxed Docker |
| Observability | Tracing | Per-request flow | OpenTelemetry |
| Observability | Logging | Every action recorded | Structured logs |
| Safety | Guardrails | Input/output validation | Prompt injection detection |
| Safety | Human-in-Loop | Critical action approval | Approval queue |
Scaling AI Agents
Section titled “Scaling AI Agents”flowchart LR subgraph SCALING["Production Scaling Considerations"] S1["🎯 Horizontal Scaling\nStateless agents behind load balancer"] S2["💾 State Management\nRedis + PostgreSQL for persistence"] S3["⚡ Caching\nResponse cache + embedding cache"] S4["🔀 Queue-Based Processing\nAsync for long-running tasks"] S5["🌐 Multi-Model Routing\nCheap model for simple, expensive for complex"] end
style S1 fill:#3b82f6,color:#fff style S2 fill:#8b5cf6,color:#fff style S3 fill:#f59e0b,color:#fff style S4 fill:#22c55e,color:#fff style S5 fill:#ef4444,color:#fffScaling Strategy
Section titled “Scaling Strategy”| Strategy | How It Works | Impact |
|---|---|---|
| Stateless agents | Agent state in Redis, not in process | Easy horizontal scaling |
| Task queues | Long tasks queued, workers pick up | Handle 10K+ concurrent tasks |
| Caching | Cache LLM responses (TTL: 1 hour) | 40-60% cost reduction |
| Model routing | GPT-4o-mini for 80% of tasks, GPT-4o for 20% | 60% cost reduction |
| Parallel execution | Independent agent steps in parallel | 2-5x speedup |
Security Architecture
Section titled “Security Architecture”flowchart TD subgraph SEC["Security Layers"] S1["🔐 Authentication\nJWT / OAuth / SSO"] S2["📋 Authorization\nRBAC + Scope-based permissions"] S3["🛡️ Prompt Protection\nInput sanitization\nInjection detection\nPII masking"] S4["🔒 Execution Isolation\nSandboxed tool execution\nDocker containers\nNo network for code exec"] S5["📝 Audit Logging\nEvery action logged\nImmutable logs\nFull traceability"] end
S1 --> S2 --> S3 --> S4 --> S5
style S1 fill:#3b82f6,color:#fff style S2 fill:#8b5cf6,color:#fff style S3 fill:#f59e0b,color:#fff style S4 fill:#ef4444,color:#fff style S5 fill:#22c55e,color:#fffBest Practices
Section titled “Best Practices”- Design for failure from day one — Every component should have a fallback. LLM down → use cache. Tool fails → retry with backoff.
- Implement observability before launch — You need tracing, logging, and metrics before you have users. Debugging blind is impossible.
- Use the cheapest model that works — Route simple tasks to GPT-4o-mini, complex tasks to GPT-4o. Save 60% on API costs.
- Always have human-in-the-loop for destructive actions — Delete, deploy, send email, pay — require human approval.
- Set hard limits — Max iterations (25), max tokens (100K), max cost per task ($1), max execution time (5 min).
Common Mistakes
Section titled “Common Mistakes”| Mistake | Impact | Fix |
|---|---|---|
| No observability | Can’t debug agent failures | Add tracing from day one |
| No rate limiting | One user exhausts API budget | 100 req/min per user |
| No sandboxing | Code execution vulnerability | Run all code in Docker containers |
| Single model | Expensive for simple tasks | Route between cheap and expensive models |
| No caching | Same query hits LLM 1000 times | Cache LLM responses with TTL |
Interview Questions
Section titled “Interview Questions”Q: What are the main components of a production AI agent architecture?
Gateway (auth, rate limiting), Orchestration (planner, router), Intelligence (LLM, memory, retriever), Execution (tools, sandbox), Observability (tracing, logs), and Safety (guardrails, human-in-the-loop).
Q: How does Cursor differ from a simple LLM code generator?
Cursor has context gathering (reads open files, imports, errors), code indexing (AST + embeddings), terminal integration (runs code, checks output), and an iterative loop (generate → run → fix → repeat). A simple LLM just generates code once.
Intermediate
Section titled “Intermediate”Q: Explain the two-phase architecture used by Devin (planner + specialized agents).
Devin uses a planner agent that creates a high-level plan: “Build a login page” → “1. Create React component 2. Add form validation 3. Style with CSS 4. Test.” Then specialized agents execute each step: Shell Agent runs npm commands, Editor Agent writes code, Browser Agent tests the UI. All agents share a common memory of project state.
Senior
Section titled “Senior”Q: Design a caching strategy for a production agent that processes 10,000 queries per day.
Three-tier cache: (1) Exact match cache — Redis, TTL 24h. If the same query was asked before, return cached response. (30% hit rate). (2) Semantic cache — Vector DB. If a new query is > 95% similar to a cached one, return adapted cached response. (15% hit rate). (3) Embedding cache — Cache query embeddings. If the same query is asked, skip embedding API call. (60% hit rate). Invalidation: Clear cache when tools or knowledge base change.
Staff Engineer
Section titled “Staff Engineer”Q: Design a multi-model routing strategy for an agent that uses GPT-4o, Claude, and a self-hosted Llama model.
Router — Intent classifier determines task complexity. Tiers: (1) Simple (greetings, FAQs) → Llama 70B (self-hosted, $0). (2) Medium (research, writing) → GPT-4o-mini ($0.15/1M tokens). (3) Complex (coding, analysis) → GPT-4o ($2.50/1M tokens). (4) Critical (medical, legal) → Claude 3.5 Opus ($15/1M tokens). Failover: If GPT-4o is down → Claude. If Claude is down → GPT-4o. Cost savings: 70% of queries hit tiers 1-2, saving 80% vs using GPT-4o for everything.
System Design
Section titled “System Design”Q: Design an agent system that can handle 1M users with < 3 second response time for simple queries.
Architecture: Stateless API servers (auto-scale based on CPU), Redis for session state, PostgreSQL for persistence. Simple queries (< 3s): Direct LLM call with cached tools. Complex queries (> 3s): Async via task queue, user polls for result. Caching: Semantic cache for 80% of queries. Model routing: GPT-4o-mini for 95% of queries, GPT-4o for 5%. Scaling: 50 API servers handle 1M users. Cost: ~$50K/mo in API costs.
Architecture
Section titled “Architecture”Q: How would you architect an agent that can research, write, and publish blog posts autonomously?
Crew setup: Research Agent (searches 5 sources) → Writer Agent (drafts post) → Editor Agent (checks quality) → Publisher Agent (uploads to CMS). Guardrails: Fact-check gate before publishing. Human-in-loop: Final approval before publishing. Memory: Past posts stored for style consistency. Monitoring: Track publish success rate, reader engagement. Scaling: 5 parallel crews for 20 posts/day.
Summary
Section titled “Summary”| Component | Purpose | Production Consideration |
|---|---|---|
| API Gateway | Auth, routing, rate limiting | Scale horizontally, set limits |
| Planner | Goal decomposition | Re-plan every 3-5 steps |
| LLM Service | Reasoning engine | Multi-model routing for cost |
| Tool Registry | Available actions | Sandbox execution, validate output |
| Memory | Cross-session context | Vector DB + Redis |
| Observability | Debugging & monitoring | Trace every step from day one |
| Guardrails | Safety & validation | Block prompt injection, PII leaks |
| Human-in-Loop | Critical action approval | Delete, deploy, pay require approval |
Navigation
Section titled “Navigation”Previous: 13 — AutoGen & OpenAI Agents SDK