Skip to content

14. Production AI Agent Architecture

Production AI agents are not just LLMs with tool calls. They are distributed systems with planning, memory, retrieval, execution, observability, guardrails, and scaling — all working together to deliver reliable, safe, and fast agent experiences.

This document shows how real products like Cursor, Devin, OpenAI Operator, and Claude Desktop are architected. These systems cost millions to build and operate — understanding their architecture is essential for building your own production agents.

flowchart TD
USER["👤 User"] --> GATEWAY["🚪 API Gateway"]
GATEWAY --> AUTH["🔐 Auth & Rate Limiting"]
AUTH --> ROUTER["🔀 Router\n(Intent Classification)"]
ROUTER --> PLAN["📋 Planner\n(Task Decomposition)"]
PLAN --> QUEUE["📨 Task Queue"]
QUEUE --> MEM["💾 Memory Layer"]
QUEUE --> RETRIEVER["🔍 Retriever\n(Knowledge Base)"]
QUEUE --> LLM["🧠 LLM Service"]
QUEUE --> TOOLS["🛠️ Tool Registry"]
TOOLS --> BROWSER["🌐 Browser"]
TOOLS --> CODE["💻 Code Execution"]
TOOLS --> FS["📁 File System"]
TOOLS --> API["🔌 External APIs"]
TOOLS --> DB["🗄️ Database"]
TOOLS --> EXEC["⚡ Execution Engine"]
EXEC --> OBS["👁️ Observer"]
OBS --> REFLECT["🪞 Reflection"]
REFLECT -->|"Continue"| PLAN
REFLECT -->|"Done"| RESPONSE["✅ Response"]
OBSERV["📊 Observability\n(Tracing, Logs, Metrics)"] -.-> ALL
GUARD["🛡️ Guardrails\n(Safety, Validation)"] -.-> ALL
style USER fill:#3b82f6,color:#fff
style GATEWAY fill:#8b5cf6,color:#fff
style PLAN fill:#f59e0b,color:#fff
style LLM fill:#22c55e,color:#fff
style EXEC fill:#ef4444,color:#fff
style RESPONSE fill:#22c55e,color:#fff
style OBSERV fill:#6366f1,color:#fff
style GUARD fill:#ec4899,color:#fff

flowchart LR
subgraph PRODUCTS["Real AI Agent Products"]
CURSOR["Cursor\nCode Editor Agent"]
DEVIN["Devin\nSoftware Engineer Agent"]
OPERATOR["OpenAI Operator\nWeb Task Agent"]
CLAUDE["Claude Desktop\nComputer Use Agent"]
end
PRODUCTS --> COMMON["Common Architecture Patterns"]
style CURSOR fill:#3b82f6,color:#fff
style DEVIN fill:#8b5cf6,color:#fff
style OPERATOR fill:#f59e0b,color:#fff
style CLAUDE fill:#22c55e,color:#fff
style COMMON fill:#6366f1,color:#fff
flowchart TD
USER_CODE["👤 User writes code"] --> AI["🤖 AI Engine"]
AI --> CONTEXT["📄 Context Builder\n(Current file, imports, errors)"]
CONTEXT --> INDEX["📑 Code Index\n(AST + Embeddings)"]
INDEX --> LLM_CURSOR["🧠 LLM\n(Code completion)"]
LLM_CURSOR --> SUGGEST["💡 Suggestion"]
USER_CODE --> DIFF["🔍 Diff Detection"]
DIFF --> ERRORS["❌ Error Detection"]
ERRORS --> AI
AI --> TERMINAL["💻 Terminal Execution"]
TERMINAL --> OUTPUT["📊 Output Analysis"]
OUTPUT --> AI
style USER_CODE fill:#3b82f6,color:#fff
style AI fill:#8b5cf6,color:#fff
style LLM_CURSOR fill:#f59e0b,color:#fff
style SUGGEST fill:#22c55e,color:#fff

How Cursor works:

  1. Context gathering — Reads current file, open tabs, imports, recent edits, compiler errors
  2. Code indexing — AST-based code understanding + embeddings for semantic search
  3. AI engine — LLM generates code suggestions based on full repository context
  4. Terminal integration — Executes commands, captures output, detects errors
  5. Iterative loop — Coder → Run → Check output → Fix errors → Repeat
flowchart TD
GOAL["🎯 User Goal:\n'Build a web app'"] --> PLANNER_D["📋 Planner\n(High-level plan)"]
PLANNER_D --> SHELL["💻 Shell Agent\n(Commands & scripts)"]
PLANNER_D --> EDITOR["📝 Editor Agent\n(Code writing)"]
PLANNER_D --> BROWSER_D["🌐 Browser Agent\n(Testing & debugging)"]
PLANNER_D --> SEARCH["🔍 Search Agent\n(Research & docs)"]
SHELL --> MEM_D["🧠 Memory\n(Project state)"]
EDITOR --> MEM_D
BROWSER_D --> MEM_D
SEARCH --> MEM_D
MEM_D --> REVIEW_D["📝 Reviewer\n(Code review)"]
REVIEW_D -->|"Issues"| EDITOR
REVIEW_D -->|"Approved"| DEPLOY["🚀 Deploy"]
style GOAL fill:#3b82f6,color:#fff
style PLANNER_D fill:#8b5cf6,color:#fff
style MEM_D fill:#f59e0b,color:#fff
style DEPLOY fill:#22c55e,color:#fff

How Devin works:

  1. Planner — Receives a high-level goal, creates a multi-step plan
  2. Specialized agents — Shell, Editor, Browser, Search agents each have specific tools
  3. Shared memory — All agents write to a shared project state (files created, commands run, errors)
  4. Reviewer — Continuously reviews code quality and progress
  5. Iterative — Loops between coding, testing, and fixing until the task is complete
flowchart LR
subgraph OPERATOR_ARCH["OpenAI Operator"]
O_USER["👤 User"] --> O_LLM["🧠 LLM"]
O_LLM --> O_BROWSER["🌐 Browser Control"]
O_BROWSER --> O_SS["📸 Screenshot Analysis"]
O_SS --> O_ACT["🖱️ Click / Type / Navigate"]
O_ACT --> O_CHECK["✅ Check Result"]
O_CHECK -->|"Failed"| O_RETRY["🔄 Retry with different approach"]
O_CHECK -->|"Success"| O_DONE["✅ Done"]
O_HITL["👤 Human Approval\n(for purchases, logins)"] -.-> O_ACT
end
subgraph CLAUDE_ARCH["Claude Desktop"]
C_USER["👤 User"] --> C_LLM["🧠 LLM"]
C_LLM --> C_SCREEN["📸 Screen Capture"]
C_SCREEN --> C_ANALYZE["👁️ Visual Analysis"]
C_ANALYZE --> C_MOUSE["🖱️ Mouse Control"]
C_ANALYZE --> C_KEYBOARD["⌨️ Keyboard Control"]
C_MOUSE --> C_RESULT["📊 Check Result"]
C_KEYBOARD --> C_RESULT
C_RESULT -->|"Continue"| C_SCREEN
C_RESULT -->|"Done"| C_DONE["✅ Task Complete"]
end
style OPERATOR_ARCH fill:#3b82f6,color:#fff
style CLAUDE_ARCH fill:#22c55e,color:#fff

flowchart TD
subgraph LAYERS["Production Agent Layers"]
L1["🔹 Layer 1: Gateway\nAPI Gateway, Auth, Rate Limiting"]
L2["🔹 Layer 2: Orchestration\nPlanner, Router, Task Queue"]
L3["🔹 Layer 3: Intelligence\nLLM, Memory, Retrieval"]
L4["🔹 Layer 4: Execution\nTool Registry, Execution Engine"]
L5["🔹 Layer 5: Observability\nTracing, Logging, Monitoring"]
L6["🔹 Layer 6: Safety\nGuardrails, Validation, Human-in-Loop"]
end
L1 --> L2 --> L3 --> L4
L5 -.-> ALL
L6 -.-> ALL
style L1 fill:#3b82f6,color:#fff
style L2 fill:#8b5cf6,color:#fff
style L3 fill:#f59e0b,color:#fff
style L4 fill:#22c55e,color:#fff
style L5 fill:#ef4444,color:#fff
style L6 fill:#ec4899,color:#fff
LayerComponentPurposeExample
GatewayAPI GatewayRoute requests, authenticateKong, Envoy
GatewayAuthUser identity & permissionsJWT, OAuth, RBAC
GatewayRate LimitingPrevent abuseToken bucket, 100 req/min
OrchestrationPlannerDecompose goals into stepsLLM-based planner
OrchestrationRouterRoute to appropriate agentIntent classifier
OrchestrationTask QueueAsync task processingRabbitMQ, Redis
IntelligenceLLM ServiceReasoning engineOpenAI, Anthropic, self-hosted
IntelligenceMemoryCross-session contextVector DB + Redis
IntelligenceRetrieverKnowledge accessRAG pipeline
ExecutionTool RegistryAvailable toolsPlugin system
ExecutionExecution EngineRun tools safelySandboxed Docker
ObservabilityTracingPer-request flowOpenTelemetry
ObservabilityLoggingEvery action recordedStructured logs
SafetyGuardrailsInput/output validationPrompt injection detection
SafetyHuman-in-LoopCritical action approvalApproval queue

flowchart LR
subgraph SCALING["Production Scaling Considerations"]
S1["🎯 Horizontal Scaling\nStateless agents behind load balancer"]
S2["💾 State Management\nRedis + PostgreSQL for persistence"]
S3["⚡ Caching\nResponse cache + embedding cache"]
S4["🔀 Queue-Based Processing\nAsync for long-running tasks"]
S5["🌐 Multi-Model Routing\nCheap model for simple, expensive for complex"]
end
style S1 fill:#3b82f6,color:#fff
style S2 fill:#8b5cf6,color:#fff
style S3 fill:#f59e0b,color:#fff
style S4 fill:#22c55e,color:#fff
style S5 fill:#ef4444,color:#fff
StrategyHow It WorksImpact
Stateless agentsAgent state in Redis, not in processEasy horizontal scaling
Task queuesLong tasks queued, workers pick upHandle 10K+ concurrent tasks
CachingCache LLM responses (TTL: 1 hour)40-60% cost reduction
Model routingGPT-4o-mini for 80% of tasks, GPT-4o for 20%60% cost reduction
Parallel executionIndependent agent steps in parallel2-5x speedup

flowchart TD
subgraph SEC["Security Layers"]
S1["🔐 Authentication\nJWT / OAuth / SSO"]
S2["📋 Authorization\nRBAC + Scope-based permissions"]
S3["🛡️ Prompt Protection\nInput sanitization\nInjection detection\nPII masking"]
S4["🔒 Execution Isolation\nSandboxed tool execution\nDocker containers\nNo network for code exec"]
S5["📝 Audit Logging\nEvery action logged\nImmutable logs\nFull traceability"]
end
S1 --> S2 --> S3 --> S4 --> S5
style S1 fill:#3b82f6,color:#fff
style S2 fill:#8b5cf6,color:#fff
style S3 fill:#f59e0b,color:#fff
style S4 fill:#ef4444,color:#fff
style S5 fill:#22c55e,color:#fff

  1. Design for failure from day one — Every component should have a fallback. LLM down → use cache. Tool fails → retry with backoff.
  2. Implement observability before launch — You need tracing, logging, and metrics before you have users. Debugging blind is impossible.
  3. Use the cheapest model that works — Route simple tasks to GPT-4o-mini, complex tasks to GPT-4o. Save 60% on API costs.
  4. Always have human-in-the-loop for destructive actions — Delete, deploy, send email, pay — require human approval.
  5. Set hard limits — Max iterations (25), max tokens (100K), max cost per task ($1), max execution time (5 min).

MistakeImpactFix
No observabilityCan’t debug agent failuresAdd tracing from day one
No rate limitingOne user exhausts API budget100 req/min per user
No sandboxingCode execution vulnerabilityRun all code in Docker containers
Single modelExpensive for simple tasksRoute between cheap and expensive models
No cachingSame query hits LLM 1000 timesCache LLM responses with TTL

Q: What are the main components of a production AI agent architecture?

Gateway (auth, rate limiting), Orchestration (planner, router), Intelligence (LLM, memory, retriever), Execution (tools, sandbox), Observability (tracing, logs), and Safety (guardrails, human-in-the-loop).

Q: How does Cursor differ from a simple LLM code generator?

Cursor has context gathering (reads open files, imports, errors), code indexing (AST + embeddings), terminal integration (runs code, checks output), and an iterative loop (generate → run → fix → repeat). A simple LLM just generates code once.

Q: Explain the two-phase architecture used by Devin (planner + specialized agents).

Devin uses a planner agent that creates a high-level plan: “Build a login page” → “1. Create React component 2. Add form validation 3. Style with CSS 4. Test.” Then specialized agents execute each step: Shell Agent runs npm commands, Editor Agent writes code, Browser Agent tests the UI. All agents share a common memory of project state.

Q: Design a caching strategy for a production agent that processes 10,000 queries per day.

Three-tier cache: (1) Exact match cache — Redis, TTL 24h. If the same query was asked before, return cached response. (30% hit rate). (2) Semantic cache — Vector DB. If a new query is > 95% similar to a cached one, return adapted cached response. (15% hit rate). (3) Embedding cache — Cache query embeddings. If the same query is asked, skip embedding API call. (60% hit rate). Invalidation: Clear cache when tools or knowledge base change.

Q: Design a multi-model routing strategy for an agent that uses GPT-4o, Claude, and a self-hosted Llama model.

Router — Intent classifier determines task complexity. Tiers: (1) Simple (greetings, FAQs) → Llama 70B (self-hosted, $0). (2) Medium (research, writing) → GPT-4o-mini ($0.15/1M tokens). (3) Complex (coding, analysis) → GPT-4o ($2.50/1M tokens). (4) Critical (medical, legal) → Claude 3.5 Opus ($15/1M tokens). Failover: If GPT-4o is down → Claude. If Claude is down → GPT-4o. Cost savings: 70% of queries hit tiers 1-2, saving 80% vs using GPT-4o for everything.

Q: Design an agent system that can handle 1M users with < 3 second response time for simple queries.

Architecture: Stateless API servers (auto-scale based on CPU), Redis for session state, PostgreSQL for persistence. Simple queries (< 3s): Direct LLM call with cached tools. Complex queries (> 3s): Async via task queue, user polls for result. Caching: Semantic cache for 80% of queries. Model routing: GPT-4o-mini for 95% of queries, GPT-4o for 5%. Scaling: 50 API servers handle 1M users. Cost: ~$50K/mo in API costs.

Q: How would you architect an agent that can research, write, and publish blog posts autonomously?

Crew setup: Research Agent (searches 5 sources) → Writer Agent (drafts post) → Editor Agent (checks quality) → Publisher Agent (uploads to CMS). Guardrails: Fact-check gate before publishing. Human-in-loop: Final approval before publishing. Memory: Past posts stored for style consistency. Monitoring: Track publish success rate, reader engagement. Scaling: 5 parallel crews for 20 posts/day.


ComponentPurposeProduction Consideration
API GatewayAuth, routing, rate limitingScale horizontally, set limits
PlannerGoal decompositionRe-plan every 3-5 steps
LLM ServiceReasoning engineMulti-model routing for cost
Tool RegistryAvailable actionsSandbox execution, validate output
MemoryCross-session contextVector DB + Redis
ObservabilityDebugging & monitoringTrace every step from day one
GuardrailsSafety & validationBlock prompt injection, PII leaks
Human-in-LoopCritical action approvalDelete, deploy, pay require approval

Previous: 13 — AutoGen & OpenAI Agents SDK

Next: 15 — AI Agent Projects & Complete Roadmap