18. Prompt Optimization
Introduction
Section titled “Introduction”A prompt that works is good. A prompt that works efficiently is production-ready.
Prompt optimization is the practice of making prompts faster, cheaper, and more reliable — reducing token count, improving accuracy, and minimizing latency while maintaining or improving output quality.
Why This Concept Exists
Section titled “Why This Concept Exists”The Story
Section titled “The Story”You have a prompt that works. It produces good results. But in production, every token costs money and every millisecond of latency matters.
Your prompt is 2,000 tokens and takes 3 seconds per call. At 10,000 calls per day, that’s $30/day and 8 hours of total latency.
After optimization, your prompt is 500 tokens and takes 1 second. At 10,000 calls/day, that’s $7.50/day and 2.8 hours of total latency.
Over a year, that’s $8,000+ saved and 2,000 hours of latency eliminated.
flowchart TD subgraph BEFORE["Before Optimization"] B1["2,000 token prompt"] --> B2["3 seconds per call"] B2 --> B3["$30/day for 10K calls"] end
subgraph AFTER["After Optimization"] A1["500 token prompt"] --> A2["1 second per call"] A2 --> A3["$7.50/day for 10K calls"] end
style BEFORE fill:#ef4444,color:#fff style AFTER fill:#22c55e,color:#fffReal-World Analogy
Section titled “Real-World Analogy”Code Optimization
Section titled “Code Optimization”A junior developer writes code that works but is slow and memory-intensive. A senior engineer refactors it to be efficient — same output, fewer resources.
Prompt optimization is the same: same output quality, fewer tokens, faster response.
Prompt optimization is code refactoring for LLM prompts.
Optimization Dimensions
Section titled “Optimization Dimensions”flowchart TD OPT["Prompt Optimization"] --> TOKEN["Token Optimization\nFewer tokens, same quality"] OPT --> LATENCY["Latency Optimization\nFaster responses"] OPT --> COST["Cost Optimization\nCheaper per call"] OPT --> QUALITY["Quality Optimization\nBetter outputs"] OPT --> RELIABILITY["Reliability Optimization\nMore consistent"]
style OPT fill:#8b5cf6,color:#fff style TOKEN fill:#3b82f6,color:#fff style LATENCY fill:#22c55e,color:#fff style COST fill:#f59e0b,color:#fff style QUALITY fill:#ec4899,color:#fff style RELIABILITY fill:#14b8a6,color:#fffToken Optimization
Section titled “Token Optimization”Techniques
Section titled “Techniques”| Technique | Token Savings | Impact on Quality |
|---|---|---|
| Remove filler words | 5-10% | None |
| Use concise instructions | 10-20% | Slightly positive |
| Replace descriptions with examples | 20-40% | Positive |
| Reduce persona description | 5-15% | Minimal |
| Remove redundant constraints | 5-10% | Neutral |
| Use abbreviations for common terms | 2-5% | Minimal risk |
Before and After
Section titled “Before and After”❌ Before (150 tokens):"I would like you to please act as a senior software engineerwho has extensive experience with the Python programming languageand specifically with the Django web framework. Could you pleasetake a look at the following code and provide me with your feedbackon any potential issues or improvements that you might notice?"
✅ After (50 tokens):"Role: Senior Django engineerTask: Review this code for issues and improvementsCode: [code]"Latency Optimization
Section titled “Latency Optimization”flowchart LR LATENCY["Latency Factors"] --> TOKENS_IN["Input Tokens\nMore input = slower"] LATENCY --> TOKENS_OUT["Output Tokens\nMore output = slower"] LATENCY --> MODEL["Model Size\nLarger model = slower"] LATENCY --> TEMPERATURE["Temperature\nHigher = more random = slower?"]
style LATENCY fill:#f59e0b,color:#fff| Technique | Latency Reduction | Trade-off |
|---|---|---|
| Limit max tokens | Significant | Shorter responses |
| Use smaller model | Significant | Lower quality |
| Reduce input context | Proportional | Might lose context |
| Prompt caching | Variable | Implementation overhead |
| Stream responses | Perceived reduction | Same total time |
Cost Optimization
Section titled “Cost Optimization”Cost Calculation
Section titled “Cost Calculation”Cost per call = (Input Tokens × Input Price) + (Output Tokens × Output Price)
Example (GPT-4o):Input: 500 tokens × $2.50/1M = $0.00125Output: 200 tokens × $10.00/1M = $0.00200Total: $0.00325 per call
At 10,000 calls/day: $32.50/day = ~$975/monthOptimization Strategies
Section titled “Optimization Strategies”| Strategy | Cost Reduction | Implementation |
|---|---|---|
| Reduce prompt size | 20-50% | Edit and compress |
| Use shorter outputs | 30-60% | Limit max_tokens |
| Cache common requests | 40-80% | Cache layer |
| Batch requests | 15-25% | Combine multiple inputs |
| Use cheaper model | 50-90% | Try smaller/older models |
Quality Optimization
Section titled “Quality Optimization”Techniques Beyond Token Reduction
Section titled “Techniques Beyond Token Reduction”flowchart TD QUALITY["Quality Optimization"] --> TEST["A/B Testing\nCompare prompt variants"] QUALITY --> FEEDBACK["User Feedback\nIncorporate real results"] QUALITY --> ITERATE["Iterative Refinement\nStep-by-step improvements"] QUALITY --> EVAL["Automated Evaluation\nMeasure quality metrics"]
style QUALITY fill:#22c55e,color:#fff| Technique | Quality Impact | Effort |
|---|---|---|
| Add examples | High | Medium |
| Better structure | Medium | Low |
| Add constraints | High | Low |
| Use system prompt | Medium | Low |
| Chain of thought | High | Medium |
Optimization Workflow
Section titled “Optimization Workflow”flowchart LR MEASURE["1. Measure\nCurrent performance"] --> IDENTIFY["2. Identify\nBottlenecks"] IDENTIFY --> OPTIMIZE["3. Optimize\nApply techniques"] OPTIMIZE --> TEST["4. Test\nVerify quality maintained"] TEST --> MEASURE TEST --> DEPLOY["5. Deploy\nTo production"] DEPLOY --> MONITOR["6. Monitor\nTrack metrics"]
style MEASURE fill:#3b82f6,color:#fff style IDENTIFY fill:#f59e0b,color:#fff style OPTIMIZE fill:#22c55e,color:#fff style TEST fill:#8b5cf6,color:#fff style DEPLOY fill:#ef4444,color:#fff style MONITOR fill:#ec4899,color:#fffReal-World Examples
Section titled “Real-World Examples”Example 1: Production Optimization
Section titled “Example 1: Production Optimization”Before (1,200 tokens):"As an AI assistant with expertise in customer support for e-commerceplatforms, I need you to analyze the following customer inquiry anddetermine the most appropriate response based on our company policiesand procedures. Please consider the customer's tone, the nature oftheir issue, their order history if available, and our current returnand refund policies..."
After (400 tokens):"System: E-commerce support agentFollow policy guidelines below for each issue type.
[Customer inquiry][policies]
Respond with: resolution, policy_ref, customer_tone, needs_escalation"→ 67% token reduction, same qualityExample 2: Format Optimization
Section titled “Example 2: Format Optimization”Before: "Please return the data in a JSON format with the following fields..."After: "Return JSON: { field1: type, field2: type }"→ 50% token reduction, more consistent outputsCommon Mistakes
Section titled “Common Mistakes”| Mistake | Why It’s Wrong |
|---|---|
| ❌ Optimizing before measuring | You don’t know what to improve without baselines |
| ❌ Reducing tokens at the expense of quality | A 50% cheaper wrong answer isn’t valuable |
| ❌ Over-optimizing for one model | Prompts optimized for GPT-4 may fail on Claude |
| ❌ Ignoring output tokens | Output tokens often cost more than input tokens |
| ❌ Not testing optimized prompts | Always verify quality after optimization |
Bad Prompt vs Good Prompt
Section titled “Bad Prompt vs Good Prompt”| Aspect | Unoptimized | Optimized |
|---|---|---|
| Length | 1,000+ tokens | 200-500 tokens |
| Clarity | Wordy, polite, indirect | Direct, concise, specific |
| Structure | Paragraph form | Structured, labeled sections |
| Instructions | Implicit | Explicit and prioritized |
| Format | Described | Shown with example |
| Redundancy | High | Minimal |
Production Examples
Section titled “Production Examples”OpenAI Prompt Optimization
Section titled “OpenAI Prompt Optimization”OpenAI’s documentation shows how to optimize prompts by removing unnecessary context, using shorter instructions, and specifying exact output formats.
Anthropic Prompt Engineering
Section titled “Anthropic Prompt Engineering”Anthropic recommends XML-style prompting and keeping prompts under 500 tokens when possible.
Interview Questions
Section titled “Interview Questions”Q: Why is prompt optimization important in production?
Because every token costs money and every millisecond of latency adds up at scale. Optimizing prompts can reduce costs by 50-80% and improve response times significantly while maintaining quality.
Intermediate
Section titled “Intermediate”Q: What techniques can reduce prompt token count without sacrificing quality?
Remove filler words, replace descriptions with examples, use structured formats instead of paragraphs, reduce persona descriptions, remove redundant constraints, and specify output format with examples.
Senior
Section titled “Senior”Q: Design a prompt optimization pipeline for a production system handling 1M+ requests per day.
I’d implement: (1) Automated prompt profiling (token count, latency, cost per call), (2) A/B testing framework for prompt variants, (3) Evaluation pipeline comparing output quality (automated + sampled human review), (4) Caching layer for identical requests, (5) Model routing (simple queries → cheaper/faster model), (6) Dynamic prompt compression based on context length, (7) Continuous monitoring with alerts for quality regression.
Summary
Section titled “Summary”| Concept | Key Point |
|---|---|
| Token Optimization | Same quality, fewer tokens |
| Latency Optimization | Faster response times |
| Cost Optimization | Cheaper per call at scale |
| Quality Optimization | Better, more consistent outputs |
| Key Principle | Measure before optimizing, test after |
Navigation
Section titled “Navigation”Previous: 17 — Meta Prompting →
Next: 19 — Prompt Testing →