Skip to content

03 — Capacity Estimation

Capacity estimation is the process of calculating the resources your system will need — servers, storage, bandwidth, memory — based on expected usage. It ensures you design for the right scale from day one.

Analogy: Estimating capacity is like planning a wedding. You need to know how many guests (traffic), how much food (storage), and how many waitstaff (servers) — book too few and it’s a disaster, too many and you waste money.


Building a system without capacity estimation leads to:

  • Under-provisioning — servers crash under load, users get timeouts
  • Over-provisioning — wasting money on unused resources
  • Wrong architecture — picking a database that can’t scale to your data volume
  • Surprise costs — unexpected bandwidth or storage bills

MetricWhy It Matters
Daily Active Users (DAU)Base for all other estimates
Requests per user per dayDetermines total traffic
Read/Write ratioShapes caching & DB strategy
Data per requestStorage & bandwidth requirements
Peak traffic multiplier5-10x average during peak hours

flowchart TB
DAU["DAU: 10M users"] --> Requests["Total Requests/Day<br/>DAU × Actions/User"]
Requests --> PerSecond["Requests/Second<br/>Total / 86400 × Peak Factor"]
PerSecond --> Servers["Server Count<br/>RPS / Capacity per Server"]
PerSecond --> Bandwidth["Bandwidth (Mbps)<br/>RPS × Data per Request"]
StoragePerUser["Storage/User: 1MB"] --> TotalStorage["Total Storage/Day<br/>Storage/User × New Users/Day"]
TotalStorage --> TotalStorage2["Total Storage/Year<br/>× 365 × Growth Rate"]
style DAU fill:#f59e0b,color:#fff
style PerSecond fill:#3b82f6,color:#fff
style Servers fill:#059669,color:#fff
style Bandwidth fill:#ef4444,color:#fff
style TotalStorage fill:#7c3aed,color:#fff

Formula: Requests per second (RPS) = (DAU × Actions per user per day) / 86400 × Peak factor

VariableExample ValueHow to Estimate
DAU10,000,000Product/marketing estimate
Actions/user/day50Avg tweets, likes, reads (Twitter: ~50)
Total actions/day500,000,000DAU × Actions/user
Avg RPS~5,800500M / 86,400 seconds
Peak factor5xTraffic spikes during events / business hours
Peak RPS~29,000Avg RPS × Peak factor

Formula: Storage per year = (Storage per user × New users per year) + (Data generated per action × Actions per year)

Data TypePer ItemAnnual Growth
User profile10 KB10M new users → 100 GB
Tweet/text post500 bytes10B tweets → 5 TB
Image (compressed)200 KB1B images → 200 TB
Video (compressed)50 MB10M videos → 500 TB
Logs1 KB per request180B requests → 180 TB
Database indexes50% overhead30% of data size
Backups3x dataFull + differential

Formula: Bandwidth (Mbps) = RPS × Average response size × 8 / 1,000,000

ComponentInbound (Mbps)Outbound (Mbps)
API requests5002,000
Image uploads1,000200 (CDN serves)
Video uploads5,00020,000 (CDN)
Logs10050
Database replication—300
Total~6,600 Mbps~22,550 Mbps

Formula: Servers needed = Peak RPS / Capacity per server

Server TypeCapacity per ServerServers Needed (29K RPS)
Web server (small)1,000 RPS~30
Web server (medium)5,000 RPS~6
Web server (large)20,000 RPS~2 (but need more for HA)
Cache node50,000 QPS~2 (then redundancy)
DB (primary)10,000 QPS~3 (1 primary + 2 replicas)

flowchart LR
Mem["Memory Per Request<br/>e.g., 50 MB"] --> Concurrent["Concurrent Requests<br/>e.g., 5,000"]
Concurrent --> TotalMem["Total Memory<br/>50 MB × 5,000 = 250 GB"]
TotalMem --> MemServers["Servers (64 GB each)<br/>250 / 64 = ~4 servers"]
CPU["CPU Time Per Request<br/>e.g., 50ms"] --> CPU_Total["Total CPU Time/sec<br/>50ms × 5,000 = 250 sec CPU"]
CPU_Total --> CPU_Servers["CPU Cores Needed<br/>250 / 0.75 utilization = ~333 cores"]
CPU_Servers --> CPU_Servers2["Servers (64 cores each)<br/>333 / 64 = ~6 servers"]
style Mem fill:#3b82f6,color:#fff
style TotalMem fill:#7c3aed,color:#fff
style CPU fill:#f59e0b,color:#fff
style CPU_Total fill:#059669,color:#fff

UnitEquivalent
1 million requests/day~12 RPS average
DAU × 10%Approximate concurrent users during peak
Peak traffic5-10x average hourly traffic
Database disk space~3x data size (data + indexes + logs)
Single web server1,000-10,000 RPS (depends on complexity)
Single DB server1,000-10,000 QPS (depends on query complexity)
1 Gbps bandwidth~125 MB/s data transfer

ChoiceBenefitCost
Over-estimateSafe, handles spikesWasted resources, higher cost
Under-estimateLower initial costPerformance issues, emergency scaling
Start small, auto-scaleCost-efficient, flexibleComplexity of scaling infrastructure
Build for peakNever slowIdle resources during off-peak

StrategyWhen
Vertical scalingQuick fix, simple apps, predictable growth
Horizontal scalingUnpredictable traffic, HA requirements
Auto-scalingVariable traffic patterns (social media, e-commerce)
Reserved capacitySteady baseline traffic (always on)
Spot/preemptibleBatch processing, non-critical workloads

  1. Estimate the storage needed for a Twitter-like app with 100M DAU.
  2. How many servers do we need for a video streaming platform handling 1M concurrent viewers?
  3. What’s the bandwidth cost for serving 10GB/hour of video content to 100K users?
  4. How do you estimate the cache size needed for a social media feed?
  5. A system crashes at 10,000 RPS on 4 servers. How do you plan for 50,000 RPS?

SystemCapacity ChallengeSolution
Twitter500M tweets/day, reads per second hugeTimeline fanout, heavy caching
Netflix200M+ users streaming simultaneouslyCDN (Open Connect), adaptive bitrate
WhatsApp100B+ messages/dayErlang-based custom server, minimal per-message storage
StripeMillions of API requests/dayIdempotency, horizontal scaling of API servers

  • Start with DAU and actions per user — everything else flows from these numbers
  • Peak traffic is 5-10x average — design your architecture for the peak, not the average
  • Storage grows forever — plan for data lifecycle (hot → warm → cold → delete)
  • Bandwidth costs money — cache aggressively at CDN and application layers
  • One server handles 1K-10K RPS — use this as your mental baseline for estimation
  • Always add a safety margin (1.5x-3x) — estimates are guesses, not guarantees