Skip to content

05 — Load Balancers

A load balancer distributes incoming network traffic across multiple backend servers. It’s a critical component for scalability, high availability, and fault tolerance.

Analogy: A load balancer is like a host at a restaurant — when customers arrive, the host directs them to the first available table, making sure no server is overwhelmed while another sits idle.


Without a load balancer:

  • One server gets all traffic and crashes
  • Users experience downtime when a server fails
  • Cannot add or remove servers without downtime
  • SSL termination and rate limiting must be handled by each server individually

flowchart TB
Users["🌍 Users"] --> DNS["DNS<br/>Round Robin"]
DNS --> LB["Load Balancer"]
subgraph LB_Internals["Load Balancer Internals"]
Algo["Scheduling Algorithm"]
HC["Health Checks"]
SSL["SSL Termination"]
Stick["Session Persistence"]
end
LB --> Backend1["Backend Server 1"]
LB --> Backend2["Backend Server 2"]
LB --> Backend3["Backend Server 3"]
LB --> BackendN["Backend Server N"]
style Users fill:#f59e0b,color:#fff
style LB fill:#7c3aed,color:#fff
style LB_Internals fill:#3b82f6,color:#fff

TypeLayerDescriptionExample
ALB (Application LB)L7HTTP/HTTPS, path-based routing, host-based routingAWS ALB
NLB (Network LB)L4TCP/UDP, extreme performance, static IPAWS NLB
GLB (Gateway LB)L3IP packet-level, for appliancesAWS GLB
Classic LBL4/L7Legacy, simple TCP/HTTP LBAWS CLB (retired)
HAProxyL4/L7Open-source, flexible, battle-testedSelf-managed
NGINXL7Reverse proxy + LB, great for HTTPSelf-managed

flowchart TB
Algo["Load Balancing Algorithms"] --> RR["Round Robin<br/>Requests distributed equally<br/>Simple, no state needed"]
Algo --> WRR["Weighted Round Robin<br/>More traffic to powerful servers<br/>Uneven capacity"]
Algo --> LeastConn["Least Connections<br/>Send to server with fewest<br/>active connections"]
Algo --> IPHash["IP Hash<br/>Same client → same server<br/>Session persistence"]
Algo --> Random["Random<br/>Random selection<br/>Simple, statistically fair"]
Algo --> Geolocation["Geolocation<br/>Route based on client location<br/>Lowest latency"]
style Algo fill:#7c3aed,color:#fff
style RR fill:#3b82f6,color:#fff
style LeastConn fill:#059669,color:#fff
style IPHash fill:#f59e0b,color:#fff
AlgorithmBest ForProsCons
Round RobinEqual-capacity serversSimple, fairDoesn’t consider load
Least ConnectionsVarying request durationsBetter resource useSlightly more complex
IP HashSticky sessionsNo extra config neededUneven distribution if clients from same IP
WeightedMixed server specsOptimizes heterogeneous clustersManual weight tuning

sequenceDiagram
participant LB as Load Balancer
participant S1 as Server 1 ✅
participant S2 as Server 2 ❌ (Down)
LB->>S1: Health check (GET /health)
S1-->>LB: 200 OK
Note over S1: Healthy → send traffic
LB->>S2: Health check (GET /health)
S2-->>LB: Timeout (no response)
Note over S2: Unhealthy → remove from pool
LB->>S2: Retry health check (after interval)
S2-->>LB: 200 OK
Note over S2: Back healthy → add to pool
loop Every N seconds
LB->>S1: Periodic check
LB->>S2: Periodic check
end
Health Check TypeWhat It ChecksInterval
TCPPort is open (connection succeeds)5-30 seconds
HTTPEndpoint returns 2005-30 seconds
HTTPSSecure endpoint returns 20010-30 seconds
Custom scriptComplex logic (e.g., DB connectivity)30-60 seconds

flowchart TB
User1["User A"] --> LB
User2["User B"] --> LB
LB -->|"User A → Server 1<br/>Session stored locally"| S1["Server 1"]
LB -->|"User B → Server 2<br/>Session stored locally"| S2["Server 2"]
subgraph Better["Better Approach: Shared Session Store"]
S_Shared["Server 1"]
S_Shared2["Server 2"]
Redis["Redis<br/>Shared session store"]
S_Shared --> Redis
S_Shared2 --> Redis
end
style User1 fill:#3b82f6,color:#fff
style User2 fill:#059669,color:#fff
style LB fill:#7c3aed,color:#fff
style Redis fill:#ef4444,color:#fff

Sticky sessions route the same client to the same server. Avoid when possible — use a shared session store (Redis) instead so any server can handle any request.


ApproachProsCons
ALB (L7)Smart routing, path-basedSlightly slower than NLB
NLB (L4)Ultra-fast, static IPNo HTTP-level features
Sticky sessionsSimple session handlingUneven load, server affinity issues
Health checksAutomatic failoverAdds complexity
Multiple LBsHigh availabilityCost, management overhead

StrategyHow It Works
DNS Round RobinMultiple LB DNS entries, client picks one
Multi-region LBGeo-routing to closest region’s LB
LB Auto ScalingAutomatically register/deregister new servers
Layer 7 RoutingRoute /api/* to API servers, /static/* to CDN
Weighted routingCanary deployments (5% → 50% → 100%)

  1. How does a load balancer improve availability?
  2. What’s the difference between L4 and L7 load balancing?
  3. How do health checks work and what should they check?
  4. What is the difference between sticky sessions and a shared session store?
  5. How would you design load balancing for a global application?

SystemLoad Balancer Strategy
AmazonMulti-region Route53 + ALB per region + NLB for internal
NetflixZuul (gateway) + Ribbon (client-side LB) → migrated to Spring Cloud Gateway
CloudflareGlobal anycast + L4 LB at edge + L7 LB at origin
KubernetesIngress Controller (L7) + Service (L4) — built-in LB abstraction

  • Load balancers distribute traffic across servers — essential for scaling and HA
  • L4 (NLB) = fast, works for any TCP/UDP traffic
  • L7 (ALB) = smarter, understands HTTP paths and headers
  • Health checks automatically remove dead servers from the pool
  • Avoid sticky sessions — use a shared cache (Redis) for session data instead
  • Multiple LB layers = multiple layers of scale and redundancy