Skip to content

14 — Monitoring and Observability

Monitoring tells you when something is wrong. Observability lets you understand why it’s wrong. Together, they form the foundation of reliable, debuggable systems.

Analogy: Monitoring is like a car’s dashboard warning lights — they tell you when the engine is overheating. Observability is the mechanic’s diagnostic tools — they tell you which hose is leaking and why.


Without proper monitoring and observability:

  • Blind deployment — no idea if the new version is working
  • Reactive debugging — users report problems before you know about them
  • Long MTTR — hours spent finding root causes
  • No performance baselines — can’t tell what’s normal vs abnormal
  • Wasted resources — no visibility into utilization

flowchart TB
Obs["Observability"] --> Logs["📝 Logs<br/>Events that happened<br/>'User 123 placed order'"]
Obs --> Metrics["📊 Metrics<br/>Aggregated measurements<br/>'500 requests/sec, 95% latency: 150ms'"]
Obs --> Traces["🔍 Traces<br/>Request lifecycle<br/>'Order took 2.3s across 4 services'"]
Logs --> LogTools["ELK Stack, Loki, CloudWatch Logs"]
Metrics --> MetricTools["Prometheus, Grafana, Datadog"]
Traces --> TraceTools["Jaeger, Zipkin, X-Ray"]
style Obs fill:#7c3aed,color:#fff
style Logs fill:#3b82f6,color:#fff
style Metrics fill:#059669,color:#fff
style Traces fill:#f59e0b,color:#fff

flowchart LR
SLI["SLI — Service Level Indicator<br/>(What we measure)<br/>Latency: P99 < 200ms"] --> SLO["SLO — Service Level Objective<br/>(Our target)<br/>99.9% of requests < 200ms"]
SLO --> SLA["SLA — Service Level Agreement<br/>(Contract with users)<br/>99.9% uptime, or refund"]
style SLI fill:#3b82f6,color:#fff
style SLO fill:#059669,color:#fff
style SLA fill:#ef4444,color:#fff
TermDefinitionExample
SLIA measured indicator of service performanceP99 latency, error rate, throughput
SLOYour target for the SLI99.9% of requests succeed within 200ms
SLAA contractual commitment to customers99.95% uptime, service credits if violated
Error BudgetHow much failure is allowed before SLO violation0.1% errors/month = 43 min downtime

flowchart TB
subgraph Sources["Data Sources"]
App["Application Logs"]
Infra["Infrastructure Metrics<br/>CPU, RAM, Disk"]
DB["Database Metrics"]
LB["Load Balancer Metrics"]
end
subgraph Collection["Collection & Storage"]
Agent["Agents / Exporters"]
TSDB["Time Series DB<br/>Prometheus / InfluxDB"]
LogDB["Log Storage<br/>Elasticsearch / Loki"]
TraceDB["Trace Storage<br/>Jaeger / Tempo"]
end
subgraph Visualization["Visualization & Alerting"]
Grafana["Grafana Dashboards"]
Alert["Alert Manager<br/>PagerDuty, Slack"]
end
Sources --> Collection
Collection --> Visualization
style Sources fill:#3b82f6,color:#fff
style Collection fill:#7c3aed,color:#fff
style Visualization fill:#059669,color:#fff

Google’s SRE book defines four golden signals of monitoring:

SignalWhat It MeasuresWhat to Watch For
LatencyTime to serve a requestIncreasing P99 = performance issue
TrafficHow many requestsSpike = DDoS or viral event; Drop = routing issue
ErrorsRate of failed requestsHTTP 5xx, exceptions, timeouts
SaturationHow “full” the service isCPU > 80%, memory > 80%, connections exhausted

sequenceDiagram
participant Client as Client
participant GW as API Gateway
participant Auth as Auth Service
participant Order as Order Service
participant DB as Database
Note over Client,DB: Trace ID: abc-123 (passed via headers)
Client->>GW: POST /orders
GW->>Auth: Validate token (trace: abc-123)
Auth-->>GW: Valid (5ms)
GW->>Order: Create order (trace: abc-123)
Order->>DB: INSERT order (trace: abc-123)
DB-->>Order: OK (20ms)
Order->>DB: INSERT order_items (trace: abc-123)
DB-->>Order: OK (15ms)
Order->>Order: Calculate total (trace: abc-123)
Order-->>GW: Order created (50ms total)
GW-->>Client: 201 Created (65ms total)

Distributed tracing tracks a request across multiple services. Each service adds its own span with timing. You can see exactly where time is spent.


PracticeWhy
Structured JSON logsMachine-parseable, queryable in ELK/Loki
Include request/correlation IDTrace across services
Log levels: DEBUG, INFO, WARN, ERRORFilter by severity
Don’t log PIIPasswords, credit cards, personal data
Centralized log aggregationSingle place to search all logs
Log on entry and exit of critical pathsDebugging complex flows
Use sampling for high-volume logsDon’t overload the logging system

LayerKey Metrics
Web/APIRPS, latency (P50/P95/P99), error rate, active connections
ApplicationGC pauses, thread pool, queue depth, cache hit rate
DatabaseQPS, slow queries, connections, replication lag
InfrastructureCPU, memory, disk, network I/O
BusinessDAU, signups, orders, revenue

DecisionProsCons
Log everythingFull visibilityExpensive storage, noise
Sample tracesLow overheadMissing rare issues
High-precision metricsAccurate dataHigher storage cost
Self-hosted (Prometheus+Grafana)Full control, no vendor lockOps overhead
Managed (Datadog, New Relic)Easy setup, great UXCost scales with data volume

StrategyDescription
SamplingRecord a percentage of requests (1-10% at scale)
AggregationPre-aggregate metrics (not every data point)
Log retention tiersHot (7 days) → Warm (30 days) → Cold (1 year)
Structured loggingJSON logs → easier to query and analyze
Sidecar patternLog/metrics agent per pod (Kubernetes)

  1. What’s the difference between monitoring and observability?
  2. What are the three pillars of observability?
  3. Explain SLI, SLO, and SLA with examples.
  4. How does distributed tracing work?
  5. What would you monitor for a payment processing system?

SystemObservability Approach
GoogleBuilt Borgmon (predecessor to Prometheus), SRE golden signals
NetflixAtlas (metrics), Spinnaker (alerting), custom distributed tracing
UberM3 (metrics), Jaeger (tracing), ELK (logs)
DatadogSaaS observability — traces, metrics, logs unified

  • Monitoring = dashboards, alerts, knowing when something is wrong
  • Observability = being able to answer why by looking at logs, metrics, and traces
  • Three pillars: Logs (events), Metrics (aggregated numbers), Traces (request flow)
  • SLI/SLO/SLA = what you measure / your target / your contract with users
  • Four golden signals: Latency, Traffic, Errors, Saturation
  • Distributed tracing tracks a request across all services it touches
  • Start simple — monitor the four golden signals before adding complex tooling