Skip to content

Observability

Observability is the ability to understand what’s happening inside your system by looking at the outputs. It has three pillars: logging, metrics, and tracing.


flowchart TB
App["📱 Application"] --> Logs["📝 Logs<br/>Events (errors, info)"]
App --> Metrics["📊 Metrics<br/>Counters, gauges, histograms"]
App --> Traces["🔍 Traces<br/>Request spans across services"]
Logs --> Dashboard["📈 Monitoring Dashboard<br/>(Grafana, Datadog)"]
Metrics --> Dashboard
Traces --> Dashboard
Dashboard --> Alerts["🔔 Alerts<br/>(PagerDuty, Slack)"]
style App fill:#7c3aed,color:#fff
style Logs fill:#4f46e5,color:#fff
style Metrics fill:#6366f1,color:#fff
style Traces fill:#8b5cf6,color:#fff
style Dashboard fill:#059669,color:#fff
style Alerts fill:#dc2626,color:#fff

Log LevelWhen to Use
ERRORA failure that needs investigation (DB connection lost, payment failed)
WARNSomething unexpected but not critical (rate limit approaching, retry attempt)
INFOImportant events (user registered, order placed)
DEBUGDetailed information for debugging (not in production)

Best practice: Log in structured format (JSON), not plain text.

{ "level": "ERROR", "service": "payment", "msg": "Payment declined",
"userId": 123, "orderId": 456, "errorCode": "INSUFFICIENT_FUNDS",
"traceId": "abc-123-def", "timestamp": "2024-01-15T10:30:00Z" }

Metric TypeExample
CounterTotal requests, total errors (only increases)
GaugeCurrent CPU, memory usage, queue size (goes up and down)
HistogramRequest latency p50, p95, p99

The “Four Golden Signals” (Google SRE):

  1. Latency — time to serve requests (p50, p95, p99)
  2. Traffic — requests per second
  3. Errors — rate of failed requests (explicit 5xx + implicit errors like wrong data)
  4. Saturation — how “full” your system is (CPU, memory, queue depth)

A trace tracks a single request as it travels through multiple services:

sequenceDiagram
participant Client as 📱 Client
participant API as 🚪 API Gateway
participant Auth as 🔐 Auth Service
participant DB as 🗄️ Database
Client->>API: POST /order (traceId: abc)
API->>Auth: Validate token (span: auth)
Auth-->>API: ✅ Valid
API->>DB: Create order (span: db)
DB-->>API: ✅ Order 123
API-->>Client: ✅ 201 Created
Note over Client,API: Trace = abc (spans: auth=15ms, db=45ms, total=80ms)

Tools: Jaeger, Zipkin, OpenTelemetry, AWS X-Ray


ComponentWhat to WatchWhy
API GatewayQPS, error rate, p99 latencyIs the system reachable?
ApplicationCPU, memory, GC pauses, thread countIs the app healthy?
DatabaseConnection count, query latency, slow queriesIs the DB the bottleneck?
CacheHit rate, memory usage, evictionsIs caching working?
QueueQueue depth, consumer lagAre consumers keeping up?

  • Observability is essential but adds cost (storage for logs, overhead for tracing).
  • Too little observability = blind debugging.
  • Too much observability = alert fatigue, high storage costs.
  • Start with logs + the 4 golden signals. Add tracing when you have microservices.

  • Observability = seeing what’s happening inside your system.
  • Three pillars: logs (events), metrics (numbers), traces (request paths).
  • Monitor the 4 golden signals: latency, traffic, errors, saturation.