Observability, Health Checks, Prometheus Metrics & Resilience

Operating mission-critical Python applications in production requires deep Observability (Structured JSON Logging, OpenTelemetry Distributed Tracing, Prometheus Metrics), Kubernetes Health Probes (/livez vs /readyz), and Resilience Patterns (Circuit Breakers, Bulkheads, Rate Limiters).

This chapter details structlog JSON formatting, OpenTelemetry trace context propagation, Prometheus RED metric collection, Kubernetes probes, and Circuit Breaker implementations.


1. Structured JSON Logging with structlog

Standard text log lines (2026-08-11 INFO User logged in) are difficult for log aggregators (Datadog, ElasticSearch) to parse. Production systems use Structured JSON Logging:

import structlog
import logging

# Configure structlog for JSON formatting
structlog.configure(
    processors=[
        structlog.contextvars.merge_contextvars,
        structlog.processors.add_log_level,
        structlog.processors.TimeStamper(fmt="iso"),
        structlog.processors.JSONRenderer(),  # Outputs valid JSON strings!
    ],
    logger_factory=structlog.PrintLoggerFactory(),
)

logger = structlog.get_logger()

# Emit structured log with context fields
logger.info(
    "user_authenticated",
    user_id=42,
    ip_address="192.168.1.1",
    execution_time_ms=14.2
)

JSON Output:

{"event": "user_authenticated", "user_id": 42, "ip_address": "192.168.1.1", "execution_time_ms": 14.2, "level": "info", "timestamp": "2026-08-11T22:00:00Z"}

2. OpenTelemetry Distributed Tracing

In a microservice architecture, a single user request spans multiple services. OpenTelemetry propagates trace context (trace_id, span_id) across HTTP boundaries using W3C Trace Context headers (traceparent):

OpenTelemetry Distributed Trace Context Propagation:

[ Client ] ──> [ API Gateway (Trace ID: abc-123) ]
                       |
                       v (Injects HTTP Header: traceparent: 00-abc-123-span1-01)
               [ User Service (Trace ID: abc-123, Span ID: span2) ]
                       |
                       v (Injects HTTP Header: traceparent: 00-abc-123-span2-01)
               [ Payment Service (Trace ID: abc-123, Span ID: span3) ]

3. Prometheus RED Metrics Collection

Monitor microservice health using the RED Method:

  • Rate: Requests processed per second (Counter).
  • Errors: Failed requests per second (Counter).
  • Duration: Request processing latency distributions (Histogram).
from prometheus_client import Counter, Histogram
import time

REQUEST_COUNT = Counter("http_requests_total", "Total HTTP Requests", ["method", "endpoint", "status"])
REQUEST_LATENCY = Histogram("http_request_duration_seconds", "HTTP Latency", ["endpoint"])

def handle_request(endpoint: str):
    start = time.perf_counter()
    status = "200"
    try:
        # Execute business logic...
        pass
    except Exception:
        status = "500"
        raise
    finally:
        duration = time.perf_counter() - start
        REQUEST_COUNT.labels(method="GET", endpoint=endpoint, status=status).inc()
        REQUEST_LATENCY.labels(endpoint=endpoint).observe(duration)

4. Kubernetes Probes & Circuit Breakers

1. Kubernetes Health Check Probes:

  • Liveness Probe (/livez): Checks if the Python process is alive. If it fails, Kubernetes restarts the container pod.
  • Readiness Probe (/readyz): Checks if the app can serve traffic (e.g. DB connection is alive). If it fails, Kubernetes removes the pod from load balancer routing without restarting it!

2. Circuit Breaker Pattern:

Prevents cascading failure by failing fast when an upstream dependency is down:

Circuit Breaker State Machine:

[ CLOSED (Normal) ] ──> (Error threshold exceeded) ──> [ OPEN (Fail Fast) ]
         ^                                                      |
         β”‚                                                      v (Sleep window expires)
         └─── (Success threshold met) ◄── [ HALF-OPEN (Testing) ]
Display Options
Appearance
Text Size
100%