KubernetesGrafana TempoDistributed Tracing

Mastering Distributed Tracing: A Production-Ready Guide to Grafana Tempo on Kubernetes

Also mirrored on Medium.

Dağıtık izleme (distributed tracing) kavramını temsil eden kapak görseli

In the modern microservices era, a single user request can traverse dozens of services, containers, and cloud regions. When something breaks, finding the root cause feels like searching for a needle in a haystack. This is where distributed tracing becomes not just useful, but essential.

Enter Grafana Tempo — the open-source, high-scale, and cost-effective distributed tracing backend that's changing how organizations handle their observability data.

Why Tempo Stands Out in the Tracing Landscape

While other tracing solutions exist, Tempo takes a unique approach:

  • Massive Scale: Built to handle billions of spans across your microservices architecture
  • Cost-Effective: Leverages object storage (S3, GCS, Azure Blob) instead of expensive proprietary databases
  • Simplicity: No complex indexing — find traces directly using trace IDs
  • OpenTelemetry Native: Seamlessly integrates with the CNCF standard for observability
  • Grafana Ecosystem: Tight integration with Grafana for a unified observability experience

Architecting Tempo for Production on Kubernetes

Let's dive into a production-ready Tempo deployment on Kubernetes. Our configuration demonstrates enterprise-grade patterns including AWS S3 integration, metrics generation, and secure service accounts.

The Core Deployment Strategy

We use the official Helm chart with a carefully crafted configuration:

# tempo.yaml - Production Configuration
serviceAccount:
  create: true
  name: tempo-sa
  annotations:
    eks.amazonaws.com/role-arn: arn:aws:iam::<account_id>:role/tempo-irsa-role

storage:
  trace:
    backend: s3
    s3:
      endpoint: "s3.eu-west-1.amazonaws.com"
      region: "eu-west-1"
      bucket: medium-tempo

traces:
  otlp:
    grpc:
      enabled: true
    http:
      enabled: true

server:
  http_listen_port: 3100

compactor:
  compaction:
    block_retention: 48h

metricsGenerator:
  enabled: true
  config:
    processor:
      service_graphs:
        dimensions: [peer_service]
  storage:
    path: /var/tempo/wal
    remote_write:
      - url: http://mimir-gateway.mimir.svc.cluster.local/prometheus/api/v1/write

service:
  type: ClusterIP

Key Production Features Explained

1. Cloud-Native Storage with AWS S3

Tempo's brilliance shines in its storage strategy. By using S3 as the primary backend, we achieve:

  • Infinite scalability — S3 handles whatever load you throw at it
  • Cost optimization — Significantly cheaper than running dedicated tracing databases
  • Durability — 11 9's of durability from AWS S3
  • Disaster recovery — Built-in replication and backup capabilities

2. IAM Roles for Service Accounts (IRSA)

Security first! Our configuration uses AWS IRSA to provide secure, short-lived credentials without storing secrets in configuration files:

serviceAccount:
  annotations:
    eks.amazonaws.com/role-arn: arn:aws:iam::<account_id>:role/tempo-irsa-role

3. Metrics Generation — The Game Changer

Tempo doesn't just store traces; it generates metrics from them:

metricsGenerator:
  enabled: true
  config:
    processor:
      service_graphs:
        dimensions: [peer_service]
  storage:
    remote_write:
      - url: http://mimir-gateway.mimir.svc.cluster.local/prometheus/api/v1/write

This feature automatically creates:

  • RED metrics (Rate, Errors, Duration) from trace data
  • Service graphs visualizing service dependencies and interactions
  • Span metrics for performance analysis across services

4. OpenTelemetry Protocol (OTLP) Support

Modern instrumentation standards are crucial:

traces:
  otlp:
    grpc:
      enabled: true
    http:
      enabled: true

This enables seamless integration with:

  • OpenTelemetry collectors
  • Instrumented applications
  • Various tracing SDKs

Deployment in Action

One-Command Deployment

Our installation script makes deployment straightforward:

#!/bin/bash
# install.sh

helm repo add grafana https://grafana.github.io/helm-charts
helm repo update

helm upgrade --install tempo grafana/tempo-distributed \
  --namespace tempo \
  --create-namespace \
  --values tempo.yaml

Validating with Real Traffic

Test your deployment with actual trace generation:

# test_trace.sh
kubectl -n tempo apply -f - <<EOF
apiVersion: batch/v1
kind: Job
metadata:
  name: tracegen
spec:
  template:
    spec:
      containers:
      - name: tracegen
        image: ghcr.io/open-telemetry/opentelemetry-collector-contrib/telemetrygen:latest
        args:
        - traces
        - --otlp-http
        - --otlp-endpoint=alloy.alloy.svc.cluster.local:4318
        - --duration=30s
        - --workers=1
        - --otlp-insecure
      restartPolicy: Never
  backoffLimit: 4
EOF

Advanced Configuration Deep Dive

Scalability and High Availability

The distributed nature of Tempo means you can scale individual components based on load:

# In our distributed configuration
ingester:
  replicas: 3
  persistence:
    enabled: true
    size: 10Gi

distributor:
  replicas: 2
  autoscaling:
    enabled: true
    minReplicas: 2
    maxReplicas: 5

querier:
  replicas: 2

Resource Optimization

Tempo's microservices architecture allows precise resource allocation:

  • Ingesters: Memory-intensive — handle trace batching and flushing
  • Distributors: CPU-intensive — manage incoming trace ingestion
  • Queriers: Balanced — handle trace lookup operations
  • Compactors: I/O intensive — manage storage optimization

Monitoring Tempo Itself

A well-instrumented Tempo deployment monitors itself:

metaMonitoring:
  serviceMonitor:
    enabled: true
  grafanaAgent:
    enabled: true
    logs:
      remote:
        url: 'http://loki:3100/loki/api/v1/push'
    metrics:
      remote:
        url: 'http://mimir:9090/api/v1/push'

Real-World Benefits and Use Cases

Troubleshooting in Complex Environments

Imagine a scenario where users report intermittent payment failures. With Tempo:

Find the trace using the trace ID from your application logs Visualize the entire flow across payment service, fraud detection, database, and external payment gateway Identify the bottleneck — maybe the fraud service occasionally times out Correlate with metrics — use the generated metrics to see the error rate pattern

Cost Savings in Practice

Compared to commercial alternatives:

  • 90% reduction in storage costs by using object storage
  • No per-span pricing — ingest as much as you need
  • Reduced operational overhead — Kubernetes-native operations

Developer Experience Transformation

Developers can:

  • Self-service troubleshooting without needing production database access
  • Understand service dependencies through automatically generated service graphs
  • Performance profiling across the entire stack

Best Practices for Production

Data Retention Strategy

compactor:
  compaction:
    block_retention: 48h  # Adjust based on compliance and debugging needs
  • Debugging: 2–7 days typically sufficient
  • Compliance: Longer retention as required
  • Cost control: Balance retention with storage costs

Security Considerations

  • Use IRSA/IAM roles for cloud credentials
  • Enable TLS for OTLP endpoints in production
  • Configure network policies to restrict access
  • Use Tempo Enterprise for advanced multi-tenancy

Performance Tuning

  • Monitor ingester memory usage carefully
  • Adjust compaction windows based on trace volume
  • Configure appropriate S3 retry policies
  • Use dedicated object storage buckets for isolation

The Future with Tempo

As you scale, Tempo grows with you:

  • Multi-cluster deployments for global applications
  • Tempo Enterprise features for advanced analytics
  • Continuous profiling integration
  • AI/ML-assisted root cause analysis

Conclusion: Your Tracing Foundation

Grafana Tempo represents the evolution of distributed tracing — from a complex, expensive luxury to a simple, scalable necessity. Our production configuration demonstrates how to leverage Tempo's strengths while maintaining operational excellence.

The combination of open standards, cloud-native design, and seamless Grafana integration makes Tempo the ideal choice for organizations serious about observability.

Start tracing, stop guessing.

← back to all articles