
The Modern Observability Trinity: Metrics, Logs, Traces, and Telemetry Collection
In today's cloud-native world, monitoring has evolved from collecting simple metrics to achieving full-stack observability. The combination of Grafana for visualization, Mimir for metrics, Loki for logs, Tempo for traces, and now Alloy for telemetry collection represents the next generation of Kubernetes monitoring. In this comprehensive guide, I'll walk you through deploying a production-ready observability stack that provides unparalleled insights into your applications.
Why This Stack Matters
Before we dive into the technical details, let's understand why this particular combination is so powerful:
- Grafana: The universal dashboard platform that ties everything together
- Mimir: Horizontally scalable, highly available Prometheus-as-a-Service
- Loki: Log aggregation system inspired by Prometheus
- Tempo: Highly scalable, cost-effective distributed tracing
- Alloy: The next-generation telemetry collector (successor to Grafana Agent)
Together, they form a complete observability platform that's both powerful and cost-effective.
Prerequisites
- Kubernetes cluster (EKS, AKS, GKE, or on-prem)
- Helm installed
- kubectl configured
- Basic understanding of Kubernetes concepts
Deployment Architecture
Let's first understand what we're building:

Version 1: Professional Architecture Diagram

Version 2: Detailed Component Architecture

Version 3: Network-Focused Diagram
Step-by-Step Deployment
1. Understanding the Grafana Configuration
Let me break down our configuration file:
grafana.yaml
# grafana.yaml
adminUser: enter_your_username
adminPassword: enter_your_password
datasources:
datasources.yaml:
apiVersion: 1
datasources:
- name: Loki
type: loki
access: proxy
uid: loki
jsonData:
maxLines: 1000
healthCheck:
enabled: false
url: http://loki-gateway.loki.svc.cluster.local
isDefault: false
- name: Tempo
type: tempo
access: proxy
uid: tempo
url: http://tempo-query-frontend.tempo.svc.cluster.local:3200
jsonData:
tracesToLogsV2:
datasourceUid: 'loki'
spanStartTimeShift: '-5m'
spanEndTimeShift: '5m'
serviceMap:
datasourceUid: 'mimir'
- name: Mimir
type: prometheus
uid: mimir
access: proxy
url: http://mimir-gateway.mimir.svc.cluster.local/prometheus
resources:
requests:
cpu: 100m
memory: 256Mi
limits:
cpu: 500m
memory: 512Mi
persistence:
type: pvc
enabled: true
size: 10Gi
ingress:
enabled: false
2. Configuration Deep Dive
Datasources Configuration
Loki Configuration:
- name: Loki
type: loki
url: http://loki-gateway.loki.svc.cluster.local
- Purpose: Centralized log aggregation
- URL Pattern: Uses Kubernetes service DNS
service.namespace.svc.cluster.local - maxLines: 1000 lines per log query for performance
- Health Check Disabled: Often needed when behind gateways
Tempo Configuration:
- name: Tempo
type: tempo
url: http://tempo-query-frontend.tempo.svc.cluster.local:3200
jsonData:
tracesToLogsV2:
datasourceUid: 'loki'
spanStartTimeShift: '-5m'
spanEndTimeShift: '5m'
serviceMap:
datasourceUid: 'mimir'
- Traces to Logs: Automatically links traces with relevant logs
- Time Shifting: Expands time range for context (±5 minutes)
- Service Map: Generates service dependency graphs using Mimir metrics
Mimir Configuration:
- name: Mimir
type: prometheus
url: http://mimir-gateway.mimir.svc.cluster.local/prometheus
- Prometheus-Compatible: Mimir exposes the Prometheus API
- Scalable Backend: Handles millions of series efficiently
- High Availability: Built-in replication and fault tolerance
Resource Management
resources:
requests:
cpu: 100m
memory: 256Mi
limits:
cpu: 500m
memory: 512Mi
- Requests: Guaranteed resources for stable operation
- Limits: Prevents Grafana from consuming excessive resources
- Balanced: Suitable for medium-sized deployments
Persistence
persistence:
type: pvc
enabled: true
size: 10Gi
- PVC: Persistent Volume Claim for data durability
- 10GB: Ample space for dashboards, users, and preferences
- Stateful: Survives pod restarts and upgrades
3. Installation Script Breakdown
Our install.sh script handles the deployment:
#!/bin/bash
# Add Grafana Helm repository
helm repo add grafana https://grafana.github.io/helm-charts
helm repo update
# Deploy Grafana with our custom configuration
helm upgrade --install grafana grafana/grafana --namespace monitoring --values grafana.yaml
What this script does:
Adds the official Grafana Helm repository Updates local repository cache Deploys Grafana using our customized values Uses upgrade --install for idempotent deployment
4. Deployment Execution
Let's deploy our observability stack:
# Create monitoring namespace
kubectl create namespace monitoring
# Make script executable
chmod +x install.sh
# Run installation
./install.sh

5. Verifying the Deployment
Check if everything was deployed correctly:
# Check Grafana pod status
kubectl get pods -n monitoring -l app.kubernetes.io/name=grafana
# Check persistent volume claim
kubectl get pvc -n monitoring
# Check service
kubectl get svc -n monitoring
# Check Grafana logs
kubectl logs -n monitoring -l app.kubernetes.io/name=grafana
6. Accessing Grafana
Since we've disabled ingress, let's use port forwarding:
# Port forward to access Grafana
kubectl port-forward -n monitoring svc/grafana 3000:80
Now access Grafana at http://localhost:3000 with credentials:
- Username: enter_your_username
- Password: enter_your_password
7. Verifying Data Sources
Once logged in, navigate to Configuration → Data Sources to verify all three data sources are connected:
- Mimir should show green "Health" status
- Loki should be connected (health check might show a warning due to the gateway)
- Tempo should be ready for trace queries
Conclusion
You've now deployed a complete observability stack that provides:
- Unified Visualization through Grafana
- Scalable Metrics with Mimir
- Centralized Logging with Loki
- Distributed Tracing with Tempo
- Modern Telemetry with Grafana Alloy
This setup transforms your Kubernetes monitoring from reactive troubleshooting to proactive observability. You can now:
- Detect issues before they impact users
- Understand system behavior across microservices
- Reduce mean time to resolution (MTTR)
- Make data-driven decisions about scaling and optimization