KubernetesGrafana MimirPrometheus

Scaling Prometheus to Infinity: A Production-Ready Guide to Grafana Mimir on Kubernetes

Also mirrored on Medium.

Prometheus'un ölçeklendirilmesi/Grafana Mimir kavramını temsil eden kapak görseli

If you've ever managed a Prometheus stack at scale, you know the pain: the single-node architecture starts to creak under the weight of your metrics. You run into cardinality explosions, storage becomes a nightmare, and high availability feels like a distant dream.

What if you could have a Prometheus-compatible system that scales horizontally, is highly available, and seamlessly stores its data in a cloud object store like AWS S3?

Enter Grafana Mimir.

Mimir is the open-source, horizontally scalable, highly available, multi-tenant Prometheus-as-a-Service project. It lets you break free from the constraints of a single Prometheus server. In this guide, we'll deploy a production-style Mimir cluster on Kubernetes, using Helm and AWS S3 for durable, scalable storage.

What is Grafana Mimir?

Before we dive into the code, let's understand what Mimir is solving. At its core, Mimir takes the Prometheus protocol you know and love and bolts on a cloud-native architecture:

  • Horizontally Scalable: Every component (ingester, distributor, querier, etc.) can be scaled out independently.
  • Highly Available: No single point of failure. Mimir uses a gossip protocol for component discovery and can tolerate zone failures.
  • Durable Storage: Uses object storage (S3, GCS, Azure Blob) for long-term metric storage, which is cheaper and more reliable than local disks.
  • Multi-Tenant: Securely serves multiple independent teams or applications from a single cluster.

Our Target Architecture

We will deploy Mimir in a "microservices" mode, where each component runs as a separate service in our Kubernetes cluster. The key components we'll be using are:

  • Distributor: Accepts metrics via the Prometheus remote_write API.
  • Ingester: Batches metrics and writes them to long-term storage (S3).
  • Querier & Query-Frontend: Handles PromQL queries, potentially reading from both ingesters and long-term storage.
  • Store-Gateway: Retrieves series data from long-term storage (S3) for queries.
  • Compactor: Compresses and deduplicates blocks in long-term storage.
  • Ruler: Evaluates Prometheus recording and alerting rules.
  • Alertmanager: Handles alerts sent from the Ruler.

All of these will be fronted by a Gateway, which acts as a unified ingress point for all Mimir APIs.

Our storage backend will be Amazon S3, and we'll use a fine-tuned Helm chart for the deployment.

The Deployment: Code Walkthrough

We'll manage our deployment with three key files. Let's break them down.

1. The Installation Script (install.sh)

This simple script handles the Helm repository setup and the installation command.

#!/bin/bash
# install.sh

helm repo add grafana https://grafana.github.io/helm-charts
helm repo update

helm upgrade --install mimir grafana/mimir-distributed \
  --namespace mimir --create-namespace \
  --values mimir.yaml

This script is straightforward:

Adds the official Grafana Helm repository. Updates the local repository cache. Uses helm upgrade --install for an idempotent deployment—it will install Mimir if it doesn't exist, or upgrade it if it does. The configuration is driven by our central mimir.yaml file.

2. The Core Configuration (mimir.yaml)

This is the heart of our deployment. We're not using the default values; we're providing a focused, production-oriented configuration.

# mimir.yaml

# 1. Service Account with IAM Role for S3 Access (IRSA)
serviceAccount:
  create: true
  name: mimir-sa
  annotations:
    eks.amazonaws.com/role-arn: arn:aws:iam::<account_id>:role/mimir-irsa-role

# 2. Mimir Structured Configuration
mimir:
  structuredConfig:
    blocks_storage:
      backend: s3
      s3:
        endpoint: "s3.eu-west-1.amazonaws.com"
        region: "eu-west-1"
        bucket_name: medium-mimir-tsdb

    alertmanager_storage:
      backend: s3
      s3:
        endpoint: "s3.eu-west-1.amazonaws.com"
        region: "eu-west-1"
        bucket_name: medium-mimir-alertmanager

    ruler_storage:
      backend: s3
      s3:
        endpoint: "s3.eu-west-1.amazonaws.com"
        region: "eu-west-1"
        bucket_name: medium-mimir-ruler

# 3. Resource Requests & Limits
resources:
  requests:
    cpu: 200m
    memory: 512Mi
  limits:
    cpu: 1
    memory: 1Gi

# 4. Disabling Bundled Services
minio:  # We use external S3, not the built-in MinIO
  enabled: false

nginx:  # We use the modern 'gateway' component instead
  enabled: false

# 5. Enabling the Unified Gateway
gateway:
  enabled: true
  enabledNonEnterprise: true

Key Configuration Insights:

  • AWS IAM Roles for Service Accounts (IRSA): This is a critical security practice on EKS. Instead of storing AWS keys in a secret, we annotate the Mimir service account with an IAM Role ARN. The pods will assume this role to read/write to our S3 buckets. The role needs permissions for s3:GetObject, s3:PutObject, s3:ListBucket, etc., on the specified buckets.
  • Structured Config: This is the cleanest way to configure Mimir via the Helm chart. We directly specify the YAML that would go into Mimir's config file. Here, we point three different storage aspects (time-series blocks, Alertmanager state, and Ruler rules) to our S3 buckets.
  • Resource Management: We set sensible initial CPU and memory requests and limits. In a real production environment, you would closely monitor and adjust these based on actual usage.
  • Gateway over Nginx: The chart previously used an Nginx sidecar for routing. The new gateway component is a more integrated and efficient replacement, providing a single endpoint for all Mimir APIs.

3. Understanding the Helm Chart (mimir_values.yaml)

The provided mimir_values.yaml is a snapshot of the default values file for the mimir-distributed Helm chart. While we don't directly use it in our install.sh command, it's our reference manual. It shows the immense flexibility of the chart:

  • Component Toggling: You can enable or disable any Mimir microservice.
  • Advanced Scaling: Configures Horizontal Pod Autoscalers (HPA) and even KEDA autoscaling rules.
  • Caching Layers: Shows how to set up Memcached for chunks, indexes, and query results to boost performance.
  • Resource Templates: Defines the default resources, probes, and strategies for every component.

Our mimir.yaml file is a deliberate override of these extensive defaults, focusing only on what we need to change.

Deploying to Your Cluster

Ready to run it? The process is simple.

1. Prerequisites:

  • A running Kubernetes cluster (EKS recommended).
  • kubectl and helm configured.
  • AWS S3 buckets created with the correct names.
# Example S3 bucket creation via Terraform!

module "medium-mimir-tsdb-bucket" {
  source  = "terraform-aws-modules/s3-bucket/aws"
  version = "5.7.0"
  # sets the name of the S3 bucket
  bucket = "medium-mimir-tsdb"
  # Versioning keeps multiple versions of objects in the bucket, which is crucial for managing and tracking changes to Terraform state files.
  versioning = {
    enabled = true
  }
  # This configuration enables server-side encryption for the S3 bucket.
  # It uses the AWS Key Management Service (KMS) for encryption
  server_side_encryption_configuration = {
    rule = {
      apply_server_side_encryption_by_default = {
        kms_master_key_id = <s3_kms_key_arn>
        sse_algorithm      = "aws:kms"
      }
    }
  }
  # This block enables object locking for the S3 bucket.
  # Object locking ensures that objects stored in the bucket are immutable and helps protect against accidental deletion or modification.
  object_lock_configuration = {
    object_lock_enabled = "Enabled"
  }

  tags = merge(local.common_tags, { Name = "medium-mimir-tsdb" })
}
  • An IAM Role configured for IRSA with access to those buckets.
# Example Iam Role Service Account (IRSA) creation via Terraform!

resource "aws_iam_policy" "mimir_s3_policy" {
  name        = "MimirS3AccessPolicy"
  description = "Least-privilege access for Mimir to access S3 bucket"

  policy = jsonencode({
    Version = "2012-10-17",
    Statement = [
      {
        Effect = "Allow",
        Action = [
          "s3:GetObject",
          "s3:PutObject",
          "s3:DeleteObject"
        ],
        Resource = [
          "arn:aws:s3:::medium-mimir-tsdb/*",
          "arn:aws:s3:::medium-mimir-ruler/*",
          "arn:aws:s3:::medium-mimir-alertmanager/*"
        ]
      },
      {
        Effect = "Allow",
        Action = [
          "s3:ListBucket",
          "s3:GetBucketLocation"
        ],
        Resource = [
          "arn:aws:s3:::medium-mimir-tsdb",
          "arn:aws:s3:::medium-mimir-ruler",
          "arn:aws:s3:::medium-mimir-alertmanager"
        ]
      },
      {
        Effect = "Allow",
        Action = [
          "kms:Decrypt",
          "kms:GenerateDataKey"
        ],
        Resource = "arn:aws:kms:eu-west-1:${local.account_id}:key/<key_id>"
      }
    ]
  })
}

resource "aws_iam_role" "mimir_irsa_role" {
  name = "mimir-irsa-role"

  assume_role_policy = data.aws_iam_policy_document.mimir_irsa_trust_policy.json
}

data "aws_iam_policy_document" "mimir_irsa_trust_policy" {
  statement {
    effect = "Allow"

    principals {
      type        = "Federated"
      identifiers = [<oidc_provider_arn>]
    }

    actions = ["sts:AssumeRoleWithWebIdentity"]

    condition {
      test     = "StringEquals"
      variable = "oidc.eks.eu-west-1.amazonaws.com/id/YOUR_OIDC_PROVIDER_ID:sub"
      values   = ["system:serviceaccount:mimir:mimir-sa"]
    }

    # Add the audience condition
    condition {
      test     = "StringEquals"
      variable = "oidc.eks.eu-west-1.amazonaws.com/id/YOUR_OIDC_PROVIDER_ID:aud"
      values   = ["sts.amazonaws.com"]
    }
  }
}

resource "aws_iam_role_policy_attachment" "mimir_s3_attach" {
  role       = aws_iam_role.mimir_irsa_role.name
  policy_arn = aws_iam_policy.mimir_s3_policy.arn
}

2. Execute the deployment:

chmod +x install.sh
./install.sh
install.sh script'inin çalıştırılıp Mimir Helm chart'ının deploy edildiğini gösteren terminal çıktısı

3. Verify the installation:

kubectl -n mimir get pods

You should see pods for all the Mimir components initializing. Once they are all Running and ready, you can proceed.

Connecting Grafana and Prometheus

With Mimir running, it's time to point your clients at it.

For Grafana:

Add a new Prometheus data source in Grafana.

  • URL: http://mimir-gateway.mimir.svc.cluster.local:80/prometheus

(This is the internal Kubernetes service DNS name for the Gateway)

  • HTTP Header: Under "Custom HTTP Headers", add:
  • Header: X-Scope-OrgID
  • Value: demo (or your preferred tenant ID)

For Prometheus Remote Write:

Configure your Prometheus instances to remote write to Mimir.

# prometheus.yml
remote_write:
  - url: http://<GATEWAY_EXTERNAL_IP>/api/v1/push
    headers:
      X-Scope-OrgID: "demo"

You can expose the mimir-gateway service via a LoadBalancer or an Ingress controller to get an external IP.

Why This Setup is Production-Ready

Cloud-Native Storage: By using S3, we have durable, scalable, and cost-effective storage that decouples compute from data. Security-First: Using IRSA for AWS credentials is more secure than static keys and allows for fine-grained permissions. Scalability: The microservices architecture means you can scale the ingest path (distributors, ingesters) independently from the query path (queriers, store-gateways) based on load. Operational Simplicity: The official Helm chart encapsulates the complexity of managing dozens of interdependent microservices, making day-2 operations like upgrades and configuration changes much safer.

Happy scaling!

← back to all articles