KubernetesAWS ALBNGINX

Zero Downtime Deployments with Kubernetes Using ALB and NGINX

Also mirrored on Medium.

Kubernetes'te ALB ve NGINX kullanarak kesintisiz (zero downtime) deployment kavramını temsil eden kapak görseli

Introduction

Ensuring zero downtime deployments is crucial for businesses relying on continuous availability. However, during our Kubernetes deployments, we encountered a major issue: even after increasing replicas, downtime persisted whenever new pods were created and old ones were drained.

This problem was caused by how AWS Application Load Balancer (ALB) managed traffic during rolling updates. ALB did not immediately detect pod shutdowns, while Kubernetes had already stopped sending traffic to terminating pods, leading to a traffic mismatch and failed requests before new pods became available.

Our Solution?

We introduced NGINX Ingress Controller to manage internal traffic flow efficiently, ensuring a seamless transition between old and new pods and eliminating downtime. This guide will break down our approach step by step, explaining how Terraform provisions an ALB that perceives NGINX as a redirect target, ensuring smooth deployments.

The Problem We Faced

Initial Setup

Our deployment architecture included:

  • AWS Application Load Balancer (ALB) as the external entry point
  • Kubernetes services running microservices

Why Did Downtime Occur?

  • ALB did not immediately recognize new pods, causing traffic routing delays.
  • Old pods were removed too quickly, leaving a traffic gap.
  • ALB continued to route traffic to terminating pods, leading to failed requests.

Result: Whenever we deployed a new version, users experienced downtime because ALB couldn't properly handle pod transitions.

Implementing the Solution Step by Step with Terraform

1. Deploying a Service Account

This service account enables AWS ALB integration with Kubernetes, allowing the Load Balancer Controller to manage AWS Application Load Balancers (ALB).

# This resource creates service account for AWS Load Balancer Controller in kube-system namespace

resource "kubernetes_service_account" "service_account" {
  metadata {
    # Name should be "aws-load-balancer-controller" and namespace should be "kube-system"
    name      = "aws-load-balancer-controller"
    # The service account is deployed in the kube-system namespace, where cluster-wide services are typically managed.
    namespace = "kube-system"
    # Labels are key-value pairs associated with the Service Account. Name and component labels should be as below
    labels = {
      "app.kubernetes.io/name"      = "aws-load-balancer-controller"
      "app.kubernetes.io/component" = "controller"
    }
    # Annotations are key-value pairs that provide additional information about the Service Account.
    # We should add annotations related to Amazon EKS, including the IAM role ARN and regional endpoint settings.
    annotations = {
      "eks.amazonaws.com/role-arn"               = enter_your_eks_lb_role_arn
      "eks.amazonaws.com/sts-regional-endpoints" = "true"
    }
  }
}

Why Is This Resource Important? Grants Kubernetes permission to manage AWS ALBs Ensures secure access using IAM roles Required for deploying AWS Load Balancer Controller in EKS Helps Kubernetes interact with AWS securely

By properly configuring this service account, we ensure that the AWS Load Balancer Controller has the necessary permissions to operate securely and efficiently.

2. Deploying a Load Balancer Controller

The AWS Load Balancer Controller automatically manages Application Load Balancers (ALBs) for Kubernetes Ingress resources.

# This resource deploys aws-load-balancer-controller
resource "helm_release" "lb" {
  name       = "aws-load-balancer-controller"
  repository = "https://aws.github.io/eks-charts"
  chart      = "aws-load-balancer-controller"
  namespace  = "kube-system"

  # This is crucial because the AWS Load Balancer Controller needs IAM permissions via the service account to manage ALBs.
  depends_on = [
    kubernetes_service_account.service_account
  ]

  # Specifies the AWS region where the EKS cluster is running.
  set {
    name  = "region"
    value = enter_your_region
  }

  # The VPC ID associated with the EKS cluster. This is required because the ALB Controller needs access to the AWS networking infrastructure.
  set {
    name  = "vpcId"
    value = enter_your_vpcid
  }

  # Specifies the AWS Elastic Container Registry (ECR) URL for the Load Balancer Controller image.
  # https://docs.aws.amazon.com/eks/latest/userguide/add-ons-images.html
  set {
    name  = "image.repository"
    value = "enter_your_image_repository_url/amazon/aws-load-balancer-controller"
  }

  # This prevents Helm from creating a new service account.
  set {
    name  = "serviceAccount.create"
    value = "false"
  }

  # Assigns an existing IAM service account (aws-load-balancer-controller) that has the required permissions to interact with ALB and NLB.
  set {
    name  = "serviceAccount.name"
    value = "aws-load-balancer-controller"
  }

  # Specifies the EKS cluster name where the Load Balancer Controller will operate.
  set {
    name  = "clusterName"
    value = enter_your_eks_cluster_name
  }

  # Sets the logging level to debug for better troubleshooting and monitoring.
  # If you don't need verbose logs, change this to info or warn.
  set {
    name  = "log-level"
    value = "debug"
  }

  # Ensures that the ALB controller pods only run on nodes labeled with application=microservices
  set {
    name  = "nodeSelector.application"
    value = "microservices"
  }
}

Why Is This Resource Important? Automatically provisions ALBs for Kubernetes Ingress resources Ensures external traffic is handled efficiently Supports both ALB, allowing flexible load balancing Ensures seamless deployment with Terraform & Helm

By deploying AWS Load Balancer Controller using Terraform, we achieve scalability, reliability, and automated ingress management for Kubernetes applications!

3. Deploying a NGINX Ingress Controller

This controller manages internal traffic routing by acting as a reverse proxy and load balancer.

# Deploys the NGINX Ingress Controller, which is responsible for managing ingress traffic (HTTP/HTTPS) in the cluster.

resource "helm_release" "nginx" {
  name       = "nginx-ingress-controller"
  repository = "https://kubernetes.github.io/ingress-nginx"
  chart      = "ingress-nginx"
  namespace  = "kube-system"
  version    = "4.12.0"
  depends_on = [
    kubernetes_service_account.service_account
  ]

  # Uses a ClusterIP service type, keeping NGINX traffic internal to the cluster.
  set {
    name  = "controller.service.type"
    value = "ClusterIP"
  }

  # Runs as a DaemonSet, ensuring availability on every node. This is useful for handling high availability and scaling traffic efficiently.
  set {
    name  = "controller.kind"
    value = "DaemonSet"
  }
}

Why Is This Resource Important? Manages internal traffic routing efficiently Prevents ALB from routing traffic to terminating pods Ensures high availability by running as a DaemonSet Improves scalability across multiple services

By deploying NGINX Ingress Controller using Terraform, we ensure that Kubernetes services have a highly available and scalable traffic routing mechanism.

4. Deploying a NGINX Ingress ALB

This Nginx Ingress ALB enables AWS Application Load Balancer (ALB) to route traffic through NGINX Ingress Controller to multiple backend Kubernetes services. This setup ensures that traffic is properly forwarded based on the request path, allowing seamless communication between external users and internal microservices.

resource "kubernetes_ingress_v1" "microservices_alb_ingress_nginx" {
  metadata {
    name      = "microservices-alb-ingress-nginx"
    namespace = "microservices"
  }

  # Ensures that ALB is fully deployed before applying the Ingress resource.
  depends_on = [helm_release.lb]

  # Tells Terraform to wait until the ALB is fully provisioned before proceeding.
  wait_for_load_balancer = true

  spec {
    # Specifies that this Ingress resource will be handled by the NGINX Ingress Controller (not AWS ALB directly).
    ingress_class_name = "nginx"

    # This Ingress resource defines three separate rules, each forwarding traffic to a different Kubernetes service based on the request path.
    # This rule forwards request for /first and /first/* path to port 80 of first kubernetes service.
    # Any request with the path /first/ (e.g., https://example.com/first/) will be forwarded to first-service on port 80.
    rule {
      http {
        path {
          backend {
            service {
              name = "first-service"
              port {
                number = 80
              }
            }
          }
          path      = "/first/"
          path_type = "Prefix"
        }
      }
    }
    # This rule forwards request for /second and /second/* path to port 80 of second kubernetes service
    rule {
      http {
        path {
          backend {
            service {
              name = "second-service"
              port {
                number = 80
              }
            }
          }
          path      = "/second/"
          path_type = "Prefix"
        }
      }
    }
    # This rule forwards request for /third and /third/* path to port 80 of third kubernetes service
    rule {
      http {
        path {
          backend {
            service {
              name = "third-service"
              port {
                number = 80
              }
            }
          }
          path      = "/third/"
          path_type = "Prefix"
        }
      }
    }
  }
}

Why Is This Resource Important? Enables external access to multiple microservices using a single ALB Efficiently routes traffic based on URL path Ensures high availability by waiting for the ALB to be provisioned Seamlessly integrates with NGINX Ingress Controller for internal routing

By defining Ingress rules with NGINX, we ensure that AWS ALB efficiently forwards requests to the correct backend service, providing scalability, flexibility, and zero downtime for microservices!

5. Deploying an Ingress ALB

This Ingress ALB enables AWS Application Load Balancer (ALB) to manage external traffic. This Ingress is configured to route all traffic to NGINX Ingress Controller, which then forwards it to different microservices within the Kubernetes cluster.

resource "kubernetes_ingress_v1" "microservices-alb-ingress" {
  metadata {
    name      = "microservices-alb-ingress"
    # Deploys the Ingress resource in the kube-system namespace, ensuring it handles global ingress traffic at the cluster level.
    namespace = "kube-system"
    # This annotations provide configurations for ALB.
    # https://kubernetes-sigs.github.io/aws-load-balancer-controller/v2.6/guide/ingress/annotations/
    annotations = {
      "alb.ingress.kubernetes.io/group.name" = "microservices-apigw-alb"
      # Specifies that the ALB is internal-only, meaning it won't be exposed to the public internet.
      "alb.ingress.kubernetes.io/scheme" = "internal"
      # Configures the ALB to listen for HTTP (port 80) and HTTPS (port 443) traffic.
      "alb.ingress.kubernetes.io/listen-ports" = "[{\"HTTP\": 80}, {\"HTTPS\": 443}]"
      # Associates an AWS ACM SSL certificate with the ALB for handling HTTPS traffic.
      "alb.ingress.kubernetes.io/certificate-arn" = enter_your_acm_certificate_arn
      # Routes traffic directly to pods using IP-based routing instead of instance-based routing.
      "alb.ingress.kubernetes.io/target-type" = "ip"
      # ALB uses this path to perform health checks on backend services.
      "alb.ingress.kubernetes.io/healthcheck-path" = "/healthz"
      # Configures ALB to perform a health check every 10 seconds.
      "alb.ingress.kubernetes.io/healthcheck-interval-seconds" = "10"
      "alb.ingress.kubernetes.io/healthcheck-timeout-seconds"  = "2"
      # ALB will consider a service healthy after 2 successful checks and unhealthy after 2 failures.
      "alb.ingress.kubernetes.io/healthy-threshold-count"   = "2"
      "alb.ingress.kubernetes.io/unhealthy-threshold-count" = "2"
      # ALB expects a 200 HTTP response from the /healthz endpoint.
      "alb.ingress.kubernetes.io/success-codes" = "200"
      "alb.ingress.kubernetes.io/tags"          = enter_your_alb_ingress_kubernetes_io_tags
      "alb.ingress.kubernetes.io/group.order"   = 1
    }
  }
  depends_on = [helm_release.lb]
  # This line indicates that Terraform should wait for the ALB to be fully created and ready before proceeding.
  wait_for_load_balancer = true
  spec {
    # Specifies that this Ingress resource should be managed by AWS ALB, instead of NGINX or another ingress controller.
    ingress_class_name = "alb"

    # Any request to / (e.g., https://example.com/) is forwarded to nginx-ingress-controller.
    # NGINX Ingress Controller then routes traffic to appropriate microservices based on internal Ingress rules.
    rule {
      http {
        path {
          backend {
            service {
              name = "nginx-ingress-controller-ingress-nginx-controller"
              port {
                number = 80
              }
            }
          }

          path      = "/"
          path_type = "Prefix"
        }
      }
    }
  }
}

Why Is This Resource Important? Allows AWS ALB to manage ingress traffic for Kubernetes workloads Provides secure traffic routing using ACM SSL certificates Performs intelligent health checks to ensure high availability Integrates seamlessly with NGINX for internal routing

By deploying this ALB Ingress resource, we enable scalable, highly available, and secure external access to Kubernetes services while maintaining zero downtime deployments.

Detailed Explanation of Deployment yaml for Kustomize

This Kubernetes Deployment YAML defines the first microservice (first-deployment) inside the microservices namespace. It uses Kustomize for managing configurations and ensures zero downtime deployments by implementing rolling update strategies, health probes, and graceful shutdown mechanisms.

  • Metadata Configuration
  • Replica Strategy & Rolling Updates
  • Pod Selector & Template
  • Dapr Sidecar Annotations for Observability
  • Node Selection & IAM Authentication
  • Secret Management with AWS Secrets Manager
  • Graceful Pod Shutdown & Cleanup
  • Container Configuration & Probes
  • Health Probes for Smooth Traffic Management
  • Graceful Shutdown with PreStop Hook

1. Metadata Configuration

apiVersion: apps/v1
kind: Deployment
metadata:
  name: first-deployment
  namespace: microservices
  labels:
    app: first
  • apiVersion: apps/v1 → Uses Kubernetes API for deploying applications.
  • kind: Deployment → Specifies that this resource is a Deployment, which manages pod updates.
  • metadata.name: first-deployment → Assigns the deployment the name first-deployment.
  • metadata.namespace: microservices → Deploys it in the microservices namespace, isolating it from other workloads.
  • labels.app: first → Labels the deployment for service discovery and Ingress routing.

2. Replica Strategy & Rolling Updates

spec:
  replicas: 3
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxUnavailable: 0  # Ensures no pods go down during updates
      maxSurge: 2        # Allows two extra pods during rollout
  revisionHistoryLimit: 1

replicas: 3 → Maintains three replicas of the pod for high availability.

  • strategy.type: RollingUpdate → Uses rolling updates instead of Recreate, which prevents downtime.
  • rollingUpdate.maxUnavailable: 0 → Ensures that all running pods stay available while new ones are deployed.
  • rollingUpdate.maxSurge: 2 → Allows up to two extra pods to be created during the rollout for smooth transitions.
  • revisionHistoryLimit: 1 → Keeps only one previous revision to reduce storage usage in the cluster.

Zero Downtime Impact:

  • Ensures no downtime by keeping old pods running until new pods are healthy.
  • The system gradually replaces old pods with new ones instead of restarting everything at once.

3. Pod Selector & Template

selector:
  matchLabels:
    app: first
template:
  metadata:
    labels:
      app: first
  • selector.matchLabels.app: first → Ensures that the Deployment manages pods labeled app: first.
  • template.metadata.labels.app: first → Assigns the same label inside the pod template for consistency.

Zero Downtime Impact:

  • Ensures proper pod replacement by keeping labels consistent.
  • Kubernetes can correctly map traffic to new pods during deployments.

4. Dapr Sidecar Annotations for Observability

annotations:
  dapr.io/enabled: "true"
  dapr.io/app-id: first
  dapr.io/app-port: "80"
  dapr.io/log-as-json: "true"
  dapr.io/log-level: "info"
  dapr.io/sidecar-cpu-limit: "50m"
  dapr.io/sidecar-memory-limit: "128Mi"
  dapr.io/sidecar-cpu-request: "5m"
  dapr.io/sidecar-memory-request: "64Mi"
  dapr.io/config: "tracing-config"
  dapr.io/trace-sampling-rate: "1"
  • Enables Dapr Sidecar to add distributed tracing, logging, and observability.
  • Limits resource usage for the sidecar process to optimize performance.

Zero Downtime Impact:

  • Improves logging and monitoring to detect issues during deployments.
  • Ensures minimal sidecar resource usage to prevent overloading the node.

5. Node Selection & IAM Authentication

spec:
  nodeSelector:
    app:
  serviceAccountName: serviceaccount

nodeSelector.app: → Restricts pods to specific nodes based on the app label.

  • serviceAccountName: serviceaccount → Uses a predefined IAM Service Account for secure AWS access.

Zero Downtime Impact:

  • Ensures pods run on dedicated nodes, avoiding congestion.
  • Secure authentication reduces risk of failed deployments due to permission issues.

6. Secret Management with AWS Secrets Manager

volumes:
  - name: secrets-store-inline
    csi:
      driver: secrets-store.csi.k8s.io
      readOnly: true
      volumeAttributes:
        secretProviderClass: <name>-aws-secrets
  • Uses Kubernetes Secrets Store CSI Driver to fetch secrets dynamically from AWS Secrets Manager.

Zero Downtime Impact:

  • Ensures secure key rotation without pod restarts.
  • Avoids hardcoding secrets in deployment files.

7. Graceful Pod Shutdown & Cleanup

terminationGracePeriodSeconds: 120
  • terminationGracePeriodSeconds: 120 → Pods wait 120 seconds before shutting down, allowing in-flight requests to complete.

Zero Downtime Impact:

  • Prevents abrupt pod termination, avoiding dropped connections.
  • Gives NGINX and ALB enough time to remove the pod from the load balancer.

8. Container Configuration & Probes

containers:
  - name: first-deployment
    image: ecr_url:latest
    resources:
      requests:
        memory: 1Gi
      limits:
        memory: 2Gi
  • Defines memory requests and limits to prevent resource starvation.
  • Uses a prebuilt image (ecr_url:latest), ensuring the latest version is deployed.

9. Health Probes for Smooth Traffic Management

startupProbe:
  httpGet:
    path: /healthz
    port: 80
  failureThreshold: 60
  periodSeconds: 5
  • startupProbe → Waits for the application to fully initialize before receiving traffic.
  • Prevents ALB and Kubernetes from routing traffic to unready pods.
readinessProbe:
  httpGet:
    path: /healthz
    port: 80
  initialDelaySeconds: 40
  periodSeconds: 5
  successThreshold: 1
  failureThreshold: 10
  • readinessProbe → Ensures traffic is only sent to fully ready pods.
  • Traffic is stopped for failing pods instead of removing them.
livenessProbe:
  httpGet:
    path: /healthz
    port: 80
  initialDelaySeconds: 10
  periodSeconds: 5
  failureThreshold: 3
  • livenessProbe → Restarts unresponsive pods automatically.

Zero Downtime Impact:

  • Ensures new pods are not overwhelmed with traffic before they are ready.
  • Prevents requests from hitting dead pods, improving reliability.

10. Graceful Shutdown with PreStop Hook

lifecycle:
  preStop:
    exec:
      command: ["/bin/sh", "-c", "sleep 60"]  # Graceful shutdown delay
  • PreStop Hook → Adds a 60-second delay before the pod shuts down, allowing existing connections to finish.

Zero Downtime Impact:

  • Gives ALB and Kubernetes enough time to remove the pod from routing tables.
  • Ensures users do not experience failed requests during rolling updates.

How We Solved the Problem

We introduced NGINX Ingress Controller to act as an intermediary between ALB and Kubernetes services. Instead of sending traffic directly to the backend services, ALB now routes requests to NGINX, which intelligently manages internal traffic.

How NGINX Fixed the Issue

  • NGINX ensures that only healthy, ready pods receive traffic.
  • Prevents ALB from sending traffic to draining pods, reducing failed requests.
  • Traffic transitions smoothly between old and new pods, ensuring seamless updates.

Understanding the Final Architecture

ALB → NGINX → Kubernetes servisleri şeklindeki nihai trafik akış mimarisini gösteren diyagram

Traffic Flow: ALB → NGINX → Kubernetes Services

1. ALB handles external traffic and forwards it to NGINX. 2. NGINX intelligently routes traffic to the correct microservice. 3. Kubernetes readiness probes ensure new pods are fully functional before receiving traffic. 4. Pod termination hooks delay shutdown, allowing ALB time to remove old pods safely.

Result: Smooth, zero downtime deployments!

Final Thoughts: Key Takeaways

  • ALB alone is not enough — NGINX provides better control over traffic flow.
  • NGINX prevents ALB from sending traffic to terminating pods, reducing downtime.
  • Readiness probes ensure new pods are ready before receiving traffic.
  • PreStop hooks and termination grace periods delay shutdown, preventing traffic gaps.
  • Rolling updates with maxUnavailable: 0 prevent downtime.
  • PreStop hook & termination grace period allow smooth shutdown.
  • Health probes ensure only healthy pods serve traffic.
  • AWS Secrets Manager integration avoids restarts for secret updates.
  • Dapr sidecar improves tracing & logging for debugging.

By integrating NGINX with ALB, we eliminated downtime during rolling updates, ensuring 100% availability for our services.

This approach has transformed our deployment strategy, providing a scalable, resilient, and high-availability Kubernetes environment.

← back to all articles