Introduction
Ensuring zero downtime deployments is crucial for businesses relying on continuous availability. However, during our Kubernetes deployments, we encountered a major issue: even after increasing replicas, downtime persisted whenever new pods were created and old ones were drained.
This problem was caused by how AWS Application Load Balancer (ALB) managed traffic during rolling updates. ALB did not immediately detect pod shutdowns, while Kubernetes had already stopped sending traffic to terminating pods, leading to a traffic mismatch and failed requests before new pods became available.
Our Solution?
We introduced NGINX Ingress Controller to manage internal traffic flow efficiently, ensuring a seamless transition between old and new pods and eliminating downtime. This guide will break down our approach step by step, explaining how Terraform provisions an ALB that perceives NGINX as a redirect target, ensuring smooth deployments.
The Problem We Faced
Initial Setup
Our deployment architecture included:
- AWS Application Load Balancer (ALB) as the external entry point
- Kubernetes services running microservices
Why Did Downtime Occur?
- ALB did not immediately recognize new pods, causing traffic routing delays.
- Old pods were removed too quickly, leaving a traffic gap.
- ALB continued to route traffic to terminating pods, leading to failed requests.
Result: Whenever we deployed a new version, users experienced downtime because ALB couldn't properly handle pod transitions.
Implementing the Solution Step by Step with Terraform
1. Deploying a Service Account
This service account enables AWS ALB integration with Kubernetes, allowing the Load Balancer Controller to manage AWS Application Load Balancers (ALB).
# This resource creates service account for AWS Load Balancer Controller in kube-system namespace
resource "kubernetes_service_account" "service_account" {
metadata {
# Name should be "aws-load-balancer-controller" and namespace should be "kube-system"
name = "aws-load-balancer-controller"
# The service account is deployed in the kube-system namespace, where cluster-wide services are typically managed.
namespace = "kube-system"
# Labels are key-value pairs associated with the Service Account. Name and component labels should be as below
labels = {
"app.kubernetes.io/name" = "aws-load-balancer-controller"
"app.kubernetes.io/component" = "controller"
}
# Annotations are key-value pairs that provide additional information about the Service Account.
# We should add annotations related to Amazon EKS, including the IAM role ARN and regional endpoint settings.
annotations = {
"eks.amazonaws.com/role-arn" = enter_your_eks_lb_role_arn
"eks.amazonaws.com/sts-regional-endpoints" = "true"
}
}
}
Why Is This Resource Important? Grants Kubernetes permission to manage AWS ALBs Ensures secure access using IAM roles Required for deploying AWS Load Balancer Controller in EKS Helps Kubernetes interact with AWS securely
By properly configuring this service account, we ensure that the AWS Load Balancer Controller has the necessary permissions to operate securely and efficiently.
2. Deploying a Load Balancer Controller
The AWS Load Balancer Controller automatically manages Application Load Balancers (ALBs) for Kubernetes Ingress resources.
# This resource deploys aws-load-balancer-controller
resource "helm_release" "lb" {
name = "aws-load-balancer-controller"
repository = "https://aws.github.io/eks-charts"
chart = "aws-load-balancer-controller"
namespace = "kube-system"
# This is crucial because the AWS Load Balancer Controller needs IAM permissions via the service account to manage ALBs.
depends_on = [
kubernetes_service_account.service_account
]
# Specifies the AWS region where the EKS cluster is running.
set {
name = "region"
value = enter_your_region
}
# The VPC ID associated with the EKS cluster. This is required because the ALB Controller needs access to the AWS networking infrastructure.
set {
name = "vpcId"
value = enter_your_vpcid
}
# Specifies the AWS Elastic Container Registry (ECR) URL for the Load Balancer Controller image.
# https://docs.aws.amazon.com/eks/latest/userguide/add-ons-images.html
set {
name = "image.repository"
value = "enter_your_image_repository_url/amazon/aws-load-balancer-controller"
}
# This prevents Helm from creating a new service account.
set {
name = "serviceAccount.create"
value = "false"
}
# Assigns an existing IAM service account (aws-load-balancer-controller) that has the required permissions to interact with ALB and NLB.
set {
name = "serviceAccount.name"
value = "aws-load-balancer-controller"
}
# Specifies the EKS cluster name where the Load Balancer Controller will operate.
set {
name = "clusterName"
value = enter_your_eks_cluster_name
}
# Sets the logging level to debug for better troubleshooting and monitoring.
# If you don't need verbose logs, change this to info or warn.
set {
name = "log-level"
value = "debug"
}
# Ensures that the ALB controller pods only run on nodes labeled with application=microservices
set {
name = "nodeSelector.application"
value = "microservices"
}
}
Why Is This Resource Important? Automatically provisions ALBs for Kubernetes Ingress resources Ensures external traffic is handled efficiently Supports both ALB, allowing flexible load balancing Ensures seamless deployment with Terraform & Helm
By deploying AWS Load Balancer Controller using Terraform, we achieve scalability, reliability, and automated ingress management for Kubernetes applications!
3. Deploying a NGINX Ingress Controller
This controller manages internal traffic routing by acting as a reverse proxy and load balancer.
# Deploys the NGINX Ingress Controller, which is responsible for managing ingress traffic (HTTP/HTTPS) in the cluster.
resource "helm_release" "nginx" {
name = "nginx-ingress-controller"
repository = "https://kubernetes.github.io/ingress-nginx"
chart = "ingress-nginx"
namespace = "kube-system"
version = "4.12.0"
depends_on = [
kubernetes_service_account.service_account
]
# Uses a ClusterIP service type, keeping NGINX traffic internal to the cluster.
set {
name = "controller.service.type"
value = "ClusterIP"
}
# Runs as a DaemonSet, ensuring availability on every node. This is useful for handling high availability and scaling traffic efficiently.
set {
name = "controller.kind"
value = "DaemonSet"
}
}
Why Is This Resource Important? Manages internal traffic routing efficiently Prevents ALB from routing traffic to terminating pods Ensures high availability by running as a DaemonSet Improves scalability across multiple services
By deploying NGINX Ingress Controller using Terraform, we ensure that Kubernetes services have a highly available and scalable traffic routing mechanism.
4. Deploying a NGINX Ingress ALB
This Nginx Ingress ALB enables AWS Application Load Balancer (ALB) to route traffic through NGINX Ingress Controller to multiple backend Kubernetes services. This setup ensures that traffic is properly forwarded based on the request path, allowing seamless communication between external users and internal microservices.
resource "kubernetes_ingress_v1" "microservices_alb_ingress_nginx" {
metadata {
name = "microservices-alb-ingress-nginx"
namespace = "microservices"
}
# Ensures that ALB is fully deployed before applying the Ingress resource.
depends_on = [helm_release.lb]
# Tells Terraform to wait until the ALB is fully provisioned before proceeding.
wait_for_load_balancer = true
spec {
# Specifies that this Ingress resource will be handled by the NGINX Ingress Controller (not AWS ALB directly).
ingress_class_name = "nginx"
# This Ingress resource defines three separate rules, each forwarding traffic to a different Kubernetes service based on the request path.
# This rule forwards request for /first and /first/* path to port 80 of first kubernetes service.
# Any request with the path /first/ (e.g., https://example.com/first/) will be forwarded to first-service on port 80.
rule {
http {
path {
backend {
service {
name = "first-service"
port {
number = 80
}
}
}
path = "/first/"
path_type = "Prefix"
}
}
}
# This rule forwards request for /second and /second/* path to port 80 of second kubernetes service
rule {
http {
path {
backend {
service {
name = "second-service"
port {
number = 80
}
}
}
path = "/second/"
path_type = "Prefix"
}
}
}
# This rule forwards request for /third and /third/* path to port 80 of third kubernetes service
rule {
http {
path {
backend {
service {
name = "third-service"
port {
number = 80
}
}
}
path = "/third/"
path_type = "Prefix"
}
}
}
}
}
Why Is This Resource Important? Enables external access to multiple microservices using a single ALB Efficiently routes traffic based on URL path Ensures high availability by waiting for the ALB to be provisioned Seamlessly integrates with NGINX Ingress Controller for internal routing
By defining Ingress rules with NGINX, we ensure that AWS ALB efficiently forwards requests to the correct backend service, providing scalability, flexibility, and zero downtime for microservices!
5. Deploying an Ingress ALB
This Ingress ALB enables AWS Application Load Balancer (ALB) to manage external traffic. This Ingress is configured to route all traffic to NGINX Ingress Controller, which then forwards it to different microservices within the Kubernetes cluster.
resource "kubernetes_ingress_v1" "microservices-alb-ingress" {
metadata {
name = "microservices-alb-ingress"
# Deploys the Ingress resource in the kube-system namespace, ensuring it handles global ingress traffic at the cluster level.
namespace = "kube-system"
# This annotations provide configurations for ALB.
# https://kubernetes-sigs.github.io/aws-load-balancer-controller/v2.6/guide/ingress/annotations/
annotations = {
"alb.ingress.kubernetes.io/group.name" = "microservices-apigw-alb"
# Specifies that the ALB is internal-only, meaning it won't be exposed to the public internet.
"alb.ingress.kubernetes.io/scheme" = "internal"
# Configures the ALB to listen for HTTP (port 80) and HTTPS (port 443) traffic.
"alb.ingress.kubernetes.io/listen-ports" = "[{\"HTTP\": 80}, {\"HTTPS\": 443}]"
# Associates an AWS ACM SSL certificate with the ALB for handling HTTPS traffic.
"alb.ingress.kubernetes.io/certificate-arn" = enter_your_acm_certificate_arn
# Routes traffic directly to pods using IP-based routing instead of instance-based routing.
"alb.ingress.kubernetes.io/target-type" = "ip"
# ALB uses this path to perform health checks on backend services.
"alb.ingress.kubernetes.io/healthcheck-path" = "/healthz"
# Configures ALB to perform a health check every 10 seconds.
"alb.ingress.kubernetes.io/healthcheck-interval-seconds" = "10"
"alb.ingress.kubernetes.io/healthcheck-timeout-seconds" = "2"
# ALB will consider a service healthy after 2 successful checks and unhealthy after 2 failures.
"alb.ingress.kubernetes.io/healthy-threshold-count" = "2"
"alb.ingress.kubernetes.io/unhealthy-threshold-count" = "2"
# ALB expects a 200 HTTP response from the /healthz endpoint.
"alb.ingress.kubernetes.io/success-codes" = "200"
"alb.ingress.kubernetes.io/tags" = enter_your_alb_ingress_kubernetes_io_tags
"alb.ingress.kubernetes.io/group.order" = 1
}
}
depends_on = [helm_release.lb]
# This line indicates that Terraform should wait for the ALB to be fully created and ready before proceeding.
wait_for_load_balancer = true
spec {
# Specifies that this Ingress resource should be managed by AWS ALB, instead of NGINX or another ingress controller.
ingress_class_name = "alb"
# Any request to / (e.g., https://example.com/) is forwarded to nginx-ingress-controller.
# NGINX Ingress Controller then routes traffic to appropriate microservices based on internal Ingress rules.
rule {
http {
path {
backend {
service {
name = "nginx-ingress-controller-ingress-nginx-controller"
port {
number = 80
}
}
}
path = "/"
path_type = "Prefix"
}
}
}
}
}
Why Is This Resource Important? Allows AWS ALB to manage ingress traffic for Kubernetes workloads Provides secure traffic routing using ACM SSL certificates Performs intelligent health checks to ensure high availability Integrates seamlessly with NGINX for internal routing
By deploying this ALB Ingress resource, we enable scalable, highly available, and secure external access to Kubernetes services while maintaining zero downtime deployments.
Detailed Explanation of Deployment yaml for Kustomize
This Kubernetes Deployment YAML defines the first microservice (first-deployment) inside the microservices namespace. It uses Kustomize for managing configurations and ensures zero downtime deployments by implementing rolling update strategies, health probes, and graceful shutdown mechanisms.
- Metadata Configuration
- Replica Strategy & Rolling Updates
- Pod Selector & Template
- Dapr Sidecar Annotations for Observability
- Node Selection & IAM Authentication
- Secret Management with AWS Secrets Manager
- Graceful Pod Shutdown & Cleanup
- Container Configuration & Probes
- Health Probes for Smooth Traffic Management
- Graceful Shutdown with PreStop Hook
1. Metadata Configuration
apiVersion: apps/v1
kind: Deployment
metadata:
name: first-deployment
namespace: microservices
labels:
app: first
- apiVersion: apps/v1 → Uses Kubernetes API for deploying applications.
- kind: Deployment → Specifies that this resource is a Deployment, which manages pod updates.
- metadata.name: first-deployment → Assigns the deployment the name first-deployment.
- metadata.namespace: microservices → Deploys it in the microservices namespace, isolating it from other workloads.
- labels.app: first → Labels the deployment for service discovery and Ingress routing.
2. Replica Strategy & Rolling Updates
spec:
replicas: 3
strategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 0 # Ensures no pods go down during updates
maxSurge: 2 # Allows two extra pods during rollout
revisionHistoryLimit: 1
replicas: 3 → Maintains three replicas of the pod for high availability.
- strategy.type: RollingUpdate → Uses rolling updates instead of Recreate, which prevents downtime.
- rollingUpdate.maxUnavailable: 0 → Ensures that all running pods stay available while new ones are deployed.
- rollingUpdate.maxSurge: 2 → Allows up to two extra pods to be created during the rollout for smooth transitions.
- revisionHistoryLimit: 1 → Keeps only one previous revision to reduce storage usage in the cluster.
Zero Downtime Impact:
- Ensures no downtime by keeping old pods running until new pods are healthy.
- The system gradually replaces old pods with new ones instead of restarting everything at once.
3. Pod Selector & Template
selector:
matchLabels:
app: first
template:
metadata:
labels:
app: first
- selector.matchLabels.app: first → Ensures that the Deployment manages pods labeled app: first.
- template.metadata.labels.app: first → Assigns the same label inside the pod template for consistency.
Zero Downtime Impact:
- Ensures proper pod replacement by keeping labels consistent.
- Kubernetes can correctly map traffic to new pods during deployments.
4. Dapr Sidecar Annotations for Observability
annotations:
dapr.io/enabled: "true"
dapr.io/app-id: first
dapr.io/app-port: "80"
dapr.io/log-as-json: "true"
dapr.io/log-level: "info"
dapr.io/sidecar-cpu-limit: "50m"
dapr.io/sidecar-memory-limit: "128Mi"
dapr.io/sidecar-cpu-request: "5m"
dapr.io/sidecar-memory-request: "64Mi"
dapr.io/config: "tracing-config"
dapr.io/trace-sampling-rate: "1"
- Enables Dapr Sidecar to add distributed tracing, logging, and observability.
- Limits resource usage for the sidecar process to optimize performance.
Zero Downtime Impact:
- Improves logging and monitoring to detect issues during deployments.
- Ensures minimal sidecar resource usage to prevent overloading the node.
5. Node Selection & IAM Authentication
spec:
nodeSelector:
app:
serviceAccountName: serviceaccount
nodeSelector.app: → Restricts pods to specific nodes based on the app label.
- serviceAccountName: serviceaccount → Uses a predefined IAM Service Account for secure AWS access.
Zero Downtime Impact:
- Ensures pods run on dedicated nodes, avoiding congestion.
- Secure authentication reduces risk of failed deployments due to permission issues.
6. Secret Management with AWS Secrets Manager
volumes:
- name: secrets-store-inline
csi:
driver: secrets-store.csi.k8s.io
readOnly: true
volumeAttributes:
secretProviderClass: <name>-aws-secrets
- Uses Kubernetes Secrets Store CSI Driver to fetch secrets dynamically from AWS Secrets Manager.
Zero Downtime Impact:
- Ensures secure key rotation without pod restarts.
- Avoids hardcoding secrets in deployment files.
7. Graceful Pod Shutdown & Cleanup
terminationGracePeriodSeconds: 120
- terminationGracePeriodSeconds: 120 → Pods wait 120 seconds before shutting down, allowing in-flight requests to complete.
Zero Downtime Impact:
- Prevents abrupt pod termination, avoiding dropped connections.
- Gives NGINX and ALB enough time to remove the pod from the load balancer.
8. Container Configuration & Probes
containers:
- name: first-deployment
image: ecr_url:latest
resources:
requests:
memory: 1Gi
limits:
memory: 2Gi
- Defines memory requests and limits to prevent resource starvation.
- Uses a prebuilt image (ecr_url:latest), ensuring the latest version is deployed.
9. Health Probes for Smooth Traffic Management
startupProbe:
httpGet:
path: /healthz
port: 80
failureThreshold: 60
periodSeconds: 5
- startupProbe → Waits for the application to fully initialize before receiving traffic.
- Prevents ALB and Kubernetes from routing traffic to unready pods.
readinessProbe:
httpGet:
path: /healthz
port: 80
initialDelaySeconds: 40
periodSeconds: 5
successThreshold: 1
failureThreshold: 10
- readinessProbe → Ensures traffic is only sent to fully ready pods.
- Traffic is stopped for failing pods instead of removing them.
livenessProbe:
httpGet:
path: /healthz
port: 80
initialDelaySeconds: 10
periodSeconds: 5
failureThreshold: 3
- livenessProbe → Restarts unresponsive pods automatically.
Zero Downtime Impact:
- Ensures new pods are not overwhelmed with traffic before they are ready.
- Prevents requests from hitting dead pods, improving reliability.
10. Graceful Shutdown with PreStop Hook
lifecycle:
preStop:
exec:
command: ["/bin/sh", "-c", "sleep 60"] # Graceful shutdown delay
- PreStop Hook → Adds a 60-second delay before the pod shuts down, allowing existing connections to finish.
Zero Downtime Impact:
- Gives ALB and Kubernetes enough time to remove the pod from routing tables.
- Ensures users do not experience failed requests during rolling updates.
How We Solved the Problem
We introduced NGINX Ingress Controller to act as an intermediary between ALB and Kubernetes services. Instead of sending traffic directly to the backend services, ALB now routes requests to NGINX, which intelligently manages internal traffic.
How NGINX Fixed the Issue
- NGINX ensures that only healthy, ready pods receive traffic.
- Prevents ALB from sending traffic to draining pods, reducing failed requests.
- Traffic transitions smoothly between old and new pods, ensuring seamless updates.
Understanding the Final Architecture

Traffic Flow: ALB → NGINX → Kubernetes Services
1. ALB handles external traffic and forwards it to NGINX. 2. NGINX intelligently routes traffic to the correct microservice. 3. Kubernetes readiness probes ensure new pods are fully functional before receiving traffic. 4. Pod termination hooks delay shutdown, allowing ALB time to remove old pods safely.
Result: Smooth, zero downtime deployments!
Final Thoughts: Key Takeaways
- ALB alone is not enough — NGINX provides better control over traffic flow.
- NGINX prevents ALB from sending traffic to terminating pods, reducing downtime.
- Readiness probes ensure new pods are ready before receiving traffic.
- PreStop hooks and termination grace periods delay shutdown, preventing traffic gaps.
- Rolling updates with maxUnavailable: 0 prevent downtime.
- PreStop hook & termination grace period allow smooth shutdown.
- Health probes ensure only healthy pods serve traffic.
- AWS Secrets Manager integration avoids restarts for secret updates.
- Dapr sidecar improves tracing & logging for debugging.
By integrating NGINX with ALB, we eliminated downtime during rolling updates, ensuring 100% availability for our services.
This approach has transformed our deployment strategy, providing a scalable, resilient, and high-availability Kubernetes environment.