A cluster sized for peak traffic wastes money at 3 a.m.; one sized for the average collapses under a traffic spike. Kubernetes autoscaling solves both problems at once. In this lesson you will build the full automatic scaling stack: the Horizontal Pod Autoscaler for replica counts, the Vertical Pod Autoscaler for container resource sizes, and the Cluster Autoscaler for node capacity, plus the Metrics Server that powers them all.

1. Learning Objectives

By the end of this lesson, you will be able to:

  • Explain the three dimensions of Kubernetes autoscaling: pods with the Horizontal Pod Autoscaler, resources with the Vertical Pod Autoscaler, and nodes with the Cluster Autoscaler
  • Install the Metrics Server and confirm it with kubectl top
  • Create and tune a Horizontal Pod Autoscaler based on CPU and memory utilization
  • Understand how the HPA control loop computes the desired replica count
  • Deploy the Vertical Pod Autoscaler and apply its recommendations safely
  • Configure the Cluster Autoscaler so your node pool grows and shrinks with demand
  • Combine all three autoscalers into one cost-effective scaling strategy

2. Why This Matters

Imagine your online store is featured on a national news site at 9 a.m. By 9:02, traffic jumps from 500 requests per second to 50,000. If you sized your cluster for the peak, you are paying for idle nodes at 3 a.m. when nobody is shopping. If you sized it for the average, your pods are saturated, latency spikes, and customers abandon their carts.

This is the classic scale-versus-cost dilemma, and it is exactly what Kubernetes autoscaling solves. A production cluster that handles real traffic must scale itself: more pods when demand rises, bigger pods when individual pods are undersized, and more nodes when the cluster simply runs out of room. In this lesson you will build all three layers of that automatic scaling stack and learn when each one should act.

3. Core Concepts

Kubernetes offers three autoscalers that work at different layers of the stack. They are complementary, not interchangeable:

  • Horizontal Pod Autoscaler (HPA) — changes the number of replicas of a workload based on observed metrics such as CPU, memory, or custom application metrics.
  • Vertical Pod Autoscaler (VPA) — changes the resource requests and limits of the containers inside a pod, right-sizing CPU and memory over time.
  • Cluster Autoscaler — changes the number of nodes in your cluster, adding nodes when pods cannot be scheduled and removing idle ones.

3.1 The Horizontal Pod Autoscaler (HPA)

The HPA is a control loop that runs in the kube-controller-manager. Every 15 seconds (the default --horizontal-pod-autoscaler-sync-period) it queries the metrics API, compares the current utilization against your target, and computes the desired replica count. The formula the controller uses is:

desiredReplicas = ceil(currentReplicas x (currentMetricValue / desiredMetricValue))
The HPA scaling formula

Because the ratio is computed on every loop, the HPA reacts smoothly: it scales up quickly when demand spikes and scales down gradually, honoring the cooldown windows (--horizontal-pod-autoscaler-downscale-stabilization-window, 5 minutes by default) so flapping does not thrash your cluster. Modern HPA objects use apiVersion: autoscaling/v2, which supports CPU, memory, and custom/external metrics in a single object.

3.2 The Vertical Pod Autoscaler (VPA)

Some workloads are hard to scale horizontally: batch jobs, stateful workers, or applications that hold large in-memory caches. For those, the VPA observes actual usage, computes recommended requests and limits using the Vertical Pod Autoscaler Recommender, and — in updateMode: Auto — evicts pods so they are recreated with the right-sized resources. The three VPA components mirror the HPA architecture: the Recommender analyzes usage history, the Updater evicts pods that need new sizes, and the Admission Plugin rewrites pod specs as they are created.

3.3 The Cluster Autoscaler

The third layer works at the infrastructure level. When a pod stays Pending because no node has enough capacity, the Cluster Autoscaler provisions a new node (subject to your minimum and maximum node counts and cloud provider quotas). When nodes are underutilized and all pods can be rescheduled elsewhere, it drains and removes them. It also respects Pod Disruption Budgets and cluster-autoscaler.kubernetes.io/safe-to-evict annotations so it never disrupts critical workloads.

cluster-autoscaler \
  --cloud-provider=aws \
  --nodes=2:10:default-worker-pool \
  --scale-down-utilization-threshold=0.5 \
  --scale-down-unneeded-time=10m \
  --balance-similar-node-groups
Key Cluster Autoscaler flags (AWS example)

3.4 The Metrics Server

All of this depends on a prerequisite: the Metrics Server. It scrapes pod and node usage from kubelet, aggregates it, and exposes it through the metrics.k8s.io API. Without it, kubectl top fails and the HPA reports <unknown> instead of a utilization percentage. Most managed clusters (EKS, AKS, GKE) still require you to install it explicitly.

4. Hands-On Practice

In this walkthrough you will build the full autoscaling stack on a test cluster. All examples assume kubectl points at a cluster with at least two schedulable worker nodes and enough quota to run 10+ pods.

4.1 Install the Metrics Server

kubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml

# verify the API service is available
kubectl get apiservices | grep metrics
Installing the Metrics Server
kubectl get pods -w
$ kubectl get apiservices | grep metrics
v1beta1.metrics.k8s.io True 5m
Metrics Server API service registered

Give the deployment a few seconds to start, then confirm it with kubectl top:

kubectl top nodes
kubectl top pods -A
Confirm resource metrics are flowing

4.2 Deploy a Workload with Resource Requests

The HPA can only compute utilization when your containers declare resources.requests. A container without CPU requests has no baseline to measure against, so the HPA reports <unknown>. Create a simple web deployment with explicit requests:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: web-app
spec:
  replicas: 2
  selector:
    matchLabels:
      app: web-app
  template:
    metadata:
      labels:
        app: web-app
    spec:
      containers:
        - name: web
          image: nginx:1.27
          ports:
            - containerPort: 80
          resources:
            requests:
              cpu: 250m
              memory: 128Mi
            limits:
              cpu: 500m
              memory: 256Mi
Deployment with CPU and memory requests
kubectl apply -f web-app.yaml
kubectl get deployment web-app
Apply the deployment

4.3 Create a Horizontal Pod Autoscaler

You can create the HPA imperatively or from YAML. The imperative form is perfect for a quick test:

kubectl autoscale deployment web-app \
  --cpu-percent=60 \
  --min=2 \
  --max=10
Create the HPA targeting 60% CPU
kubectl get pods -w
$ kubectl get hpa web-app
NAME REFERENCE TARGETS MINPODS MAXPODS REPLICAS AGE
web-app Deployment/web-app 2%/60% 2 10 2 20s
HPA shows a healthy 2% CPU utilization

The declarative autoscaling/v2 form lets you target memory and custom metrics in the same object:

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: web-app-hpa
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: web-app
  minReplicas: 2
  maxReplicas: 10
  metrics:
    - type: Resource
      resource:
        name: cpu
        target:
          type: Utilization
          averageUtilization: 60
    - type: Resource
      resource:
        name: memory
        target:
          type: Utilization
          averageUtilization: 70
Declarative HPA with CPU and memory targets

4.4 Generate Load and Watch It Scale

Open a second terminal and generate traffic against the service, then watch the HPA react. Within a minute the replica count should climb past the initial two:

kubectl run -i --tty load-generator --rm \
  --image=busybox --restart=Never -- \
  /bin/sh -c "while true; do wget -q -O- http://web-app-service; done"
Generate load with a busybox pod
kubectl get hpa web-app --watch
Watch the HPA in real time
kubectl get pods -w
$ kubectl get hpa web-app --watch
NAME REFERENCE TARGETS MINPODS MAXPODS REPLICAS AGE
web-app Deployment/web-app 38%/60% 2 10 2 1m
web-app Deployment/web-app 74%/60% 2 10 4 2m
web-app Deployment/web-app 61%/60% 2 10 6 3m
web-app Deployment/web-app 58%/60% 2 10 6 4m
HPA scales replicas from 2 to 6 under load

Stop the load generator with Ctrl+C and delete it. Notice that scale-down is much slower than scale-up: the HPA waits out the 5-minute downscale stabilization window so a brief traffic dip does not collapse your replica count.

4.5 Set Up the Vertical Pod Autoscaler

The VPA ships in the kubernetes/autoscaler repository. On a non-managed cluster, deploy all three components with the convenience script:

git clone https://github.com/kubernetes/autoscaler.git
cd autoscaler/vertical-pod-autoscaler
./hack/vpa-up.sh
Deploy the VPA components

Now create a VPA object that targets the web-app deployment. Use updateMode: Off first so you can inspect recommendations before letting the VPA evict anything:

apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
  name: web-app-vpa
spec:
  targetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: web-app
  updatePolicy:
    updateMode: Off
  resourcePolicy:
    containerPolicies:
      - containerName: "*"
        minAllowed:
          cpu: 100m
          memory: 64Mi
        maxAllowed:
          cpu: "2"
          memory: 2Gi
VPA in Off mode to collect recommendations first
kubectl apply -f web-app-vpa.yaml
kubectl describe vpa web-app-vpa
Apply the VPA and inspect recommendations
kubectl get pods -w
$ kubectl describe vpa web-app-vpa
Recommendation:
Container: web
Lower Bound: 100m CPU, 64Mi memory
Target: 412m CPU, 180Mi memory
Uncapped Target: 450m CPU, 190Mi memory
Upper Bound: 1 CPU, 1Gi memory
Status: RecommendationProvided
VPA Recommender output with target values

After observing recommendations for a while, flip updateMode to Auto to let the VPA evict and recreate pods with the recommended sizes. Always start in Off mode in production: the VPA cannot see every workload nuance, and its evictions restart your pods.

4.6 Deploy the Cluster Autoscaler

The Cluster Autoscaler is cloud-provider specific. On AWS EKS it runs as a Deployment using IRSA; on AKS it is enabled through the az CLI; on GKE it is a checkbox on the node pool. The universal signal that it is needed is a pod stuck in Pending with a scheduler message about insufficient CPU or memory:

kubectl describe pod pending-app | grep -A5 Events

# after the Cluster Autoscaler adds a node:
kubectl get nodes
kubectl get pods -o wide
Detect pending pods, then confirm the new node
kubectl get pods -w
$ kubectl get nodes
NAME STATUS ROLES AGE VERSION
ip-10-0-1-101 Ready <none> 45m v1.30.0
ip-10-0-1-102 Ready <none> 45m v1.30.0
ip-10-0-1-201 Ready <none> 2m v1.30.0 <-- added by autoscaler
Cluster Autoscaler provisions a third node

Set sane minimum and maximum node counts on the node group. The Cluster Autoscaler will never scale below the minimum — that is your guard rail for reserved capacity — and will never exceed the maximum, which is your cost ceiling.

5. Common Errors & Solutions

  • HPA shows <unknown>/60% in the TARGETS column — The container has no CPU requests, so the controller cannot compute utilization. Fix: declare resources.requests.cpu on every container, then delete and recreate the HPA.
  • kubectl top nodes fails with "error: metrics not available yet" — The Metrics Server is not running or its aggregated API has not registered. Fix: verify the deployment and kubectl get apiservices | grep metrics; on some distributions you must add --kubelet-insecure-tls to the Metrics Server args for self-signed kubelet certs.
  • HPA scales up forever and never scales down — The downscale stabilization window and scale-up can interact with a too-low target. Fix: raise the target utilization or check kubectl describe hpa for the actual metric; also confirm your application is not simply CPU-bound at the request level you set.
  • VPA shows no recommendation (status "NoRecommendation") — The Recommender needs at least 24 hours of usage history, and containers must declare requests. Fix: wait, or seed usage with kubectl run load; check vpa-recommender logs for MissingRequest warnings.
  • Cluster Autoscaler never scales down — Idle nodes are protected by Pod Disruption Budgets, the safe-to-evict annotation defaulting to false on pods with local storage, or the node count already being at the configured minimum. Fix: add cluster-autoscaler.kubernetes.io/safe-to-evict: "true" to disposable workloads and review --scale-down-utilization-threshold.
  • Custom metrics do not trigger the HPA — Resource metrics come from the Metrics Server, but custom metrics require an adapter such as Prometheus Adapter. Fix: deploy the adapter and point the HPA at type: Pods or type: Object metrics with the correct metric name.

6. Summary Checklist

  • I can explain what each of the three autoscalers scales: replicas (HPA), resources (VPA), and nodes (Cluster Autoscaler).
  • I installed the Metrics Server and verified it with kubectl top nodes and kubectl top pods.
  • I created a Deployment whose containers declare CPU and memory requests.
  • I created an HPA with autoscaling/v2 targeting CPU and memory.
  • I generated load and watched the HPA scale replicas up, then stabilize before scaling down.
  • I deployed the VPA, ran it in Off mode, and inspected its recommendations before enabling Auto.
  • I know how to enable the Cluster Autoscaler on my cloud provider and how to set node group min/max bounds.
  • I can diagnose the most common autoscaling failures: missing requests, missing Metrics Server, and blocked scale-downs.

7. Practice Exercise

Build a complete autoscaling demo in your own cluster:

  1. Create a Deployment named api running nginx:1.27 with 3 replicas, requests.cpu: 200m, and requests.memory: 128Mi.
  2. Expose it with a ClusterIP service named api.
  3. Create an HPA targeting 50% CPU with minReplicas: 3 and maxReplicas: 12.
  4. Run the busybox load generator from section 4.4 for 5 minutes and record how the replica count changes over time.
  5. Create a VPA in Off mode for the api Deployment and read its recommendation after 10 minutes.
  6. Answer these questions in a short write-up: Why did the HPA scale up in steps rather than jumping straight to 12? Why is the VPA recommendation usually larger than the request you originally set? What happens to the Cluster Autoscaler if the node pool maximum is reached while pods are still pending?

8. Next Steps

You have now automated the full scaling stack: the HPA keeps replica counts in line with demand, the VPA right-sizes individual containers, and the Cluster Autoscaler matches your node pool to the workload. Together they are what let a modern platform run at 40% of the cost of a peak-sized static cluster.

In the next lesson, we will explore Kubernetes Ingress and Ingress Controllers — routing external traffic to your services with path-based rules, TLS termination, and annotations. That is the piece that connects your autoscaled, secured workloads to the outside world.

Gataya Med

DevOps Engineer & Backend Developer. Sharing insights on cloud, automation, and scalable systems.

Comments (0)

Sarah Chen August 9, 2026

This is exactly what I needed! The initContainer approach solved our migration issues completely. Thanks for the detailed guide!

Reply

Leave a Comment