A cluster sized for peak traffic wastes money at 3 a.m.; one sized for the average collapses under a traffic spike. Kubernetes autoscaling solves both problems at once. In this lesson you will build the full automatic scaling stack: the Horizontal Pod Autoscaler for replica counts, the Vertical Pod Autoscaler for container resource sizes, and the Cluster Autoscaler for node capacity, plus the Metrics Server that powers them all.
1. Learning Objectives
By the end of this lesson, you will be able to:
- Explain the three dimensions of Kubernetes autoscaling: pods with the Horizontal Pod Autoscaler, resources with the Vertical Pod Autoscaler, and nodes with the Cluster Autoscaler
- Install the Metrics Server and confirm it with
kubectl top - Create and tune a Horizontal Pod Autoscaler based on CPU and memory utilization
- Understand how the HPA control loop computes the desired replica count
- Deploy the Vertical Pod Autoscaler and apply its recommendations safely
- Configure the Cluster Autoscaler so your node pool grows and shrinks with demand
- Combine all three autoscalers into one cost-effective scaling strategy
2. Why This Matters
Imagine your online store is featured on a national news site at 9 a.m. By 9:02, traffic jumps from 500 requests per second to 50,000. If you sized your cluster for the peak, you are paying for idle nodes at 3 a.m. when nobody is shopping. If you sized it for the average, your pods are saturated, latency spikes, and customers abandon their carts.
This is the classic scale-versus-cost dilemma, and it is exactly what Kubernetes autoscaling solves. A production cluster that handles real traffic must scale itself: more pods when demand rises, bigger pods when individual pods are undersized, and more nodes when the cluster simply runs out of room. In this lesson you will build all three layers of that automatic scaling stack and learn when each one should act.
3. Core Concepts
Kubernetes offers three autoscalers that work at different layers of the stack. They are complementary, not interchangeable:
- Horizontal Pod Autoscaler (HPA) — changes the number of replicas of a workload based on observed metrics such as CPU, memory, or custom application metrics.
- Vertical Pod Autoscaler (VPA) — changes the resource requests and limits of the containers inside a pod, right-sizing CPU and memory over time.
- Cluster Autoscaler — changes the number of nodes in your cluster, adding nodes when pods cannot be scheduled and removing idle ones.
3.1 The Horizontal Pod Autoscaler (HPA)
The HPA is a control loop that runs in the kube-controller-manager. Every 15 seconds (the default --horizontal-pod-autoscaler-sync-period) it queries the metrics API, compares the current utilization against your target, and computes the desired replica count. The formula the controller uses is:
desiredReplicas = ceil(currentReplicas x (currentMetricValue / desiredMetricValue))
Because the ratio is computed on every loop, the HPA reacts smoothly: it scales up quickly when demand spikes and scales down gradually, honoring the cooldown windows (--horizontal-pod-autoscaler-downscale-stabilization-window, 5 minutes by default) so flapping does not thrash your cluster. Modern HPA objects use apiVersion: autoscaling/v2, which supports CPU, memory, and custom/external metrics in a single object.
3.2 The Vertical Pod Autoscaler (VPA)
Some workloads are hard to scale horizontally: batch jobs, stateful workers, or applications that hold large in-memory caches. For those, the VPA observes actual usage, computes recommended requests and limits using the Vertical Pod Autoscaler Recommender, and — in updateMode: Auto — evicts pods so they are recreated with the right-sized resources. The three VPA components mirror the HPA architecture: the Recommender analyzes usage history, the Updater evicts pods that need new sizes, and the Admission Plugin rewrites pod specs as they are created.
3.3 The Cluster Autoscaler
The third layer works at the infrastructure level. When a pod stays Pending because no node has enough capacity, the Cluster Autoscaler provisions a new node (subject to your minimum and maximum node counts and cloud provider quotas). When nodes are underutilized and all pods can be rescheduled elsewhere, it drains and removes them. It also respects Pod Disruption Budgets and cluster-autoscaler.kubernetes.io/safe-to-evict annotations so it never disrupts critical workloads.
cluster-autoscaler \
--cloud-provider=aws \
--nodes=2:10:default-worker-pool \
--scale-down-utilization-threshold=0.5 \
--scale-down-unneeded-time=10m \
--balance-similar-node-groups
3.4 The Metrics Server
All of this depends on a prerequisite: the Metrics Server. It scrapes pod and node usage from kubelet, aggregates it, and exposes it through the metrics.k8s.io API. Without it, kubectl top fails and the HPA reports <unknown> instead of a utilization percentage. Most managed clusters (EKS, AKS, GKE) still require you to install it explicitly.
4. Hands-On Practice
In this walkthrough you will build the full autoscaling stack on a test cluster. All examples assume kubectl points at a cluster with at least two schedulable worker nodes and enough quota to run 10+ pods.
4.1 Install the Metrics Server
kubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml
# verify the API service is available
kubectl get apiservices | grep metrics
$ kubectl get apiservices | grep metrics
v1beta1.metrics.k8s.io True 5m
Give the deployment a few seconds to start, then confirm it with kubectl top:
kubectl top nodes
kubectl top pods -A
4.2 Deploy a Workload with Resource Requests
The HPA can only compute utilization when your containers declare resources.requests. A container without CPU requests has no baseline to measure against, so the HPA reports <unknown>. Create a simple web deployment with explicit requests:
apiVersion: apps/v1
kind: Deployment
metadata:
name: web-app
spec:
replicas: 2
selector:
matchLabels:
app: web-app
template:
metadata:
labels:
app: web-app
spec:
containers:
- name: web
image: nginx:1.27
ports:
- containerPort: 80
resources:
requests:
cpu: 250m
memory: 128Mi
limits:
cpu: 500m
memory: 256Mi
kubectl apply -f web-app.yaml
kubectl get deployment web-app
4.3 Create a Horizontal Pod Autoscaler
You can create the HPA imperatively or from YAML. The imperative form is perfect for a quick test:
kubectl autoscale deployment web-app \
--cpu-percent=60 \
--min=2 \
--max=10
$ kubectl get hpa web-app
NAME REFERENCE TARGETS MINPODS MAXPODS REPLICAS AGE
web-app Deployment/web-app 2%/60% 2 10 2 20s
The declarative autoscaling/v2 form lets you target memory and custom metrics in the same object:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: web-app-hpa
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: web-app
minReplicas: 2
maxReplicas: 10
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 60
- type: Resource
resource:
name: memory
target:
type: Utilization
averageUtilization: 70
4.4 Generate Load and Watch It Scale
Open a second terminal and generate traffic against the service, then watch the HPA react. Within a minute the replica count should climb past the initial two:
kubectl run -i --tty load-generator --rm \
--image=busybox --restart=Never -- \
/bin/sh -c "while true; do wget -q -O- http://web-app-service; done"
kubectl get hpa web-app --watch
$ kubectl get hpa web-app --watch
NAME REFERENCE TARGETS MINPODS MAXPODS REPLICAS AGE
web-app Deployment/web-app 38%/60% 2 10 2 1m
web-app Deployment/web-app 74%/60% 2 10 4 2m
web-app Deployment/web-app 61%/60% 2 10 6 3m
web-app Deployment/web-app 58%/60% 2 10 6 4m
Stop the load generator with Ctrl+C and delete it. Notice that scale-down is much slower than scale-up: the HPA waits out the 5-minute downscale stabilization window so a brief traffic dip does not collapse your replica count.
4.5 Set Up the Vertical Pod Autoscaler
The VPA ships in the kubernetes/autoscaler repository. On a non-managed cluster, deploy all three components with the convenience script:
git clone https://github.com/kubernetes/autoscaler.git
cd autoscaler/vertical-pod-autoscaler
./hack/vpa-up.sh
Now create a VPA object that targets the web-app deployment. Use updateMode: Off first so you can inspect recommendations before letting the VPA evict anything:
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: web-app-vpa
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: web-app
updatePolicy:
updateMode: Off
resourcePolicy:
containerPolicies:
- containerName: "*"
minAllowed:
cpu: 100m
memory: 64Mi
maxAllowed:
cpu: "2"
memory: 2Gi
kubectl apply -f web-app-vpa.yaml
kubectl describe vpa web-app-vpa
$ kubectl describe vpa web-app-vpa
Recommendation:
Container: web
Lower Bound: 100m CPU, 64Mi memory
Target: 412m CPU, 180Mi memory
Uncapped Target: 450m CPU, 190Mi memory
Upper Bound: 1 CPU, 1Gi memory
Status: RecommendationProvided
After observing recommendations for a while, flip updateMode to Auto to let the VPA evict and recreate pods with the recommended sizes. Always start in Off mode in production: the VPA cannot see every workload nuance, and its evictions restart your pods.
4.6 Deploy the Cluster Autoscaler
The Cluster Autoscaler is cloud-provider specific. On AWS EKS it runs as a Deployment using IRSA; on AKS it is enabled through the az CLI; on GKE it is a checkbox on the node pool. The universal signal that it is needed is a pod stuck in Pending with a scheduler message about insufficient CPU or memory:
kubectl describe pod pending-app | grep -A5 Events
# after the Cluster Autoscaler adds a node:
kubectl get nodes
kubectl get pods -o wide
$ kubectl get nodes
NAME STATUS ROLES AGE VERSION
ip-10-0-1-101 Ready <none> 45m v1.30.0
ip-10-0-1-102 Ready <none> 45m v1.30.0
ip-10-0-1-201 Ready <none> 2m v1.30.0 <-- added by autoscaler
Set sane minimum and maximum node counts on the node group. The Cluster Autoscaler will never scale below the minimum — that is your guard rail for reserved capacity — and will never exceed the maximum, which is your cost ceiling.
5. Common Errors & Solutions
- HPA shows
<unknown>/60%in the TARGETS column — The container has no CPUrequests, so the controller cannot compute utilization. Fix: declareresources.requests.cpuon every container, then delete and recreate the HPA. - kubectl top nodes fails with "error: metrics not available yet" — The Metrics Server is not running or its aggregated API has not registered. Fix: verify the deployment and
kubectl get apiservices | grep metrics; on some distributions you must add--kubelet-insecure-tlsto the Metrics Server args for self-signed kubelet certs. - HPA scales up forever and never scales down — The downscale stabilization window and scale-up can interact with a too-low target. Fix: raise the target utilization or check
kubectl describe hpafor the actual metric; also confirm your application is not simply CPU-bound at the request level you set. - VPA shows no recommendation (status "NoRecommendation") — The Recommender needs at least 24 hours of usage history, and containers must declare requests. Fix: wait, or seed usage with
kubectl runload; checkvpa-recommenderlogs for MissingRequest warnings. - Cluster Autoscaler never scales down — Idle nodes are protected by Pod Disruption Budgets, the
safe-to-evictannotation defaulting to false on pods with local storage, or the node count already being at the configured minimum. Fix: addcluster-autoscaler.kubernetes.io/safe-to-evict: "true"to disposable workloads and review--scale-down-utilization-threshold. - Custom metrics do not trigger the HPA — Resource metrics come from the Metrics Server, but custom metrics require an adapter such as Prometheus Adapter. Fix: deploy the adapter and point the HPA at
type: Podsortype: Objectmetrics with the correct metric name.
6. Summary Checklist
- I can explain what each of the three autoscalers scales: replicas (HPA), resources (VPA), and nodes (Cluster Autoscaler).
- I installed the Metrics Server and verified it with
kubectl top nodesandkubectl top pods. - I created a Deployment whose containers declare CPU and memory requests.
- I created an HPA with
autoscaling/v2targeting CPU and memory. - I generated load and watched the HPA scale replicas up, then stabilize before scaling down.
- I deployed the VPA, ran it in
Offmode, and inspected its recommendations before enablingAuto. - I know how to enable the Cluster Autoscaler on my cloud provider and how to set node group min/max bounds.
- I can diagnose the most common autoscaling failures: missing requests, missing Metrics Server, and blocked scale-downs.
7. Practice Exercise
Build a complete autoscaling demo in your own cluster:
- Create a Deployment named
apirunningnginx:1.27with 3 replicas,requests.cpu: 200m, andrequests.memory: 128Mi. - Expose it with a ClusterIP service named
api. - Create an HPA targeting 50% CPU with
minReplicas: 3andmaxReplicas: 12. - Run the busybox load generator from section 4.4 for 5 minutes and record how the replica count changes over time.
- Create a VPA in
Offmode for theapiDeployment and read its recommendation after 10 minutes. - Answer these questions in a short write-up: Why did the HPA scale up in steps rather than jumping straight to 12? Why is the VPA recommendation usually larger than the request you originally set? What happens to the Cluster Autoscaler if the node pool maximum is reached while pods are still pending?
8. Next Steps
You have now automated the full scaling stack: the HPA keeps replica counts in line with demand, the VPA right-sizes individual containers, and the Cluster Autoscaler matches your node pool to the workload. Together they are what let a modern platform run at 40% of the cost of a peak-sized static cluster.
In the next lesson, we will explore Kubernetes Ingress and Ingress Controllers — routing external traffic to your services with path-based rules, TLS termination, and annotations. That is the piece that connects your autoscaled, secured workloads to the outside world.
Comments (0)
This is exactly what I needed! The initContainer approach solved our migration issues completely. Thanks for the detailed guide!
ReplyLeave a Comment