Fat images slow your CI, bloat your registry, and widen your attack surface. In this lesson you will refactor a naive Dockerfile into a lean multi-stage build, unlock BuildKit caching and secret mounts, and produce images that are smaller, safer, and ready for production Kubernetes clusters.
1. Learning Objectives
By the end of this lesson, you will be able to:
- Explain why single-stage Dockerfiles produce bloated, insecure images
- Refactor a Dockerfile into named builder and runtime stages using
COPY --from - Cut image size using Alpine slim and distroless base images
- Speed up rebuilds with BuildKit cache mounts and dependency-first layer ordering
- Keep secrets out of images with
--mount=type=secret - Verify image size, contents, and vulnerabilities with
docker history, dive, and Trivy
2. Why This Matters
Your team ships a Python microservice in a single-stage image built from python:3.12. The image weighs 1.24 GB because it carries a compiler, git, curl, and a package manager that the running application never touches. Every Kubernetes rollout pulls that image onto every node, so a deployment that should take seconds takes minutes. Your registry bill grows with every push, and the CVE scanner reports hundreds of vulnerabilities - most of them in build tools that are not even present in the final container.
Multi-stage builds solve all of this at once. You build in one stage and copy only the artifacts you need into a minimal runtime stage. The result is an image that is 10 to 20 times smaller, has a dramatically smaller attack surface, pulls faster, and is the standard pattern used in production Dockerfiles today. This lesson walks you through the refactor, the BuildKit features that make it fast, and the verification commands that prove the image is actually lean.
3. Core Concepts
3.1 The Single-Stage Anti-Pattern
A single-stage Dockerfile starts from a full development image, installs everything, and copies the whole world into the final layer. Every command adds a layer, and every layer ships to production:
FROM python:3.12
WORKDIR /app
COPY requirements.txt .
RUN pip install -r requirements.txt
COPY . .
# Build tools the app never needs at runtime - still shipped
RUN apt-get update && apt-get install -y build-essential curl git vim
EXPOSE 8000
CMD ["python", "app.py"]
3.2 Multi-Stage Builds: Build Once, Ship Artifacts
Multi-stage builds let a single Dockerfile contain several FROM instructions. Only the last stage becomes the image, but earlier stages can be named with AS and copied from with COPY --from=<stage>. The compiler and the package manager live in the builder stage, which is thrown away after the build. The runtime stage ships only the compiled result:
# Stage 1: builder - everything needed to assemble the app
FROM python:3.12-slim AS builder
WORKDIR /build
COPY requirements.txt .
RUN pip install --no-cache-dir --prefix=/install -r requirements.txt
# Stage 2: runtime - only what the app needs to run
FROM python:3.12-slim AS runtime
WORKDIR /app
COPY --from=builder /install /usr/local
COPY app.py .
RUN useradd --create-home appuser
USER appuser
EXPOSE 8000
CMD ["python", "app.py"]
3.3 Choosing a Runtime Base Image
The base image of the runtime stage decides how small and how safe the final image is:
- python:3.12 - full Debian with compilers and headers, hundreds of megabytes
- python:3.12-slim - Debian without the build toolchain, a good default for Python apps
- python:3.12-alpine - musl-based and tiny, but some prebuilt wheels (pandas, numpy) need glibc and fail to import
- gcr.io/distroless - no shell, no package manager, no
curl; nothing for an attacker to call after an exploit
Slim and distroless are glibc-based, so they run the same manylinux wheels your CI already builds. Distroless gives the smallest attack surface; the trade-off is that debugging requires the :debug tag variant, which adds a busybox shell.
3.4 BuildKit: Cache Mounts and Secret Mounts
BuildKit is the modern Docker build engine. Two of its features change how you write Dockerfiles. Cache mounts keep package-manager caches between builds, so an unchanged pip install step is instant instead of re-downloading everything:
# syntax=docker/dockerfile:1
FROM python:3.12-slim AS builder
WORKDIR /build
COPY requirements.txt .
RUN --mount=type=cache,target=/root/.cache/pip \
pip install --no-cache-dir --prefix=/install -r requirements.txt
Secret mounts pass build-time credentials (private PyPI tokens, SSH keys) into a build step without writing them into any layer. The secret is available only inside the RUN instruction that mounts it:
# syntax=docker/dockerfile:1
FROM python:3.12-slim AS builder
WORKDIR /build
COPY requirements.txt .
RUN --mount=type=secret,id=pip_token \
PIP_INDEX_URL="https://$(cat /run/secrets/pip_token)@repo.example.com/simple" \
pip install --no-cache-dir --prefix=/install -r requirements.txt
Build with the secret supplied from a local file:
docker build --secret id=pip_token,src=./pip_token.txt -t myapp:latest .
3.5 Layer Caching: Order Matters
Docker reuses a layer only if every layer before it is unchanged. The single most common cache mistake is copying application code before installing dependencies - then every code change re-runs pip install. Copy requirements.txt first, install, and only then copy the code:
COPY requirements.txt .
RUN pip install --no-cache-dir --prefix=/install -r requirements.txt
# Add the app last - code changes do not bust the dependency cache
COPY app.py .
Pin your dependencies to exact versions. A floating >= range makes the layer change on every upstream release, silently invalidating the cache even when you did not touch the file.
3.6 The Humble .dockerignore
Without a .dockerignore, the build context - and every COPY . . - includes local venv folders, .git history, and test artifacts. These can add hundreds of megabytes to the context and can overwrite files inside the image:
.git
__pycache__
*.pyc
.env
venv
.venv
tests
*.md
Dockerfile*
.dockerignore
4. Hands-On Practice
Step 1: Measure the Naive Image
Start from the naive Dockerfile above and build it so you can see the starting point. Save it as naive.Dockerfile, then build and inspect:
docker build -t myapp:naive -f naive.Dockerfile .
docker images myapp
$ docker build -t myapp:naive -f naive.Dockerfile .
[+] Building 84.2s (12/12) FINISHED
$ docker images myapp
REPOSITORY TAG IMAGE ID CREATED SIZE
myapp naive a1b2c3d4e5f6 15 seconds ago 1.24GB
Step 2: Refactor to Multi-Stage
Apply the two-stage structure from section 3.2. The builder stage compiles and installs dependencies into /install; the runtime stage copies only that prefix plus your application code:
# Stage 1: builder
FROM python:3.12-slim AS builder
WORKDIR /build
COPY requirements.txt .
RUN pip install --no-cache-dir --prefix=/install -r requirements.txt
# Stage 2: runtime
FROM python:3.12-slim AS runtime
WORKDIR /app
COPY --from=builder /install /usr/local
COPY app.py .
RUN useradd --create-home appuser
USER appuser
EXPOSE 8000
CMD ["python", "app.py"]
Step 3: Harden the Runtime Stage
Swap the runtime base for distroless to remove the shell and package manager entirely. Python's distroless image already sets python3 as its entrypoint, so the command becomes just the script name:
FROM python:3.12-slim AS builder
WORKDIR /build
COPY requirements.txt .
RUN pip install --no-cache-dir --prefix=/install -r requirements.txt
# Runtime: no shell, no package manager, no curl
FROM gcr.io/distroless/python3-debian12
WORKDIR /app
COPY --from=builder /install /usr/local
COPY app.py .
EXPOSE 8000
CMD ["app.py"]
Step 4: Add BuildKit Cache Mounts
Enable BuildKit and add the cache mount so repeated builds reuse the pip cache. Add the # syntax=docker/dockerfile:1 line at the top of the file, then rebuild twice - the second build skips the download entirely:
# syntax=docker/dockerfile:1
FROM python:3.12-slim AS builder
WORKDIR /build
COPY requirements.txt .
RUN --mount=type=cache,target=/root/.cache/pip \
pip install --no-cache-dir --prefix=/install -r requirements.txt
Step 5: Build and Verify the Final Image
Build the hardened image and compare sizes. Then inspect the layer history to confirm the build toolchain is gone and scan for vulnerabilities:
DOCKER_BUILDKIT=1 docker build -t myapp:distroless -f distroless.Dockerfile .
docker images myapp
docker history myapp:distroless
docker run --rm aquasec/trivy image myapp:distroless
$ DOCKER_BUILDKIT=1 docker build -t myapp:distroless -f distroless.Dockerfile .
[+] Building 12.3s (9/9) FINISHED
$ docker images myapp
REPOSITORY TAG IMAGE ID CREATED SIZE
myapp distroless 9f8e7d6c5b4a 8 seconds ago 78.1MB
myapp naive a1b2c3d4e5f6 15 seconds ago 1.24GB
5. Common Errors & Solutions
ERROR: failed to solve: ... no such stage: build - the alias in COPY --from=build does not match any AS name, usually because of a typo or because the stage uses a different alias.
FIX: Name every stage explicitly (FROM python:3.12-slim AS builder) and reference the exact alias in COPY --from=builder. Prefer copying from a named stage, not from a base image name or an index number.
ERROR: Error loading shared library libgomp.so.1 or ImportError: libc.musl-x86_64.so.1: cannot open shared object file - prebuilt wheels built for glibc are being run on an Alpine (musl) image.
FIX: Use a glibc-based runtime such as python:3.12-slim or a distroless Debian image, or install the Alpine-compatible package set. Match the runtime libc to the one your wheels were built against.
ERROR: The dependency layer rebuilds on every commit even though requirements.txt did not change - the COPY . . of application code sits before the install step, so any file change busts the whole cache chain.
FIX: Copy requirements.txt first, run the install, then copy the code. Pin exact versions so the requirements file only changes when you intend it to, and add a BuildKit cache mount for the package cache.
ERROR: The image is still huge after multi-stage - the build context includes a local venv or .git directory, or pip install ran without --no-cache-dir and left downloaded wheels in the image.
FIX: Add a .dockerignore, install with --no-cache-dir into an isolated prefix, and verify with docker history that build tools appear only in the discarded builder stage.
ERROR: dockerfile parse error or unknown flag: --mount when using cache or secret mounts - the host is running an older Docker engine without BuildKit enabled.
FIX: Put # syntax=docker/dockerfile:1 at the top of the Dockerfile, set DOCKER_BUILDKIT=1 or enable BuildKit in Docker Desktop settings, and upgrade the engine. For debugging distroless images, use the :debug tag variant, which adds a busybox shell.
6. Summary Checklist
- Dockerfile uses at least two stages: a named builder stage and a minimal runtime stage
- Runtime base is
-slimor distroless, not the full development image - Application runs as a non-root user (
USER appuser) .dockerignoreexcludes.git, virtualenvs, and test artifacts- Dependencies are pinned and copied before application code
- BuildKit cache mounts speed up package installs; secrets use
--mount=type=secret - Final size and layer contents verified with
docker imagesanddocker history - Image scanned with Trivy and has no build-time tools in the runtime stage
7. Practice Exercise
Refactor the Dockerfile below into a multi-stage build. Requirements: the final image must run the app as a non-root user, must not include the compiler or git, and must be under 200 MB when built. The app is a Flask service that reads app.py and requirements.txt:
FROM python:3.12
WORKDIR /app
COPY requirements.txt .
RUN pip install -r requirements.txt
COPY . .
RUN apt-get update && apt-get install -y build-essential git
EXPOSE 8000
CMD ["python", "app.py"]
Success criteria:
docker imagesshows your image under 200 MB (distroless should land around 80 MB)docker historyshows noapt-getlayer in the final stage- The container starts with
docker run -p 8000:8000 myapp:finaland serves a request - Trivy reports no critical vulnerabilities in the runtime stage
Stuck? The distroless variant in Step 3 is a complete solution - adapt it and verify each criterion.
8. Next Steps
You have now covered the full Docker series: fundamentals, Dockerfiles, Compose, networking and volumes, security, registries, and production operations. Multi-stage builds and image optimization are the capstone - the skill that makes every other lesson cheaper to run. The next module in the DevOps roadmap is Kubernetes: Container Orchestration at Scale. The lean images you build here are exactly what you will deploy to clusters: small images pull fast during rolling updates, distroless runtimes shrink the blast radius of container escapes, and digest pinning gives you reproducible, immutable deployments. Before moving on, try the advanced follow-up: docker buildx build --platform linux/amd64,linux/arm64 to produce multi-architecture images from the same multi-stage Dockerfile.
Comments (0)
This is exactly what I needed! The initContainer approach solved our migration issues completely. Thanks for the detailed guide!
ReplyLeave a Comment