Module 02: Dockerfiles
Build, Layers & the Cache
Sam has been shipping their first containerized app, and so far Sam only ran images other people built. This module is where you build your own, and, more importantly, build them so they're small, fast to rebuild, and safe. The build cache is the single concept that separates a 12-minute rebuild from a 4-second one. Get it into your bones here.
Where we left off
Module 01 ended on a cliff: "data you write inside a container is not in the image, and a package you
install with docker exec vanishes on docker rm." Sam hit this
wall hard. Every time their app's container restarted, the dependencies Sam had hand-installed were gone. The fix to
that, the only proper one, is to bake what you need into the image at build time. That is what a
Dockerfile does, and it's exactly the tool Sam needs to ship for real.
Image = layers
A stack of read-only layers. Every Dockerfile instruction that changes the filesystem adds one layer.
Dockerfile = recipe
A plain text file of instructions. Run through docker build it produces an image, repeatably.
Cache = speed
Unchanged layers are reused from the last build. Order your instructions and rebuilds go from minutes to seconds.
"Write a Dockerfile for this app," "why is your image 1.2 GB," and "your CI rebuild takes 10 minutes, fix it" are three of the most common practical screens you will hit. All three are answered by understanding layers and the cache. This is the module that makes you useful, not just conversant.
What a Dockerfile actually is
A Dockerfile is a text file (literally named Dockerfile, no extension) listing the steps
to assemble an image, top to bottom. docker build reads it, runs each instruction in order,
and snapshots the filesystem after each one into a layer. The final stack of layers is your image.
Here is the smallest useful, real-world example, the Python web app Sam is shipping. Read it once; we dissect every line below.
# 1. Start from an official base image (pinned tag)
FROM python:3.12-slim
# 2. Set the working directory inside the image
WORKDIR /app
# 3. Copy ONLY the dependency manifest first (cache trick: section 4)
COPY requirements.txt .
# 4. Install dependencies
RUN pip install --no-cache-dir -r requirements.txt
# 5. Now copy the rest of the source code
COPY . .
# 6. Document the port the app listens on
EXPOSE 8000
# 7. The default command when a container starts
CMD ["python", "app.py"]
Build it and run it:
# -t names (tags) the image; the "." is the build context (section 5)
docker build -t myapp:1.0 .
# run a container from the image you just built
docker run -p 8000:8000 myapp:1.0
A Dockerfile is a declarative, ordered recipe. docker build executes it
instruction by instruction, caching each step. The output is an immutable image. The same Dockerfile + same
context produces the same image, anywhere. This is the death of "works on my machine."
The instructions you must know cold
There are ~18 Dockerfile instructions. You use about 10 daily. Know these by heart, because you'll use the distinctions constantly (especially the ones marked).
| Instruction | What it does | New layer? |
|---|---|---|
FROM | Base image to build on. Must be the first real instruction. One per stage. | base |
WORKDIR | Sets the current directory for following instructions (creates it if missing). Use instead of RUN cd. | metadata |
COPY | Copies files/dirs from the build context into the image. | yes |
ADD | Like COPY, but also auto-extracts local tarballs and can fetch URLs. Prefer COPY unless you need those. | yes |
RUN | Executes a command at build time (install packages, compile). Each RUN = a layer. | yes |
ENV | Sets an environment variable, persisted into the image and into running containers. | metadata |
ARG | Build-time variable (--build-arg). Not available at run time. Gone once built. | metadata |
EXPOSE | Documents which port the app listens on. Does not publish it (that's -p at run). | metadata |
USER | Sets the user for following instructions and the running container. Drop root (section 7). | metadata |
CMD | Default command/args when a container starts. Overridable on the CLI. One effective CMD. | metadata |
ENTRYPOINT | The fixed executable a container always runs. CMD becomes its default args. | metadata |
LABEL | Metadata key/values (maintainer, version, source repo). | metadata |
Only instructions that change the filesystem (RUN, COPY,
ADD) create a real, sized layer. The rest write image metadata (config the engine
reads). Modern Docker/BuildKit squashes metadata efficiently, so don't lose sleep over an extra
ENV, but every RUN and COPY is real
weight and a real cache boundary.
ADD vs COPY (common gotcha)
Both copy files in. ADD has two extra "magic" behaviours: it auto-extracts a local
.tar and can download a remote URL. That magic causes surprises and security issues, so the
rule the docs themselves give: use COPY always, reach for ADD
only when you specifically want tar auto-extraction. Knowing this signals you've read the
guidelines.
CMD vs ENTRYPOINT, the classic trap
This pair confuses everyone at first and is asked constantly. Sam tripped here too: their app ignored
docker stop until they understood the split. The clean way to hold it:
ENTRYPOINT = what runs
The fixed executable. Think "this container is a curl" / "this container is nginx." Hard to override (needs --entrypoint).
CMD = default arguments
The default args passed to that executable (or the whole default command if there's no ENTRYPOINT). Easily overridden; just type args after the image name.
They combine. With both set, the container runs ENTRYPOINT + CMD. Anything you type
after the image name on docker run replaces the CMD part:
ENTRYPOINT ["ping"]
CMD ["localhost"]
# docker run myimg -> ping localhost (uses default CMD)
# docker run myimg google.com -> ping google.com (CMD replaced by arg)
Always prefer the exec form, a JSON array: CMD ["python", "app.py"]. The
shell form CMD python app.py wraps your command in /bin/sh -c,
which means your process is not PID 1. It becomes a child of the shell, and Unix signals
(SIGTERM from docker stop) don't reach it. Result: containers that
ignore stop and get SIGKILL'd. Exec form makes your app PID 1 and signals work. This ties
directly back to Module 01's "container lives as long as PID 1."
ENTRYPOINT defines the executable the container always runs; CMD provides default arguments that the user can override on the command line. Use ENTRYPOINT when the image is a single-purpose tool; use CMD alone when you want the command fully overridable.
Layers & the build cache, the core skill
This is the most valuable thing in the module, and the part that finally made Sam's builds bearable. Sam's
rebuild was crawling at 12 minutes per code change, killing their shipping momentum. The cache is the fix.
docker build caches each instruction's
result. On the next build it walks down the Dockerfile and, for each instruction, asks: "has this instruction
or any instruction above it changed?"
- No change → reuse the cached layer instantly (you'll see
CACHEDin the output). - Changed → rebuild this layer and every layer below it. The cache is invalidated downward.
That second rule, a cache miss cascades to everything underneath, is the whole game. It means instruction order decides rebuild speed. Put the things that rarely change near the top, the things that change constantly (your source code) near the bottom.
The classic mistake vs the fix
FROM python:3.12-slim
WORKDIR /app
COPY . .
RUN pip install -r requirements.txt
CMD ["python", "app.py"]
FROM python:3.12-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install -r requirements.txt
COPY . .
CMD ["python", "app.py"]
In the left version, editing one line of source changes the COPY . .
layer, which sits above the pip install, so dependencies reinstall on every build (minutes). In the right
version, requirements.txt rarely changes, so its layer and the pip install stay cached; only
the final COPY . . reruns (seconds).
What the cache actually does on a code-only edit, layer by layer:
1) apt-get update in its own RUN caches a stale package list,
so always chain it: RUN apt-get update && apt-get install -y curl. 2) A changing
ARG or a COPY of a file whose timestamp/content changed busts
everything below. 3) Force a clean build with docker build --no-cache when you suspect a
stale layer.
Build context & .dockerignore
The . at the end of docker build -t myapp . is the
build context: the directory the daemon receives to build from. COPY can
only see files inside it. Crucially, the entire context is sent to the daemon first, so if your folder has a
500 MB node_modules, a .git history, and your virtualenv, all of it
ships before the build even starts. Slow builds, fat images, leaked secrets.
Fix: a .dockerignore file (same syntax as .gitignore) next to the
Dockerfile. It excludes paths from the context entirely.
.git
node_modules
__pycache__
*.pyc
.env
.venv
Dockerfile
README.md
Never COPY . . without a .dockerignore. A stray
.env, .git with old secrets in history, or cloud credentials file
gets baked into an image layer, and image layers are readable by anyone who pulls the image. Deleting the
file in a later layer does not remove it; it still exists in the earlier layer. Keep secrets out of the
context entirely.
Multi-stage builds, the size killer
Building an app often needs heavy tooling (compilers, build deps, dev packages) that the final running image
does not need. Shipping all of it makes a 1 GB image where 80 MB would do. Multi-stage builds solve this: use
multiple FROM stages, do the heavy work in an early "builder" stage, then copy only the
finished artifact into a clean, tiny final stage. Everything in the builder stage is thrown away.
# ── stage 1: build (heavy toolchain) ──
FROM golang:1.22 AS builder
WORKDIR /src
COPY . .
RUN go build -o /bin/app ./cmd/app
# ── stage 2: final (tiny, runtime only) ──
FROM gcr.io/distroless/base-debian12
COPY --from=builder /bin/app /app
ENTRYPOINT ["/app"]
The magic line is COPY --from=builder /bin/app /app: it pulls just the compiled binary
out of the builder stage. The Go compiler, source code, and module cache never reach the final image. A 900 MB
builder collapses into a ~20 MB final image.
Smaller image = faster pulls, faster deploys, smaller attack surface (fewer packages = fewer CVEs), lower
registry cost. This is a top "how would you optimize this image" answer: multi-stage build + slim/distroless
base + .dockerignore. The same pattern applies to Node (build with full toolchain, copy dist/
+ production node_modules into a slim runtime) and Python.
Small & safe images, the best-practice checklist
The habits that separate a professional Dockerfile from a tutorial one. Each is a habit worth building.
| Practice | Why |
|---|---|
| Pick a slim/alpine/distroless base | python:3.12-slim not python:3.12. Hundreds of MB and dozens of CVEs avoided. |
| Pin the tag (or digest) | :3.12-slim not :latest. Reproducible, no surprise breakage (ties to Module 01). |
| Chain RUN commands | apt-get update && install ... && rm -rf /var/lib/apt/lists/* in one RUN: fewer layers, no stale cache, cleanup in same layer. |
| Clean up in the same layer | Deleting a cache in a later layer doesn't shrink the image; the bytes live in the earlier layer. Delete where you created. |
| Run as a non-root USER | Default container user is root. A breakout then = host root-ish. Add a user and USER appuser. (Hardening depth = Module 07.) |
| Use a .dockerignore | Smaller context, faster builds, no leaked secrets (section 5). |
| Multi-stage for compiled/bundled apps | Ship the artifact, not the toolchain (section 6). |
| Order for the cache | Stable layers up top, volatile source at the bottom (section 4). |
RUN apt-get update \
&& apt-get install -y --no-install-recommends curl \
&& rm -rf /var/lib/apt/lists/*
# create an unprivileged user and switch to it
RUN useradd --create-home appuser
USER appuser
Hands-on, do this now
Reading this module teaches you nothing until you watch the cache hit and miss yourself. This is the exact loop that turned Sam's 12-minute rebuild into seconds. Spend 30 minutes here.
- Make a folder with a trivial app (a
requirements.txtwith one package + a 3-lineapp.py). Write the "slow" Dockerfile from section 4. - Build it twice in a row. Watch the second build: every line says
CACHED. Time it. - Change one line in app.py and rebuild. Notice
pip installreruns even though deps didn't change. That's the bug. - Switch to the "fast" Dockerfile (copy requirements first). Rebuild, edit app.py, rebuild again. Now pip install stays
CACHED. Feel the difference. - Run
docker history myapp:1.0anddocker imagesto see each layer's size and the total image size. - Add a
.dockerignore, then add a fake.envto the folder. Confirm withdocker buildoutput that the context shrank. - Convert any compiled-language toy (or a Node app) to a multi-stage build. Compare image sizes before/after with
docker images. - Break the cache on purpose:
docker build --no-cacheand watch every layer rebuild.
Write the app.py yourself, even three lines. The point of this course is that you can produce these files from memory on the job, not that you pasted them. If you get stuck, sketch the steps in plain English first, then translate each step to a Dockerfile instruction.
How much further to "I know Docker well"?
You asked the right question. Here is the honest map. You are 3 of ~8 Docker modules in. The good news: you can start a real project sooner than you can claim full mastery. Those are two different milestones. Sam is shipping already, hardening comes later.
| Module | Topic | Status |
|---|---|---|
| 00 | Foundations: what a container is (namespaces, cgroups) | done ✓ |
| 01 | Images & containers: lifecycle, core CLI | done ✓ |
| 02 | Dockerfiles: build, layers, cache, multi-stage | this one ✓ |
| 03 | Volumes & storage: persisting data, bind mounts vs volumes | next |
| 04 | Networking: bridge, ports, container-to-container DNS | upcoming |
| 05 | Docker Compose: multi-container apps from one YAML file | upcoming |
| 06 | Registries: push/pull, tagging, Docker Hub / GHCR | upcoming |
| 07 | Security & hardening: non-root, capabilities, scanning, secrets | upcoming |
Start a project after Module 05
Build + Volumes + Networking + Compose = enough to containerize a real multi-service app (API + DB + cache) end to end. That's 2 modules away. A great portfolio piece, and where most learning sticks.
Say "I know Docker well" after Module 07
Once security/hardening is in, you cover everything the "Docker" toolkit demands on the job. That's 4 modules away. After that comes orchestration (Kubernetes), a separate phase, not "Docker" anymore.
Don't wait until Module 07 to touch a project. The fastest way to "know Docker well" is to build a real multi-container app at Module 05 and then circle back to harden it. Modules 03–05 are short and practical; you could be project-ready within a week of focused evenings. Reading without building is how people "finish a course" and still freeze the moment they sit down to do the real work.