← Learn
Docker Course  /  Phase 1  /  Module 03
Day 3 ~2h · hands-on
Phase 1 · Docker · Day 3

Module 03: Volumes & Storage
Persisting Data in Docker

Sam just shipped their first containerized app and felt great about it. Then a redeploy wiped every signup overnight. Modules 01 and 02 taught Sam to run and build containers, but every database row, every uploaded file, every log line written inside a container disappears when you remove it, unless you understand storage. This module fixes that.

~2h · hands-on heavy builds on Module 01 & 02 volumes · bind mounts · tmpfs Docker 29.5 · CachyOS
0

Where we left off

In Module 01 you learned that a running container gets a thin, writable layer on top of the read-only image layers. In Module 02 you learned how to bake things into the image at build time with a Dockerfile. There is, however, a category of data that you simply cannot bake into an image: runtime state. A database writing rows, an app saving uploads, a service writing logs. All of this happens while the container is running, long after build time. This is exactly the gap Sam fell into.

Writable layer

Every container gets one at startup. Writes go here by default. It is tied to the container lifecycle and disappears on docker rm.

Persistent data

Database files, user uploads, config state. This data needs to outlive any single container. Volumes and bind mounts give it a home outside the writable layer.

Three mount types

Docker gives you three tools: named volumes, bind mounts, and tmpfs. Each solves a different problem. This module is how you pick the right one.

Why this module matters

"How would you run a database in Docker without losing data?" is the core question this module answers. Every real production Docker setup uses volumes. You cannot skip this and claim to know Docker.

1

The problem: why container filesystems are disposable

When Docker starts a container, it stacks the read-only image layers and then adds a fresh, empty writable layer on top. The container writes everything new into that writable layer: new files, modified configs, database rows. The image layers underneath are never touched. This is where Sam's signups were going.

Here is what the storage stack looks like at runtime:

Writable container layercreated fresh at docker run, destroyed at docker rm
COPY . . (your app code)read-only image layer
RUN pip install ...read-only image layer
FROM python:3.12-slimread-only base image layer

When you run docker rm, Docker removes the container and destroys its writable layer. Everything the running process wrote in there is gone. The image is unaffected. You can create a new container from it, which gets a fresh empty writable layer.

The database scenario

Imagine you spin up a Postgres container without any storage configuration. You create a database, add tables, insert rows. Then you run docker rm. Every single row you ever inserted is gone. When you run a new Postgres container from the same image, it starts completely empty.

terminal: the data-loss scenario (no volume)
# start postgres without any volume
docker run -d --name db postgres:16

# connect, create a table, insert a row...
docker exec -it db psql -U postgres

# remove the container
docker rm -f db

# start a fresh container from the same image
docker run -d --name db postgres:16

# all your data is gone: fresh empty database
# (this is the problem modules solve)
This is not a bug, it is by design

Containers are meant to be ephemeral and replaceable. The writable layer is intentionally thrown away so you can safely delete and recreate containers. The problem only arises when you put stateful data in a place designed for stateless operation. The solution is to move that data outside the container using a mount.

There is also a secondary problem: writing through the Union Filesystem (the overlay driver that manages those stacked layers) is slower than writing to native host storage. For I/O-heavy workloads like databases, you want data written directly to host disk, not through the overlay. Volumes solve this too. They bypass the Union Filesystem entirely.

2

The three mount types

Docker provides three distinct ways to give a container access to storage that lives outside its writable layer. Understanding when to reach for each one is the core of this module.

TypeWhere data livesManaged byBest for
Named volume Under /var/lib/docker/volumes/ on the host, in a Docker-managed directory Docker Persistent app data (database files, app state). The default recommendation for production.
Bind mount Any absolute path you specify on the host, e.g. /home/user/myproject You (the user) Development: mounting source code so edits on the host reflect live in the container.
tmpfs In host RAM only, never touches disk Docker (in-memory) Sensitive ephemeral data (secrets, scratch space) that must never be written to disk.

Named volumes

Docker creates and manages the storage location. You just give it a name. Works on any platform. Easiest to back up and migrate. Preferred for production data.

Bind mounts

You pick the exact host path. The container sees that directory directly. Ideal for dev workflows where you edit files on the host and want the container to pick up changes instantly.

tmpfs mounts

Stored in memory only. Automatically wiped when the container stops. Use for credentials or intermediate scratch data that must never land on disk, not even swap.

The decision heuristic

Ask yourself three questions: (1) Does this data need to survive container removal? Use a named volume. (2) Do I need to edit these files from my host (dev workflow, config files)? Use a bind mount. (3) Is this sensitive data that must never hit disk (token cache, session scratch)? Use tmpfs.

3

The volume CLI

Before you can use a named volume in a container, you need to understand the commands that manage the volume lifecycle. You can also let Docker create a volume implicitly on docker run, but knowing the explicit commands gives you control and insight into what is happening.

docker volume: the full lifecycle
# create a named volume
docker volume create pgdata

# list all volumes on this host
docker volume ls
DRIVER    VOLUME NAME
local     pgdata

# inspect a volume (see where it actually lives on the host)
docker volume inspect pgdata
[
  {
    "Name": "pgdata",
    "Driver": "local",
    "Mountpoint": "/var/lib/docker/volumes/pgdata/_data",
    ...
  }
]

# remove a specific volume (only if no container uses it)
docker volume rm pgdata

# remove ALL volumes not attached to a container
docker volume prune

Attaching a volume: -v vs --mount

There are two syntaxes for mounting a volume when you run a container. Both do the same thing; --mount is more verbose but makes the intent explicit and catches typos as errors rather than silently creating bind mounts. Docker's own documentation recommends --mount for new usage.

-v syntax (short, older)
# format: -v volume_name:/path/inside/container
docker run -v pgdata:/var/lib/postgresql/data postgres:16
--mount syntax (explicit, preferred)
# format: --mount type=volume,source=name,target=/path
docker run \
  --mount type=volume,source=pgdata,target=/var/lib/postgresql/data \
  postgres:16

With the --mount syntax, if you mistype the volume name or the path, Docker returns an error immediately. With -v, if Docker cannot find an existing volume named pgdata, it silently creates a new one. That can produce confusing "where did my data go?" situations in dev, the same trap that bit Sam.

docker run  -v  pgdata:/var/lib/postgresql/data  postgres:16 command    flag    volume name    mount point inside container    image
Volume created implicitly

If you reference a volume name in -v or --mount that does not yet exist, Docker creates it automatically. You do not need to run docker volume create first. Explicit creation is useful when you want to inspect or pre-configure a volume before attaching it.

4

Bind mounts in practice

A bind mount takes a directory or file from your host machine and maps it directly into a path inside the container. The container and the host see the exact same filesystem location, so writes from either side are immediately visible to the other. This is what makes bind mounts the natural choice for development workflows.

The live-reload dev pattern

Without bind mounts, every code change during development means rebuilding the image, which is slow. With a bind mount, you mount your source code directory into the container and any edit you save in your editor is immediately visible inside the running container, no rebuild needed.

dev workflow: bind-mount source code
# mount the current directory into /app inside the container
# $(pwd) expands to the absolute path of your current folder
docker run -it -v $(pwd):/app -p 8000:8000 myapp:dev

# same thing using --mount (explicit path)
docker run -it \
  --mount type=bind,source=$(pwd),target=/app \
  -p 8000:8000 myapp:dev

Now when your app inside the container uses a file-watcher (like Nodemon for Node.js, or Flask's debug reloader), it will automatically pick up the changes you save on your host.

Read-only bind mounts

Sometimes you want the container to be able to read a host file (a config file, a certificate) but you do not want it to be able to modify it. Appending :ro to the -v syntax, or adding readonly to --mount, makes the mount read-only from the container's perspective.

read-only bind mounts
# :ro makes the mount read-only inside the container
docker run -v $(pwd)/config:/app/config:ro myapp

# --mount equivalent
docker run \
  --mount type=bind,source=$(pwd)/config,target=/app/config,readonly \
  myapp

# if the container tries to write, it gets permission denied
docker exec myapp touch /app/config/newfile
touch: cannot touch '/app/config/newfile': Read-only file system
Bind mounts are host-path dependent

Bind mounts reference absolute paths on your specific host machine. That means a docker run command with a bind mount is not portable. It works on your machine but will break on another machine where that path doesn't exist, or when the paths differ. Named volumes are portable; bind mounts are not. This is why you use bind mounts only for dev, not for production deployments.

5

Volumes in practice: the Postgres example

The canonical, motivating example for named volumes is a database. Postgres stores all its data files in /var/lib/postgresql/data inside the container. If you mount a named volume there, the data files live in Docker-managed storage on the host, surviving container removal and recreation. This is the fix Sam needed.

postgres with a named volume: data survives docker rm
# step 1: run postgres, mounting the named volume pgdata
docker run -d \
  --name mydb \
  -e POSTGRES_PASSWORD=secret \
  -v pgdata:/var/lib/postgresql/data \
  postgres:16

# step 2: connect and create a table + insert a row
docker exec -it mydb psql -U postgres
postgres=# CREATE TABLE users (id serial, name text);
postgres=# INSERT INTO users (name) VALUES ('Sahil');
postgres=# \q

# step 3: remove the container completely
docker rm -f mydb

# step 4: create a brand new container using the SAME volume
docker run -d \
  --name mydb2 \
  -e POSTGRES_PASSWORD=secret \
  -v pgdata:/var/lib/postgresql/data \
  postgres:16

# step 5: your data is still there
docker exec -it mydb2 psql -U postgres
postgres=# SELECT * FROM users;
 id | name
----+-------
  1 | Sahil
(1 row)

This is the key insight: the volume pgdata is a completely separate object from the container. You can docker rm the container as many times as you like. The volume (and the data in it) is untouched until you explicitly run docker volume rm pgdata. Sam reran the demo three times just to believe the rows stayed.

Container lifecycle vs volume lifecycle

These are independent. Containers are ephemeral: create, run, stop, remove, recreate. Volumes are persistent: they outlive any individual container, can be attached to a new container, and must be explicitly deleted. Never confuse "removing a container" with "deleting its data". The data lives in the volume.

Sharing a volume between containers

Multiple containers can mount the same named volume at the same time. This is how you share files between a web server and a sidecar container, for instance. Be aware that concurrent writes without application-level coordination can lead to corruption. Most databases handle their own locking, but arbitrary shared volumes need careful thought.

sharing a volume between two containers
# container A writes to the shared volume
docker run -d --name writer -v shared:/data writer-image

# container B reads from the same volume
docker run -d --name reader -v shared:/data reader-image
6

Permissions & gotchas

Storage in Docker has some non-obvious behaviour that trips up almost everyone the first time. Knowing these going in saves you hours of debugging.

The UID mismatch problem with bind mounts

When a container process runs as a non-root user (which is the right thing to do, as you learned in Module 07), and you bind-mount a host directory into the container, the container's process might not have permission to write to that directory. Why? Because filesystem permissions are based on numeric user IDs (UIDs), not names. A user named appuser with UID 1001 inside the container and a user named sahil with UID 1000 on the host are different users as far as the filesystem is concerned.

diagnosing and fixing UID mismatch
# check what UID the container process runs as
docker exec myapp id
uid=1001(appuser) gid=1001(appuser)

# check the host directory's owner
ls -la ./myproject
drwxr-xr-x 1000 sahil sahil ...

# fix option 1: chown the host directory to match container UID
sudo chown -R 1001:1001 ./myproject

# fix option 2: run the container as your host UID
docker run -u $(id -u):$(id -g) -v $(pwd):/app myapp

Anonymous volumes: the silent surprise

Some images declare a VOLUME instruction in their Dockerfile (Postgres does this, for example). When you run such an image without specifying a mount, Docker automatically creates an anonymous volume, a named volume with a long random UUID as its name. This ensures the data directory is on a proper volume rather than the writable layer, which is thoughtful.

The surprise comes when you run docker volume ls and see a list of cryptic UUIDs that you didn't create. These are anonymous volumes. They are easy to orphan. When you docker rm the container, the anonymous volume stays behind unless you run docker rm -v (which deletes associated anonymous volumes together with the container). Named volumes are never deleted by docker rm -v, only anonymous ones.

anonymous volumes: listing and cleaning up
# run postgres without specifying a volume
docker run -d --name db postgres:16

# a random-UUID volume was created automatically
docker volume ls
DRIVER    VOLUME NAME
local     a1b2c3d4e5f6...  (anonymous)

# remove the container AND its anonymous volume together
docker rm -v db

# clean up all orphaned volumes at once
docker volume prune

Volume mount over a populated image directory

This is one of the most confusing behaviours, and it applies differently to volumes vs bind mounts. When you mount a named volume over a directory that already has files in the image at that path, Docker copies the image's existing files into the volume on the first use (if the volume is empty). This is why Postgres works at all: the image's pre-initialized data directory gets seeded into the volume the first time. On subsequent runs with the same volume, Docker uses whatever is in the volume (your actual data), not the image directory.

With a bind mount, this seeding does NOT happen. A bind mount over a populated image directory hides the image's files entirely, showing you only the host directory's contents. If that host directory is empty, the container sees an empty directory, even if the image had files there. This is a common source of confusion when bind-mounting over /app in an image that had files copied there at build time.

Volume vs bind mount: the mount-over behaviour

Named volume over populated dir: on first use (empty volume), Docker seeds the volume from the image directory. Your data persists there.
Bind mount over populated dir: image directory is completely hidden. You see only the host path. Empty host dir = empty container dir. No seeding ever happens.

7

Backup & inspecting volume data

Because named volumes live under /var/lib/docker/volumes/ (a path owned by root and managed by Docker), you cannot just cp from there directly in normal usage. The standard pattern is to use a throwaway container that mounts the volume and runs a tar backup.

Backing up a volume

backup a volume to a tar archive on your host
# spin up a throwaway container that:
#   - mounts the volume to /data (read-only)
#   - mounts the current host dir to /backup
#   - tars /data into /backup/pgdata-backup.tar.gz
#   - then exits and is removed (--rm)

docker run --rm \
  -v pgdata:/data:ro \
  -v $(pwd):/backup \
  alpine \
  tar czf /backup/pgdata-backup.tar.gz -C /data .

# the backup file now exists on your host
ls -lh pgdata-backup.tar.gz

Restoring a volume from backup

restore a volume from a tar archive
# create a fresh target volume
docker volume create pgdata-restored

# extract the tar into the new volume
docker run --rm \
  -v pgdata-restored:/data \
  -v $(pwd):/backup \
  alpine \
  tar xzf /backup/pgdata-backup.tar.gz -C /data

Inspecting volume metadata

docker volume inspect
docker volume inspect pgdata
[
  {
    "CreatedAt": "2026-06-01T10:00:00Z",
    "Driver": "local",
    "Labels": {},
    "Mountpoint": "/var/lib/docker/volumes/pgdata/_data",
    "Name": "pgdata",
    "Options": {},
    "Scope": "local"
  }
]

# the Mountpoint is where data lives on the host
# you can ls it as root: sudo ls /var/lib/docker/volumes/pgdata/_data
The throwaway container pattern

The pattern of running a minimal image (like alpine) with --rm to perform a one-off operation on a volume is fundamental to Docker workflows. The same approach is used to migrate data, seed databases, convert file formats, or inspect volume contents without modifying them. Remember it, because it comes up frequently.

8

Hands-on: do this now

Reading about data persistence means nothing until you watch a docker rm not destroy your data with your own eyes. This checklist is your minimum before moving on.

  • Create a named volume with docker volume create mydata. Run docker volume ls to see it. Run docker volume inspect mydata to find where it lives on the host.
  • Run a Postgres container with -v mydata:/var/lib/postgresql/data. Connect with docker exec and psql. Create a table. Insert a row. Exit.
  • Run docker rm -f on that container. Confirm it is gone with docker ps -a. Confirm the volume is still there with docker volume ls.
  • Create a new container from the same image, attaching the same volume. Connect to Postgres and run SELECT * FROM your_table;. Verify your row is still there. This is the moment that makes volumes click.
  • Make a folder on your host with a simple text file in it. Bind-mount it into an alpine container. Edit the file on your host and immediately cat it inside the container, then confirm the change appears live.
  • Add :ro to that bind mount. Try to write to the file from inside the container. You should see Read-only file system. This is the safety guard for config files.
  • Use the throwaway-container pattern to back up your mydata volume to a .tar.gz file on your host. Confirm the tar file exists and has a non-zero size.
  • Run docker volume prune after removing all containers that use mydata. Understand which volumes it removes and which it leaves. Try to remove a volume that a running container is still using, and note the error message.
Sketch before you type

Before running each command, write on paper what you expect to happen. "I expect the row to still be there because the volume persists independently." Then run it and see if you were right. This active prediction habit is what converts reading into actual understanding, and it is what real understanding looks like. It is also how Sam went from losing data to trusting the deploy.