Module 03: Volumes & Storage
Persisting Data in Docker
Sam just shipped their first containerized app and felt great about it. Then a redeploy wiped every signup overnight. Modules 01 and 02 taught Sam to run and build containers, but every database row, every uploaded file, every log line written inside a container disappears when you remove it, unless you understand storage. This module fixes that.
Where we left off
In Module 01 you learned that a running container gets a thin, writable layer on top of the read-only image layers. In Module 02 you learned how to bake things into the image at build time with a Dockerfile. There is, however, a category of data that you simply cannot bake into an image: runtime state. A database writing rows, an app saving uploads, a service writing logs. All of this happens while the container is running, long after build time. This is exactly the gap Sam fell into.
Writable layer
Every container gets one at startup. Writes go here by default. It is tied to the container lifecycle and disappears on docker rm.
Persistent data
Database files, user uploads, config state. This data needs to outlive any single container. Volumes and bind mounts give it a home outside the writable layer.
Three mount types
Docker gives you three tools: named volumes, bind mounts, and tmpfs. Each solves a different problem. This module is how you pick the right one.
"How would you run a database in Docker without losing data?" is the core question this module answers. Every real production Docker setup uses volumes. You cannot skip this and claim to know Docker.
The problem: why container filesystems are disposable
When Docker starts a container, it stacks the read-only image layers and then adds a fresh, empty writable layer on top. The container writes everything new into that writable layer: new files, modified configs, database rows. The image layers underneath are never touched. This is where Sam's signups were going.
Here is what the storage stack looks like at runtime:
When you run docker rm, Docker removes the container and destroys its writable layer. Everything the running process wrote in there is gone. The image is unaffected. You can create a new container from it, which gets a fresh empty writable layer.
The database scenario
Imagine you spin up a Postgres container without any storage configuration. You create a database, add tables, insert rows. Then you run docker rm. Every single row you ever inserted is gone. When you run a new Postgres container from the same image, it starts completely empty.
# start postgres without any volume
docker run -d --name db postgres:16
# connect, create a table, insert a row...
docker exec -it db psql -U postgres
# remove the container
docker rm -f db
# start a fresh container from the same image
docker run -d --name db postgres:16
# all your data is gone: fresh empty database
# (this is the problem modules solve)
Containers are meant to be ephemeral and replaceable. The writable layer is intentionally thrown away so you can safely delete and recreate containers. The problem only arises when you put stateful data in a place designed for stateless operation. The solution is to move that data outside the container using a mount.
There is also a secondary problem: writing through the Union Filesystem (the overlay driver that manages those stacked layers) is slower than writing to native host storage. For I/O-heavy workloads like databases, you want data written directly to host disk, not through the overlay. Volumes solve this too. They bypass the Union Filesystem entirely.
The three mount types
Docker provides three distinct ways to give a container access to storage that lives outside its writable layer. Understanding when to reach for each one is the core of this module.
| Type | Where data lives | Managed by | Best for |
|---|---|---|---|
| Named volume | Under /var/lib/docker/volumes/ on the host, in a Docker-managed directory |
Docker | Persistent app data (database files, app state). The default recommendation for production. |
| Bind mount | Any absolute path you specify on the host, e.g. /home/user/myproject |
You (the user) | Development: mounting source code so edits on the host reflect live in the container. |
| tmpfs | In host RAM only, never touches disk | Docker (in-memory) | Sensitive ephemeral data (secrets, scratch space) that must never be written to disk. |
Named volumes
Docker creates and manages the storage location. You just give it a name. Works on any platform. Easiest to back up and migrate. Preferred for production data.
Bind mounts
You pick the exact host path. The container sees that directory directly. Ideal for dev workflows where you edit files on the host and want the container to pick up changes instantly.
tmpfs mounts
Stored in memory only. Automatically wiped when the container stops. Use for credentials or intermediate scratch data that must never land on disk, not even swap.
Ask yourself three questions: (1) Does this data need to survive container removal? Use a named volume. (2) Do I need to edit these files from my host (dev workflow, config files)? Use a bind mount. (3) Is this sensitive data that must never hit disk (token cache, session scratch)? Use tmpfs.
The volume CLI
Before you can use a named volume in a container, you need to understand the commands that manage the volume lifecycle. You can also let Docker create a volume implicitly on docker run, but knowing the explicit commands gives you control and insight into what is happening.
# create a named volume
docker volume create pgdata
# list all volumes on this host
docker volume ls
DRIVER VOLUME NAME
local pgdata
# inspect a volume (see where it actually lives on the host)
docker volume inspect pgdata
[
{
"Name": "pgdata",
"Driver": "local",
"Mountpoint": "/var/lib/docker/volumes/pgdata/_data",
...
}
]
# remove a specific volume (only if no container uses it)
docker volume rm pgdata
# remove ALL volumes not attached to a container
docker volume prune
Attaching a volume: -v vs --mount
There are two syntaxes for mounting a volume when you run a container. Both do the same thing; --mount is more verbose but makes the intent explicit and catches typos as errors rather than silently creating bind mounts. Docker's own documentation recommends --mount for new usage.
# format: -v volume_name:/path/inside/container
docker run -v pgdata:/var/lib/postgresql/data postgres:16
# format: --mount type=volume,source=name,target=/path
docker run \
--mount type=volume,source=pgdata,target=/var/lib/postgresql/data \
postgres:16
With the --mount syntax, if you mistype the volume name or the path, Docker returns an error immediately. With -v, if Docker cannot find an existing volume named pgdata, it silently creates a new one. That can produce confusing "where did my data go?" situations in dev, the same trap that bit Sam.
If you reference a volume name in -v or --mount that does not yet exist, Docker creates it automatically. You do not need to run docker volume create first. Explicit creation is useful when you want to inspect or pre-configure a volume before attaching it.
Bind mounts in practice
A bind mount takes a directory or file from your host machine and maps it directly into a path inside the container. The container and the host see the exact same filesystem location, so writes from either side are immediately visible to the other. This is what makes bind mounts the natural choice for development workflows.
The live-reload dev pattern
Without bind mounts, every code change during development means rebuilding the image, which is slow. With a bind mount, you mount your source code directory into the container and any edit you save in your editor is immediately visible inside the running container, no rebuild needed.
# mount the current directory into /app inside the container
# $(pwd) expands to the absolute path of your current folder
docker run -it -v $(pwd):/app -p 8000:8000 myapp:dev
# same thing using --mount (explicit path)
docker run -it \
--mount type=bind,source=$(pwd),target=/app \
-p 8000:8000 myapp:dev
Now when your app inside the container uses a file-watcher (like Nodemon for Node.js, or Flask's debug reloader), it will automatically pick up the changes you save on your host.
Read-only bind mounts
Sometimes you want the container to be able to read a host file (a config file, a certificate) but you do not want it to be able to modify it. Appending :ro to the -v syntax, or adding readonly to --mount, makes the mount read-only from the container's perspective.
# :ro makes the mount read-only inside the container
docker run -v $(pwd)/config:/app/config:ro myapp
# --mount equivalent
docker run \
--mount type=bind,source=$(pwd)/config,target=/app/config,readonly \
myapp
# if the container tries to write, it gets permission denied
docker exec myapp touch /app/config/newfile
touch: cannot touch '/app/config/newfile': Read-only file system
Bind mounts reference absolute paths on your specific host machine. That means a docker run command with a bind mount is not portable. It works on your machine but will break on another machine where that path doesn't exist, or when the paths differ. Named volumes are portable; bind mounts are not. This is why you use bind mounts only for dev, not for production deployments.
Volumes in practice: the Postgres example
The canonical, motivating example for named volumes is a database. Postgres stores all its data files in /var/lib/postgresql/data inside the container. If you mount a named volume there, the data files live in Docker-managed storage on the host, surviving container removal and recreation. This is the fix Sam needed.
# step 1: run postgres, mounting the named volume pgdata
docker run -d \
--name mydb \
-e POSTGRES_PASSWORD=secret \
-v pgdata:/var/lib/postgresql/data \
postgres:16
# step 2: connect and create a table + insert a row
docker exec -it mydb psql -U postgres
postgres=# CREATE TABLE users (id serial, name text);
postgres=# INSERT INTO users (name) VALUES ('Sahil');
postgres=# \q
# step 3: remove the container completely
docker rm -f mydb
# step 4: create a brand new container using the SAME volume
docker run -d \
--name mydb2 \
-e POSTGRES_PASSWORD=secret \
-v pgdata:/var/lib/postgresql/data \
postgres:16
# step 5: your data is still there
docker exec -it mydb2 psql -U postgres
postgres=# SELECT * FROM users;
id | name
----+-------
1 | Sahil
(1 row)
This is the key insight: the volume pgdata is a completely separate object from the container. You can docker rm the container as many times as you like. The volume (and the data in it) is untouched until you explicitly run docker volume rm pgdata. Sam reran the demo three times just to believe the rows stayed.
These are independent. Containers are ephemeral: create, run, stop, remove, recreate. Volumes are persistent: they outlive any individual container, can be attached to a new container, and must be explicitly deleted. Never confuse "removing a container" with "deleting its data". The data lives in the volume.
Sharing a volume between containers
Multiple containers can mount the same named volume at the same time. This is how you share files between a web server and a sidecar container, for instance. Be aware that concurrent writes without application-level coordination can lead to corruption. Most databases handle their own locking, but arbitrary shared volumes need careful thought.
# container A writes to the shared volume
docker run -d --name writer -v shared:/data writer-image
# container B reads from the same volume
docker run -d --name reader -v shared:/data reader-image
Permissions & gotchas
Storage in Docker has some non-obvious behaviour that trips up almost everyone the first time. Knowing these going in saves you hours of debugging.
The UID mismatch problem with bind mounts
When a container process runs as a non-root user (which is the right thing to do, as you learned in Module 07), and you bind-mount a host directory into the container, the container's process might not have permission to write to that directory. Why? Because filesystem permissions are based on numeric user IDs (UIDs), not names. A user named appuser with UID 1001 inside the container and a user named sahil with UID 1000 on the host are different users as far as the filesystem is concerned.
# check what UID the container process runs as
docker exec myapp id
uid=1001(appuser) gid=1001(appuser)
# check the host directory's owner
ls -la ./myproject
drwxr-xr-x 1000 sahil sahil ...
# fix option 1: chown the host directory to match container UID
sudo chown -R 1001:1001 ./myproject
# fix option 2: run the container as your host UID
docker run -u $(id -u):$(id -g) -v $(pwd):/app myapp
Anonymous volumes: the silent surprise
Some images declare a VOLUME instruction in their Dockerfile (Postgres does this, for example). When you run such an image without specifying a mount, Docker automatically creates an anonymous volume, a named volume with a long random UUID as its name. This ensures the data directory is on a proper volume rather than the writable layer, which is thoughtful.
The surprise comes when you run docker volume ls and see a list of cryptic UUIDs that you didn't create. These are anonymous volumes. They are easy to orphan. When you docker rm the container, the anonymous volume stays behind unless you run docker rm -v (which deletes associated anonymous volumes together with the container). Named volumes are never deleted by docker rm -v, only anonymous ones.
# run postgres without specifying a volume
docker run -d --name db postgres:16
# a random-UUID volume was created automatically
docker volume ls
DRIVER VOLUME NAME
local a1b2c3d4e5f6... (anonymous)
# remove the container AND its anonymous volume together
docker rm -v db
# clean up all orphaned volumes at once
docker volume prune
Volume mount over a populated image directory
This is one of the most confusing behaviours, and it applies differently to volumes vs bind mounts. When you mount a named volume over a directory that already has files in the image at that path, Docker copies the image's existing files into the volume on the first use (if the volume is empty). This is why Postgres works at all: the image's pre-initialized data directory gets seeded into the volume the first time. On subsequent runs with the same volume, Docker uses whatever is in the volume (your actual data), not the image directory.
With a bind mount, this seeding does NOT happen. A bind mount over a populated image directory hides the image's files entirely, showing you only the host directory's contents. If that host directory is empty, the container sees an empty directory, even if the image had files there. This is a common source of confusion when bind-mounting over /app in an image that had files copied there at build time.
Named volume over populated dir: on first use (empty volume), Docker seeds the volume from the image directory. Your data persists there.
Bind mount over populated dir: image directory is completely hidden. You see only the host path. Empty host dir = empty container dir. No seeding ever happens.
Backup & inspecting volume data
Because named volumes live under /var/lib/docker/volumes/ (a path owned by root and managed by Docker), you cannot just cp from there directly in normal usage. The standard pattern is to use a throwaway container that mounts the volume and runs a tar backup.
Backing up a volume
# spin up a throwaway container that:
# - mounts the volume to /data (read-only)
# - mounts the current host dir to /backup
# - tars /data into /backup/pgdata-backup.tar.gz
# - then exits and is removed (--rm)
docker run --rm \
-v pgdata:/data:ro \
-v $(pwd):/backup \
alpine \
tar czf /backup/pgdata-backup.tar.gz -C /data .
# the backup file now exists on your host
ls -lh pgdata-backup.tar.gz
Restoring a volume from backup
# create a fresh target volume
docker volume create pgdata-restored
# extract the tar into the new volume
docker run --rm \
-v pgdata-restored:/data \
-v $(pwd):/backup \
alpine \
tar xzf /backup/pgdata-backup.tar.gz -C /data
Inspecting volume metadata
docker volume inspect pgdata
[
{
"CreatedAt": "2026-06-01T10:00:00Z",
"Driver": "local",
"Labels": {},
"Mountpoint": "/var/lib/docker/volumes/pgdata/_data",
"Name": "pgdata",
"Options": {},
"Scope": "local"
}
]
# the Mountpoint is where data lives on the host
# you can ls it as root: sudo ls /var/lib/docker/volumes/pgdata/_data
The pattern of running a minimal image (like alpine) with --rm to perform a one-off operation on a volume is fundamental to Docker workflows. The same approach is used to migrate data, seed databases, convert file formats, or inspect volume contents without modifying them. Remember it, because it comes up frequently.
Hands-on: do this now
Reading about data persistence means nothing until you watch a docker rm not destroy your data with your own eyes. This checklist is your minimum before moving on.
- Create a named volume with
docker volume create mydata. Rundocker volume lsto see it. Rundocker volume inspect mydatato find where it lives on the host. - Run a Postgres container with
-v mydata:/var/lib/postgresql/data. Connect withdocker execandpsql. Create a table. Insert a row. Exit. - Run
docker rm -fon that container. Confirm it is gone withdocker ps -a. Confirm the volume is still there withdocker volume ls. - Create a new container from the same image, attaching the same volume. Connect to Postgres and run
SELECT * FROM your_table;. Verify your row is still there. This is the moment that makes volumes click. - Make a folder on your host with a simple text file in it. Bind-mount it into an
alpinecontainer. Edit the file on your host and immediatelycatit inside the container, then confirm the change appears live. - Add
:roto that bind mount. Try to write to the file from inside the container. You should seeRead-only file system. This is the safety guard for config files. - Use the throwaway-container pattern to back up your
mydatavolume to a.tar.gzfile on your host. Confirm the tar file exists and has a non-zero size. - Run
docker volume pruneafter removing all containers that usemydata. Understand which volumes it removes and which it leaves. Try to remove a volume that a running container is still using, and note the error message.
Before running each command, write on paper what you expect to happen. "I expect the row to still be there because the volume persists independently." Then run it and see if you were right. This active prediction habit is what converts reading into actual understanding, and it is what real understanding looks like. It is also how Sam went from losing data to trusting the deploy.