Skip to content
Huseyin Babal
Go back

A container is a process: seven checks you can run on your own machine

This post is the short version of my video. Watch it here: Docker Explained in 8 Minutes

The video draws the model on a board: a container is a Linux process, namespaces decide what it sees, cgroups decide what it gets, and its files come from stacked image layers. This page is the lab sheet that goes with it. Each section is one claim and the command that lets you check it, with the output I got on Docker Engine 29.2 (cgroup v2, overlay2).

Two of the checks are not in the video at all: comparing namespace IDs by hand, and watching the kernel kill a container that goes over its memory limit.

The claims, and where to look

ClaimWhere the evidence is
A container is an ordinary processps on the host
It lives in its own namespaces/proc/<pid>/ns
Its limits are files in a cgroup/sys/fs/cgroup/.../memory.max, cpu.max
Go over the memory limit and the kernel kills itexit code 137, OOMKilled
Its filesystem is a stack of layersdocker inspect, GraphDriver
Changes land in one writable layerdocker diff
Build steps are cached in orderdocker build, the CACHED lines

First: where is “the host”?

On Linux, the host is your machine. On macOS and Windows it is not: Docker Desktop runs the engine inside a small Linux VM, so ps in your own terminal will never show a container’s process.

You can still get a host view. This starts a throwaway container that shares the VM’s process table:

docker run --rm --privileged --pid=host alpine ps -eo pid,user,args

I use that prefix for every “host” command below. On a Linux machine you can drop it and run the inner command directly. It is a privileged container, so keep it to your own laptop.

1. It is a process

docker run -d --name web nginx:alpine
docker run --rm --privileged --pid=host alpine ps -eo pid,user,args | grep 'nginx: master'
  345 root     nginx: master process nginx -g daemon off;output

PID 345, owned by root, sitting in the host’s process list like anything else. Docker will tell you the same number:

docker inspect -f '{{.State.Pid}}' web
345output

Inside the container, the same process has a different number:

docker exec web ps
PID   USER     TIME  COMMAND
    1 root      0:00 nginx: master process nginx -g daemon off;
   30 nginx     0:00 nginx: worker process
   …  (one worker per CPU core)output

345 outside, 1 inside. Nothing was copied or emulated; the kernel keeps two numbering schemes for one process.

2. Namespaces are just IDs you can compare

Every process has a folder, /proc/<pid>/ns, with one link per namespace. Two processes that show the same ID share that namespace. Here is PID 1 of the host next to our nginx:

docker run --rm --privileged --pid=host alpine sh -c 'ls -l /proc/1/ns /proc/345/ns'
NamespaceHost PID 1nginx (PID 345)
pid40265318364026532821separate
net40265318404026532823separate
mnt40265318414026532818separate
uts40265318384026532819separate
ipc40265318394026532820separate
cgroup40265318354026532822separate
user40265318374026531837shared
time40265318344026531834shared

Six of the eight are private. The two shared ones are worth knowing about:

You can also step into a single namespace of a running container. This enters only the network namespace and asks for its address:

docker run --rm --privileged --pid=host alpine nsenter -t 345 -n ip -4 addr show eth0
    inet 172.17.0.3/16 brd 172.17.255.255 scope global eth0output

That is the private address on Docker’s default bridge. The hostname works the same way through the UTS namespace: docker exec web hostname prints the container ID, 845c25be4c97, not the name of the machine.

3. Limits are plain files

docker run -d --name limited --memory=256m --cpus=0.5 nginx:alpine
CG=/sys/fs/cgroup/docker/$(docker inspect -f '{{.Id}}' limited)
cat $CG/memory.max
cat $CG/cpu.max
268435456
50000 100000output

268,435,456 bytes is 256 MiB. The second line reads “50,000 microseconds of CPU in every 100,000”, which is half a core.

The path depends on the cgroup driver. Mine is cgroupfs; many Linux distributions use systemd, and there the folder is /sys/fs/cgroup/system.slice/docker-<id>.scope. One command tells you what you have:

docker info -f '{{.Driver}} / cgroup v{{.CgroupVersion}} / {{.CgroupDriver}}'
overlay2 / cgroup v2 / cgroupfsoutput

4. Watch the limit being enforced

A limit you never hit is easy to forget. This container gets 64 MB and then tries to read an endless stream into memory:

docker run --name oom --memory=64m alpine sh -c 'tail /dev/zero'; echo "exit code: $?"
docker inspect -f '{{.State.OOMKilled}}' oom
exit code: 137
trueoutput

137 is 128 + 9: the process was ended with signal 9 by the kernel’s out-of-memory killer. If a service of yours keeps restarting with exit code 137, this flag is the first thing to check.

5. The filesystem is a stack

Docker records how it assembled the container’s root filesystem:

docker inspect -f '{{json .GraphDriver.Data}}' web

The interesting keys are:

KeyMeaning
LowerDirthe image layers, read-only, joined with :
UpperDirthis container’s own writable layer
MergedDirwhat the process sees as /

For nginx:alpine my LowerDir listed more than six folders under /var/lib/docker/overlay2/. Start a second container from the same image and it gets the same lower folders and a fresh UpperDir. That sharing is the reason ten containers of one image do not cost ten times the disk.

6. Changes go into one thin layer

docker exec web sh -c 'echo hello > /tmp/note.txt && rm /etc/nginx/conf.d/default.conf'
docker diff web
C /etc/nginx/conf.d
D /etc/nginx/conf.d/default.conf
A /var/cache/nginx/client_temp
C /tmp
A /tmp/note.txt
A /run/nginx.pidoutput

(Shortened; nginx creates a few more cache folders.) A is added, C is changed, D is deleted. Two details:

7. The build cache follows the order of your Dockerfile

FROM node:22-alpine
WORKDIR /app
COPY package.json ./
RUN npm install --omit=dev
COPY . .
CMD ["node", "server.js"]

docker history shows what each instruction added:

CREATED BY                                      SIZE
CMD ["node" "server.js"]                        0B
COPY . . # buildkit                             333B
RUN /bin/sh -c npm install --omit=dev # buil…   8.83MB
COPY package.json ./ # buildkit                 87B
WORKDIR /app                                    0Boutput

Now change one line in server.js and build again:

docker build -t hello . 2>&1 | grep -E '\[[0-9]/5\]|CACHED'
#6 [2/5] WORKDIR /app
#6 CACHED
#7 [3/5] COPY package.json ./
#7 CACHED
#8 [4/5] RUN npm install --omit=dev
#8 CACHED
#9 [5/5] COPY . .output

The 8.83 MB install step came from the cache; only the 333-byte copy ran. Swap the two COPY lines and every code change reinstalls your dependencies.

Things that quietly break the cache:

What these checks do not show

Clean up

docker rm -f web limited oom
docker rmi hello

Sources


Share this post:

Previous Post
1000 trick questions for four decision models: what the numbers hide
Next Post
Java garbage collection in its own log: eight checks you can run