DockernamespacescgroupsOverlayFSmulti-stage-buildsrootless-dockerseccompDocker-Swarmcontainer-securityDevOps
TL;DR When a container is misbehaving and you can't run a shell inside it (no shell available, or it's crashed), nsenter lets you enter its namespaces directly from the host. First find the container's PID: docker inspect --format='{{.State.Pid}}' container_name. Then: nsenter -t $PID -m -u -i -n -p -- /bin/sh. You're now inside the container's namespaces with full host access — you can read its filesystem, examine its network state, and run any host binary.
You know the commands. You've written Dockerfiles. You've done docker run and docker-compose up and called it containerization. But when production falls over at 2am because of a cgroup OOM killer, or your image is 1.8GB pulling across a flaky connection, or someone exploits your root-running container — that's when you find out whether you actually know Docker. This guide takes you all the way in.
Read the Deep Dive ↓ Open the Lab 🐳 Namespaces cgroups Overlay2 Multi-Stage Overlay Net iptables Rootless Seccomp Swarm Table of Contents
Imagine you're debugging a container that's using 100% CPU even though you set --cpus=0.5. You restart it, same problem. You inspect the container — nothing. Then you realize: the constraint was set on the wrong cgroup because your Docker daemon version doesn't support the v2 unified hierarchy. Without understanding the kernel primitives underneath Docker, this takes hours to debug. With that understanding, it takes five minutes.
Docker containers aren't magic — they're a tightly orchestrated combination of six Linux kernel features. Linux namespaces provide isolation: PID namespace so the container has its own process tree (PID 1 inside the container is just another process from the host's perspective), Network namespace so the container gets its own network stack, Mount namespace for an isolated filesystem view, User namespace to map container UIDs to unprivileged host UIDs, and UTS namespace to give containers their own hostname. None of this is Docker-specific — these namespaces existed in the Linux kernel before Docker was written.
Control Groups (cgroups) handle resource accounting and enforcement. cgroups v1 used a hierarchical tree per controller (one tree for CPU, another for memory), creating well-known inconsistencies. cgroups v2 (now default in modern kernels) provides a unified hierarchy where all controllers for a process are in a single tree. When you run docker run --memory=512m --cpus=0.5, Docker creates entries under /sys/fs/cgroup/ and writes limits to the appropriate controller files. The OOM killer is a cgroup mechanism — when a container exceeds its memory limit, the kernel kills processes inside that cgroup.
The storage driver determines how container layers are stacked. OverlayFS2 (the default) uses the kernel's overlay filesystem: each image layer is a directory, and the driver overlays them in a union mount. The lowerdir contains read-only image layers, the upperdir holds the container's writable layer, and the merged directory is what the container actually sees. When you write a file that exists in a lower layer, the overlay driver performs a copy-on-write (CoW) — copying the file to the upper layer before writing. This is why writes to large files in containers can be unexpectedly slow: you're paying the CoW cost on the first write. The dockerd → containerd → runc stack is another critical piece: containerd is the daemon that manages container lifecycle, and runc is the low-level OCI runtime that actually calls the kernel to create namespaces and cgroups.
💡 nsenter: The Expert's Debugging ToolWhen a container is misbehaving and you can't run a shell inside it (no shell available, or it's crashed), nsenter lets you enter its namespaces directly from the host. First find the container's PID: docker inspect --format='{{.State.Pid}}' container_name. Then: nsenter -t $PID -m -u -i -n -p -- /bin/sh. You're now inside the container's namespaces with full host access — you can read its filesystem, examine its network state, and run any host binary. This is the tool that saves production incidents.
# Get container PID
$ CPID=$(docker inspect --format='{{.State.Pid}}' mycontainer)
# View container namespaces (they're symlinks to kernel namespace IDs)
$ ls -la /proc/$CPID/ns/
# lrwxrwxrwx pid -> pid:[4026532297]
# lrwxrwxrwx net -> net:[4026532300] ← container's private network
# View cgroup limits imposed by Docker
$ cat /sys/fs/cgroup/system.slice/docker-$CONTAINERID.scope/memory.max
# 536870912 (512 MB in bytes)
# Inspect OverlayFS2 layer structure
$ docker inspect mycontainer | python3 -c "
import json,sys
d=json.load(sys.stdin)[0]
gd=d['GraphDriver']['Data']
print('LowerDir:', gd['LowerDir'][:80])
print('UpperDir:', gd['UpperDir'])
print('MergedDir:', gd['MergedDir'])
"
# Enter container namespaces from host (no shell required in container)
$ nsenter -t $CPID -m -u -i -n -p -- /bin/sh
A 1.8GB Docker image is a symptom, not a problem. The problems it causes are real: 45-second pull times in CI/CD, $400/month in registry storage costs, a CVE surface area the size of a small country, and on-call engineers waiting three minutes during a rollout while production traffic piles up. Expert Docker engineers build images that pull in under 10 seconds and contain only what's needed to run the process.
Multi-stage builds are the highest-leverage optimization available. The pattern: use a full SDK/compiler image to build your application, then copy only the binary (or compiled artifacts) into a minimal base image. The build tools, source code, test dependencies, and intermediate compilation artifacts never make it into the production image. For a Go service, this typically means going from a 800MB builder stage to a 5MB Distroless final image. For Node.js, you can separate the npm install/build stage from the production runtime stage, keeping devDependencies completely out of production.
Layer caching is where most teams leave performance on the table. Each instruction in a Dockerfile creates a layer. Docker caches layers and only rebuilds from the first instruction that changed. The critical rule: put instructions that change frequently (copying application code) after instructions that change rarely (installing system dependencies). If you COPY . . before RUN npm install, every code change invalidates the npm install cache. Reverse the order — COPY package*.json ., RUN npm install, then COPY . . — and npm install only reruns when dependencies actually change.
Here's the thing most tutorials miss about base image selection: Alpine's reputation for smallest images is increasingly challenged by Distroless images for production. Alpine uses musl libc instead of glibc, which can cause subtle compatibility issues with native modules and C libraries. Google's Distroless images contain only the application runtime and its direct dependencies — no shell, no package manager, no coreutils. They're smaller than Alpine for many workloads and have a dramatically smaller attack surface. The tradeoff: debugging is harder (no shell to exec into), which is fine for production but requires separate debug-variant images for troubleshooting.
✅ Use dive to Audit Every Layer Before PushingThe dive tool (github.com/wagoodman/dive) lets you inspect every layer in a Docker image interactively, showing exactly which files were added/modified/removed and how much each layer contributes to the total image size. Run it in CI with dive --ci --lowestEfficiency=0.9 to fail builds that have less than 90% layer efficiency. Common waste sources: package manager caches not cleaned in the same RUN instruction, build tools included in production layers, .git directories copied in, test data included, and documentation/man pages from apt-get packages.
# ── Stage 1: Builder ────────────────────────────────────── FROM golang:1.22-alpine AS builder WORKDIR /app # Copy dependency files FIRST (cache layer reused unless deps change) COPY go.mod go.sum ./ RUN go mod download # NOW copy source (this layer invalidates on code change) COPY . . # Build static binary (no CGO = no libc dependency) RUN CGO_ENABLED=0 GOOS=linux go build -ldflags="-s -w" -o /app/server . # ── Stage 2: Production (Distroless) ────────────────────── FROM gcr.io/distroless/static-debian12 # Run as non-root user (uid 65532 in distroless) USER nonroot:nonroot # Copy ONLY the compiled binary from builder stage COPY --from=builder --chown=nonroot:nonroot /app/server /server # No SHELL. No package manager. No attack surface. ENTRYPOINT ["/server"] # Result: 5.2MB image vs 800MB+ with full golang image # Build: docker build -t myapp:prod .
Networking is where the gap between Docker users and Docker experts is widest. Understanding why a container can reach another container, why a port forwarding rule works the way it does, and how traffic actually flows from the internet to your application requires understanding the Docker networking model all the way down to iptables.
The bridge network driver (default) creates a virtual switch (docker0) on the host. Containers connected to the same bridge can communicate via their internal IP addresses. When you publish a port (-p 8080:80), Docker adds an iptables DNAT rule that rewrites packets destined for host:8080 to containerIP:80. The DOCKER iptables chain manages these rules. This is why adding a firewall rule to block a published port might not work — you need to add it to the DOCKER-USER chain, which Docker preserves across restarts. The Host driver skips this entirely — the container shares the host's network namespace directly, getting maximum performance at the cost of isolation.
Overlay networks span multiple Docker hosts (typically in a Swarm cluster). They encapsulate container-to-container traffic in VXLAN packets that traverse the host network, making containers on different machines appear to be on the same network. The Docker daemon's embedded DNS server handles service discovery: in a user-defined network, you can reach a container named "api" from another container using just the hostname "api" — Docker's internal DNS resolves container names to their current IP addresses. This is essential because container IPs change on every restart.
Macvlan is the most powerful — and least understood — driver. It assigns a real MAC address to each container and connects it directly to a physical network interface, making containers appear as physical devices on the network. Use this when containers need to be directly reachable on your LAN, or when you need specific MAC addresses for license enforcement. The downside: your switch must support promiscuous mode, and containers on the same host can't communicate with each other through the Macvlan interface (they can via a separate bridge or Macvlan sub-interface).
⚠️ Docker Modifies iptables by Default — This Can Break Your FirewallDocker inserts rules into iptables that allow all published ports to be reachable regardless of your UFW or firewalld rules. If you run docker run -p 5432:5432 postgres, your database is publicly accessible even if UFW says "deny all incoming." The fix: add rules to the DOCKER-USER chain (Docker preserves these): iptables -I DOCKER-USER -p tcp --dport 5432 -j DROP. Or set "iptables": false in Docker daemon config and manage iptables manually — but then you lose automatic port publishing. Always audit published ports with docker ps and verify iptables rules after any Docker update.
# Inspect a network's subnets and connected containers
$ docker network inspect myapp_net --format '{{json .IPAM.Config}}'
# [{"Subnet":"172.20.0.0/16","Gateway":"172.20.0.1"}]
# See all iptables rules Docker manages
$ iptables -L DOCKER -n --line-numbers
$ iptables -L DOCKER-USER -n # ← safe place to add custom rules
# Create overlay network for multi-host (Swarm)
$ docker network create \
--driver overlay \
--subnet 10.10.0.0/24 \
--opt encrypted \ # encrypt inter-node traffic with AES
production_overlay
# Static IP for a container (requires custom subnet)
$ docker network create --subnet 192.168.1.0/24 mynet
$ docker run --network mynet --ip 192.168.1.100 nginx
# Test service discovery (DNS resolution inside user-defined network)
$ docker exec webapp ping -c 1 api # resolves "api" container name
$ docker exec webapp nslookup api # shows Docker's internal DNS: 127.0.0.11
The three Docker storage mechanisms — volumes, bind mounts, and tmpfs — look similar from the outside but have radically different performance characteristics and use cases. Choosing wrong costs you either performance, portability, or data durability.
Docker volumes are managed by the Docker daemon, stored under /var/lib/docker/volumes/, and completely portable. They survive container deletion, can be shared between containers, and can use pluggable storage drivers for cloud-native storage. For production database data, use volumes — they're the safest option. Bind mounts map a specific host directory directly into the container. They're fast (no Docker overhead), but tightly couple the container to the host's directory structure. Use bind mounts in development (for live code reloading) and in CI/CD (to inject build artifacts). Never use bind mounts for production database storage — you're betting your data durability on host filesystem hygiene.
Tmpfs mounts are in-memory filesystems — the fastest possible storage, but ephemeral. Use tmpfs for temporary files that must not touch disk (secrets during processing, session files for in-memory caches, or any sensitive intermediate data). The data vanishes when the container stops. On Linux, tmpfs mounts can also be size-limited to prevent runaway processes from exhausting host RAM.
🚨 Dangling Volumes Are a Silent Storage LeakEvery time you remove a container without -v, its associated volumes remain on disk — orphaned, unnamed, consuming space. These "dangling volumes" accumulate silently on busy Docker hosts until disk space runs out (often at 3am in production). Run docker volume ls -f dangling=true to see them. Clean up safely: docker volume prune removes all dangling volumes. In production, build volume lifecycle management into your container cleanup scripts, and use named volumes rather than anonymous ones — named volumes are easier to audit and intentionally preserve.
# Named volume (production database) — survived container restarts/deletion $ docker volume create --driver local pgdata $ docker run -v pgdata:/var/lib/postgresql/data postgres:16 # Bind mount (dev hot-reload) — fast, host-coupled $ docker run -v $(pwd)/src:/app/src:ro node:20 npm run dev # Tmpfs (in-memory secrets processing) $ docker run --tmpfs /tmp:rw,noexec,nosuid,size=100m myapp # Cloud volume driver (AWS EBS) for stateful production workloads $ docker volume create \ --driver rexray/ebs \ --opt volumetype=gp3 \ --opt size=100 \ production_db_volume # Garbage collection audit $ docker system df # total: images, containers, volumes, cache $ docker volume ls -f dangling=true | wc -l # count orphans $ docker volume prune --filter "label!=keep" # preserve labeled volumes
The single most common Docker security mistake isn't using a vulnerable image or exposing a port — it's running containers as root. By default, a process inside a Docker container runs as UID 0. If that process escapes the container (via a kernel exploit, a misconfigured volume mount, or a privileged operation), it's root on the host. Every container you run as root is a potential host takeover waiting for the right CVE.
Rootless Docker solves this at the daemon level: the entire Docker daemon runs as an unprivileged user, using user namespaces to map container UIDs to unprivileged host UIDs. Even if a container process escapes, it's still running as an unprivileged user on the host. Set this up with dockerd-rootless-setuptool.sh install. At the container level, always add USER nonroot to your Dockerfiles and run with --read-only (plus --tmpfs /tmp for temporary file locations). Capabilities provide fine-grained privilege control: --cap-drop=ALL --cap-add=NET_BIND_SERVICE gives the container only the ability to bind low-numbered ports — nothing else.
Seccomp profiles limit which kernel syscalls a container can make. Docker's default seccomp profile already blocks ~44 syscalls including dangerous ones like ptrace, kexec_load, and clock modification. Custom profiles (JSON format, loaded via --security-opt seccomp=/path/to/profile.json) can restrict this further to only the syscalls your application actually needs. This is your defense-in-depth against container escape via kernel exploits — even if a vulnerability exists, the syscall it requires is blocked.
Secrets management is where most teams still have embarrassing gaps. Environment variables for secrets are readable by any process in the container, logged by process managers, visible in docker inspect, and often end up in logs and error messages. Docker Secrets (in Swarm mode) or external solutions (HashiCorp Vault with a sidecar agent) mount secrets as in-memory tmpfs files accessible only to authorized containers. If you're not in Swarm, use BuildKit's --secret flag for build-time secrets to prevent them from ever appearing in image layers.
Integrate Trivy into every CI pipeline to scan images before they reach a registry. A one-liner: trivy image --exit-code 1 --severity CRITICAL myapp:latest — fails the build if any CRITICAL vulnerabilities are found. For shift-left security, also scan your Dockerfile itself: trivy config Dockerfile. Configure Trivy to ignore accepted risks: trivy image --ignorefile .trivyignore myapp. Schedule weekly scans of all images in your registry to catch newly discovered CVEs in images you built months ago — a critical CVE can appear in a previously-clean image when new vulnerability data is published.
# Maximum hardening for a production container $ docker run \ --read-only \ # immutable filesystem --tmpfs /tmp:noexec,nosuid,size=50m \ # writable temp only --cap-drop=ALL \ # drop ALL capabilities --cap-add=NET_BIND_SERVICE \ # add only what's needed --security-opt no-new-privileges \ # prevent privilege escalation --security-opt seccomp=./seccomp.json \ # custom syscall filter --user 10001:10001 \ # explicit non-root UID --memory=512m --memory-swap=512m \ # no swap (prevent OOM escape) --pids-limit=100 \ # prevent fork bombs --network mynet \ # explicit network (not default bridge) myapp:prod # Enable Docker Content Trust (image signing) $ export DOCKER_CONTENT_TRUST=1 $ docker pull myregistry/myapp:prod # fails if not signed # Vulnerability scanning in CI $ trivy image --exit-code 1 --severity CRITICAL myapp:latest $ trivy config Dockerfile # Dockerfile security checks # Docker Secrets (Swarm mode) $ echo "supersecretpassword" | docker secret create db_password - $ docker service create \ --secret db_password \ # mounted at /run/secrets/db_password myapp:prod
The counterintuitive truth about Kubernetes: if your infrastructure team skips directly from docker run to Kubernetes, they usually don't understand why Kubernetes makes the decisions it does. Docker Compose and Docker Swarm teach you the core orchestration concepts — service definitions, rolling updates, health checks, secret management, service discovery — in a dramatically simpler environment. The engineers who understand Swarm understand Kubernetes faster and debug it more effectively.
Advanced Docker Compose goes far beyond a single docker-compose.yml. The docker-compose.override.yml file is automatically merged with the base file, letting you override specific values for local development without modifying the production configuration. For multiple environments, use explicit override files: docker-compose -f docker-compose.yml -f docker-compose.prod.yml up. Variable substitution with .env files and environment-specific overrides lets you maintain a single service definition with environment-specific resource limits, image tags, replica counts, and secrets.
Docker Swarm transforms a group of Docker hosts into a cluster with a manager/worker node architecture. Swarm handles service scheduling (placing containers across workers based on resource availability), rolling updates with configurable parallelism and rollback triggers, distributed secret management, and overlay networking across hosts. A Swarm service definition is remarkably similar to a Kubernetes Deployment — both define replicas, update strategies, health checks, and resource limits. The key concepts — desired state reconciliation, service mesh DNS, stateful vs. stateless services — transfer directly to Kubernetes.
💡 Use Compose Override Files for Every EnvironmentMaintain three files: docker-compose.yml (base, shared across environments), docker-compose.dev.yml (bind mounts, debug ports, hot reload, no resource limits), and docker-compose.prod.yml (named volumes, resource limits, replica counts, production image tags, health check intervals). Makefile targets wrap the combinations: make dev runs docker-compose -f docker-compose.yml -f docker-compose.dev.yml up. This pattern scales to QA and staging environments without duplicating the entire service definition.
# Base: docker-compose.yml (services, networks, volumes)
# Override: docker-compose.prod.yml (prod-specific values)
version: '3.9'
services:
api:
image: myregistry/api:${IMAGE_TAG:-latest}
deploy:
replicas: 3
update_config:
parallelism: 1 # roll one at a time
delay: 10s
failure_action: rollback
resources:
limits: {cpus: '0.5', memory: 512M}
reservations: {cpus: '0.25', memory: 256M}
secrets: [db_password, api_key]
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost/health"]
interval: 30s
retries: 3
environment:
DATABASE_URL: "postgresql://user:$(cat /run/secrets/db_password)@db:5432"
secrets:
db_password:
external: true # managed by Swarm, not in this file
api_key:
external: true
Docker expertise isn't about knowing more commands. It's about having a layered mental model that connects kernel primitives to production behavior. When a container consumes unexpected CPU, you know to check the cgroup controller. When image pulls are slow, you know to audit layer order and base image choice. When a port isn't reachable, you know to inspect iptables DOCKER chain. When a container escape is reported, you check for root execution, excessive capabilities, and kernel version.
The security hardening checklist, the networking patterns, the storage choices, and the image optimization techniques are all independent concerns — but they interact. A rootless Docker daemon changes how user namespaces work, which affects how bind mounts resolve UIDs. Multi-stage builds affect which CVEs appear in Trivy scans. Overlay networks require Swarm mode, which requires thinking about stateful vs. stateless service design. An expert holds all these layers in mind simultaneously and makes tradeoffs deliberately rather than accidentally.
# Day 1: Understand what you're running
$ docker system info # daemon config, storage driver, cgroup version
$ docker ps --format "{{.Names}}: {{.Image}} ({{.Status}})"
$ docker system df # how much disk space is Docker using?
$ docker images --filter dangling=true # find orphaned images
# Day 2: Security audit
$ docker ps --format "{{.Names}}: {{.ID}}" | while read line; do
name=$(echo $line | cut -d: -f1)
id=$(echo $line | cut -d: -f2 | tr -d ' ')
user=$(docker inspect $id --format='{{.Config.User}}')
echo "$name: user='${user:-root}'" # empty = root!
done
$ trivy image myapp:prod --severity CRITICAL,HIGH
# Day 3: Network audit
$ docker network ls # what networks exist?
$ iptables -L DOCKER -n # what ports are exposed?
$ docker ps --format "{{.Ports}}" # port mappings
# Day 4: Image optimization
$ dive myapp:prod # inspect layer efficiency
$ docker images myapp --format "{{.Size}}" # current size
# Day 5: Apply multi-stage build, rootless, and seccomp
# → See code examples in sections above
Four experiments: Dockerfile cache simulator, security audit, image size optimizer, and namespace visualizer.
Dockerfile Layer Cache Simulator Each row is a Dockerfile instruction. Drag to reorder. Hit "Simulate Change" to see cache invalidation cascade. 0 Cache hits 0 Cache misses — Build time — Optimized? Build Output
● CACHED — layer reused from previous build
● MISS — layer invalidated, rebuilding...
● SOURCE — changed instruction triggers all below
Configure a container and see your security score. Each setting affects attack surface.
0/100 Security score F Grade 0 Critical issues 0 Controls passed Generated docker run Command FindingsImage size comparison across base images and optimization techniques
Image Size Calculator Application type Node.js API Optimizations applied — Before (MB) — After (MB) — Size reduction — Pull time (100Mb/s)Linux namespace isolation — click containers to inspect their namespace view
Namespace ExplorerEach container lives in its own set of namespaces. Click a container in the diagram to see what it can and cannot see.
Add container Selected container namespaces 0 Containers — Isolation risk