Docker · Lesson 8 of 11
Debugging and Inspecting Containers
Debug Docker containers that crash or misbehave using docker logs, exec, inspect, stats, events, exit codes, health checks and docker debug.
- Intermediate
- 16 min read
- 4 objectives
Before this lessonLesson 7: Environment Variables, Config and Secrets
What you will learn
- Read logs and exit codes to find why a container stopped
- Open a shell or run tools inside a running container
- Use docker inspect, stats and events
- Debug images that have no shell
Your Progress
0 of 11 lessons 0%
- Lessons0 / 11
- Completed0
- Est. time left~ 3 hours
Create a free account to keep your progress on every device.
Tip: pressing Next marks this lesson complete automatically.
Sooner or later a container will exit the moment it starts, restart in a loop, or answer every request with a timeout. Because the app runs in an isolated box, the usual approach of "look at the terminal" does not work. Docker gives you a small set of tools that cover almost every case, and a repeatable order to use them in.
The checklist this lesson builds: status and exit code, then logs, then inspect, then get inside. Most problems are solved by step two.
Step 1: status and exit code
docker ps -a --filter name=apiCONTAINER ID IMAGE COMMAND CREATED STATUS PORTS NAMES 7d3f9a2c1b8e stackcone-api:1.0.0 "uvicorn main:app --…" 40 seconds ago Exited (1) 38 seconds ago api
The exit code is the first clue:
0: the process finished normally. For a server, that usually means the command was wrong (it ran and ended instead of staying in the foreground).1(or other small numbers): the application raised an error. Check the logs.125,126,127: Docker could not run the command at all (bad flag, not executable, command not found).137: killed with SIGKILL, very often by the out-of-memory killer, or bydocker stopafter the grace period.143: stopped with SIGTERM, a normal graceful shutdown.
Step 2: logs
Docker captures anything the main process writes to stdout and stderr. Logs survive after the container exits (until you remove it):
docker logs --tail 20 apiINFO: Started server process [1]
INFO: Waiting for application startup.
ERROR: Traceback (most recent call last):
File "/app/main.py", line 14, in lifespan
await db.connect()
...
OSError: Multiple exceptions: [Errno 111] Connect call failed ('127.0.0.1', 5432)
ERROR: Application startup failed. Exiting.Found it: the app tried to reach Postgres at 127.0.0.1, which inside a container is the container itself (see the volumes and networks lesson). Useful flags: -f to follow, --since 10m, and -t for timestamps. With Compose, docker compose logs -f api does the same per service.
Step 3: inspect the configuration
docker inspect prints everything Docker knows about a container as JSON: environment, mounts, networks, restart count, health status. Use --format with a Go template to pull out one piece:
docker inspect --format '{{json .Config.Env}}' api
docker inspect --format '{{.State.OOMKilled}} restarts={{.RestartCount}}' api
docker inspect --format '{{range .Mounts}}{{.Source}} -> {{.Destination}}{{println}}{{end}}' api["DATABASE_URL=postgresql://app:app@localhost:5432/stackcone","PATH=/venv/bin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin","PYTHONUNBUFFERED=1"] false restarts=0 /Users/ada/stackcone/api -> /app
This confirms the bad DATABASE_URL came from configuration, not code. Change localhost to db and put both containers on the same network.
Step 4: get inside
For a running container, docker exec starts an extra process next to the app. Open a shell, check files, test DNS and connectivity from the container's point of view:
docker exec -it api sh
# inside the container:
cat /etc/hosts
env | grep DATABASE
python -c "import socket; print(socket.gethostbyname('db'))"127.0.0.1 localhost 172.19.0.4 8b1c0e7f2a3d DATABASE_URL=postgresql://app:app@db:5432/stackcone 172.19.0.2
If the container exits immediately, you cannot exec into it. Instead, start a new container from the same image with a shell as the command so it stays alive:
docker run --rm -it --entrypoint sh stackcone-api:1.0.0
# now explore: ls -la /app, run the start command by hand, read the tracebackResource usage and events
For slow or restarting containers, check live CPU and memory:
docker stats --no-streamCONTAINER ID NAME CPU % MEM USAGE / LIMIT MEM % NET I/O BLOCK I/O PIDS 8b1c0e7f2a3d api 187.42% 498.2MiB / 512MiB 97.30% 1.2MB / 3.4MB 0B / 0B 23 3a9e2d1c7b40 db 0.35% 42.1MiB / 7.66GiB 0.54% 2.8MB / 1.1MB 12.3MB / 45MB 9
The api container is at 97% of its 512 MiB limit and about to be OOM-killed (exit 137). Either raise the limit or find the leak. docker events streams what the daemon is doing, which makes restart loops obvious:
docker events --filter container=api --since 5m2026-09-20T11:40:02.114Z container oom 8b1c0e7f2a3d (image=stackcone-api:1.0.0, name=api) 2026-09-20T11:40:02.231Z container die 8b1c0e7f2a3d (exitCode=137, image=stackcone-api:1.0.0, name=api) 2026-09-20T11:40:02.735Z container start 8b1c0e7f2a3d (image=stackcone-api:1.0.0, name=api)
Health checks
A process can be running yet broken (deadlocked, lost its DB connection). A HEALTHCHECK lets Docker test it and report healthy or unhealthy:
HEALTHCHECK --interval=10s --timeout=3s --retries=3 \
CMD python -c "import urllib.request; urllib.request.urlopen('http://localhost:8000/health')" || exit 1docker inspect --format '{{json .State.Health}}' api{"Status":"unhealthy","FailingStreak":4,"Log":[{"Start":"2026-09-20T11:45:10Z","End":"2026-09-20T11:45:13Z","ExitCode":1,"Output":"urllib.error.URLError: <urlopen error timed out>\n"}]}Debugging images with no shell
Distroless and scratch images (from the multi-stage lesson) have no sh, so docker exec -it api sh fails. Two options work well. docker debug (included with Docker Desktop) attaches a toolbox shell with common utilities to any container or image without changing it. On any Docker host, you can also run a tools container that shares the target's network and process namespaces:
docker run --rm -it --network container:api --pid container:api nicolaka/netshoot
# inside: ps aux, curl localhost:8000/health, ss -tlnp, dig dbFinally, docker cp api:/app/config.yaml . copies files out of any container, running or stopped, which is handy for grabbing a generated config or crash dump.
Recap
- Start with
docker ps -a: the exit code (0, 1, 127, 137, 143) tells you what kind of failure it was. docker logssolves most problems; make sure your app logs unbuffered to stdout.docker inspect --formatshows the real environment, mounts, health and OOM status.- Use
docker execfor running containers anddocker run --entrypoint shfor ones that crash on start. - For shell-less images, use
docker debugor a tools container sharing the target's namespaces.
# Write your solution here
Finished reading? Mark this lesson complete to track your progress.
