Programming and IT

What is Docker

Docker packages a program together with everything it needs to run into a single image — and that image starts the same way on a developer's laptop and on a server. Below — what an image is, how it differs from a container, how to write your first Dockerfile and where the technology is overkill.

Updated
In this article

The pain it removes#

A program rarely consists of a single file. It needs a particular language version, a dozen libraries, environment variables and sometimes system packages. On the author's machine all of this is already installed, on a colleague's it is a different version, and on the server a third one. Hence the eternal "it works on my machine".

Docker solves this by moving the boundary: what gets packaged is not the program but a whole snapshot of a file system with the program already installed and configured. That snapshot starts the same way anywhere Docker is available.

A virtual machine solves the same problem differently — it carries a complete operating system with its own kernel, weighs tens of gigabytes and takes minutes to start. A container uses the host system's kernel and is isolated by the kernel's own mechanisms, so it weighs tens of megabytes and starts in a fraction of a second. On macOS and Windows, Docker does keep one small Linux virtual machine — the containers run inside it.

Image and container#

Mixing up these two words gets in the way of understanding everything else, so let us settle it right away.

An image is an immutable template: a set of file-system layers plus a description of what to run. You build it once, store it in a registry and download it. You cannot change an image, only build a new one.

A container is a running instance of an image with a writable layer added on top. You can start ten containers from one image, and they do not see each other's files.

The closest analogy from programming: an image is a class, a container is an object. An installation disc works the same way: you can install the system on as many machines as you like, and the disc itself does not change.

The practical consequence: everything a container writes to its own layer disappears when it is removed. Almost everyone's first bad experience is the same — a database lived in a container for a week and vanished with it. Volumes fix this; more on them below.

First commands#

The first line downloads an image and runs the program inside, the second starts Ubuntu and drops you into its shell, and the third runs a web server in the background:

docker run hello-world
docker run -it ubuntu bash
docker run -d -p 8080:80 --name web nginx

The options on the last line: -d — run in the background, -p 8080:80 — port 8080 on your machine leads to port 80 inside the container, --name web — the name you will use to refer to it, nginx — the image.

Remember the order in -p like this: outside on the left, inside on the right. Swap them, and the browser will find nothing at localhost:8080.

docker ps                 # what is running now
docker ps -a              # every container, including stopped ones
docker logs -f web        # the program's output, -f follows new lines
docker exec -it web sh    # a shell inside a running container
docker stop web           # stop it
docker rm web             # remove a stopped container
docker image ls           # downloaded and built images

docker logs is the first thing to check when a container started and immediately died. The second clue is nearby: docker ps -a shows the exit code in the status column.

Images pile up unnoticed and take gigabytes. docker system df shows how much space they use, and docker system prune frees it; the command removes stopped containers and unused images, so read its question before you confirm.

Dockerfile#

An image is described in a text file called Dockerfile. Each instruction is a separate layer.

FROM python:3.12-slim

WORKDIR /app

COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt

COPY . .

ENV PORT=8000
EXPOSE 8000

CMD ["python", "app.py"]

Line by line:

  • FROM — the base, a ready-made image from a registry. Variants tagged slim or alpine are several times smaller than the full ones.
  • WORKDIR — the working directory inside the image; all following paths are relative to it.
  • COPY — copy files from your folder into the image.
  • RUN — run a command during the build.
  • ENV — an environment variable.
  • EXPOSE — a note that a port is used; on its own it opens nothing, the port is published by -p at run time.
  • CMD — what to run when the container starts.

Why the dependency list is copied separately and before the rest of the code: Docker caches layers and rebuilds only those after the one that changed. With this order, editing one line of code does not force all the libraries to be downloaded again. Change the order, and every build takes minutes instead of seconds.

Next to it you put .dockerignore — a list of what must not go into the image: the .git directory, local virtual environments, files with keys. It works like .gitignore and speeds up the build as well.

Build and run:

docker build -t myapp:1.0 .
docker run -d -p 8000:8000 --name api myapp:1.0

The dot at the end of build is the build context, the directory the COPY files come from. The -t tag gives the image a name and version; without a version you get latest, and a month later you will not know what exactly is running on the server.

Data: volumes#

For data to survive the removal of a container, you keep it outside.

A named volume is created and stored by Docker itself:

docker volume create pgdata
docker run -d --name db \
  -e POSTGRES_PASSWORD=secret \
  -v pgdata:/var/lib/postgresql/data \
  postgres:16

The second way is to mount a directory from your machine inside:

docker run -d -p 8000:8000 -v "$PWD":/app myapp:1.0

The first kind, a named volume, suits databases, is fast and lives in Docker's own storage. The second, a bind mount, makes any edit to a file on your side visible inside immediately, which is why it is used during development.

A practice database started with a command like this is the simplest way to learn SQL without installing a server on your system: when you are done, remove the container and the volume.

Once you have more than one service, the commands stop fitting in your head, and you describe them in a compose.yaml file:

docker compose up -d      # start everything described
docker compose ps         # state of the services
docker compose logs -f    # combined log
docker compose down       # stop and clean up

Services from one file see each other by name: the application connects to the database at the address db, not by IP.

When you do not need Docker#

The technology takes time to learn and adds a layer that can break too. There are tasks where it is overkill.

  • A single script on your own machine. The language's virtual environment solves dependencies and needs nothing extra installed.
  • A static website. Ready pages are deployed as files; there is nothing to package.
  • A program with a graphical window. You can forward the interface, but it is more hassle than it is worth.
  • Heavy disk work on macOS and Windows. Mounted directories there go through a virtual machine and are noticeably slower.

And the flip side: a container does not make a program secure automatically. Running as root inside, keys baked into the image, --privileged "just in case" — all of that reduces isolation to nothing.

Common mistakes#

The container starts and stops at once. A container lives exactly as long as its main process runs. The program exited — the container closed. Check docker logs.

The port is taken. bind: address already in use means the port on the left of -p is already used by something; ss -tulpn finds the culprit — the command is explained in the Linux commands cheat sheet.

Code changes are not visible. The image was built with an old copy of the files. Either rebuild it or mount the directory with -v.

Keys and passwords inside the image. Anything that made it into a layer can be pulled out of the image by anyone, even if the next instruction deleted it. Secrets are passed as environment variables at run time — the same principle as with API keys.

Step-by-step plan

  1. Run someone else's imageStart nginx with docker run, publish a port and open the page in your browser.
  2. Look around insideEnter a running container with docker exec, look at the files, exit and check the container is still alive.
  3. Build your own imageWrite a five-instruction Dockerfile for a small program and build it with a version tag.
  4. Check the layer cacheChange a line of code, rebuild and compare the time with the first build; then reorder the COPY lines and repeat.
  5. Keep your dataStart a database with a named volume, write a row, remove the container, start it again and find the row still there.

Start learning this in your own space

The plan goes into your repository: tick off stages, keep notes — the change history shows how far you have come.

Start the plan

Check yourself

1.Which command lists running containers?

2.A container was started with -p 8080:80. Which port do you use to reach it from your machine?

3.What does docker build produce from a Dockerfile?

4.What do you need for a database's data to survive the removal of its container?

Sources

Was this helpful?

More articles

Programming and IT What is an API An API is an agreed way for one program to ask another for something. No screens and no buttons — the request travels as text over the network, and the answer comes back in a machine-readable form. Once you understand the four parts of an HTTP request, you can read the documentation of any service. Programming and IT C++ from scratch C++ is a compiled language used wherever speed and direct access to memory matter — game engines, browsers, databases, firmware. It is harder to get into than Python, but you can build your first working program on the very first evening. Programming and IT How to learn Python from scratch Python is a good first programming language: code reads almost like text, and the standard library covers most everyday tasks. This plan takes you from installing the interpreter to your own scripts covered by tests in about four months, at roughly an hour a day. Programming and IT How to learn SQL from scratch SQL is the query language of relational databases. Developers, analysts, testers and managers who want to pull numbers themselves all need it. Basic queries take a few weeks to learn; working confidently with complex reports takes two or three months of practice. Below is the order of topics and ways to train on a real database. Programming and IT How to learn Linux from scratch Linux runs most servers, containers and countless devices, so developers, testers, analysts and system administrators all need the command line. The easiest way to learn it is not by reading lists of commands but by working in the terminal every day and solving small practical tasks. Below is a sequence of topics for two to three months. Programming and IT How to learn Java from scratch Java is a strictly typed language behind banking systems, the servers of large services and Android apps. The strictness slows you down at first, but the compiler catches many mistakes before the program ever runs. This plan takes about six months at an hour a day and leads from your first program to a small backend application.

More solutions