Home › Guides › Docker best practices
Tech Explained · 2026Docker Best Practices in 2026: 10 Mistakes That Bloat Your Images and How to Fix Each One
Docker best practices in 2026 come down to one habit: ship only what actually runs. Pick a slim base image, order your layers so dependencies stay cached, move build tools into a separate stage, and add a .dockerignore. Done properly that takes a typical Python service from roughly 1 GB to under 100 MB.
-
The base image is most of the fight.
python:3.11topython:3.11-slimdrops around 870 MB before you touch any code. -
Layer order decides build time.
COPY . .above your dependency install means every commit reinstalls every package. - Docker Hub meters you now. Anonymous pulls cap at 10 per hour per IP, so an unauthenticated CI runner fails before your code does.
- Alpine is not the default for Python. Debian slim is usually the better trade, and this guide says when it is not.
-
Secrets baked into a layer stay there after a later delete.
docker historyproves it. - Multi-stage builds cost nothing at runtime, the best return available once your base image is sorted.
Your build worked fine in March. Now it takes eleven minutes, the image is 1.02 GB, deploys time out on staging, and last Tuesday the CI runner started returning toomanyrequests: You have reached your pull rate limit at 4pm for no reason anyone could explain. Nothing broke. The Dockerfile just accumulated ten small decisions, every one reasonable at the time.
Docker Best Practices Start With the Base Image
Keep one team in mind throughout: two platform engineers at a logistics SaaS in Pune, running nine Python services on a small Kubernetes cluster. Nobody owns the Dockerfiles, so each was copied from the last one. Their three costs are the ones bloat always charges. Registry throttling, because Docker's usage documentation caps unauthenticated pulls at 10 per hour per IP address and a free Docker Personal account at 100 per hour, with unlimited pulls under fair use only on Pro, Team and Business. CI minutes, because a research survey of Dockerfile smells found image bloat added 48.06 MB on average and produced images up to 89% larger than necessary. And attack surface, because every package shipped is a package somebody patches. If you are working toward AWS DevOps certification, that last one is the reasoning DOP-C02 wants: not "make images small" but "justify what is in the image".
Also read: Docker Tutorial for Beginners 2026: Containerise a Python API in 7 Steps, which builds the working image this guide takes apart.
Mistakes 1 to 3: the base image you picked without thinking
What the tag suffix is worth in megabytes
Same language, same interpreter, very different images. The cheapest change on the list.
Sizes as reported by the pythonspeed base image comparison (February 2026) and Docker Hub tag listings, checked 27 September 2026. Tags differ by interpreter version and Debian release, so read these as orders of magnitude and measure with docker image ls.
Mistake 1: shipping the full language image
The plain python:3.11 tag carries a complete Debian userland plus the build toolchain, which is why it lands near a gigabyte. Almost no production service needs gcc at runtime. Change one word in your FROM line and measure, before anything else.
Mistake 2: pinning to :latest
Two engineers build the same commit a week apart and get different images. That is not a mystery, it is :latest. Pin a real tag, and pin the digest if the build must be reproducible for an audit. For scale: the current Docker Engine release is 29.8.1 as of 16 September 2026, with 29.8 shipping on 3 September, and Docker Desktop 4.91.0 bundles Engine 29.8.0. Engine moves that fast. Base images move faster.
Mistake 3: reaching for Alpine reflexively
Here is the opinion, trade-off stated out loud: for Python, default to Debian slim rather than Alpine. Alpine uses musl libc instead of glibc, so any dependency with a C extension needs a musllinux wheel. When one is missing, pip compiles from source, your build goes from 40 seconds to 6 minutes, and you install gcc into the final image to make it work. The small image is gone and the build is slow too. Comparisons published in 2026 note the gap often narrows from around 70 MB to 20 or 30 MB once real dependencies are installed. Choose Alpine for a pure-Python or pure-Go dependency tree you control, and only after measuring.
| Base image family | Rough size | Pick it when | The gotcha nobody mentions |
|---|---|---|---|
Full language image (python:3.x, node:20) |
~1 GB | Local experiments and build stages only | Ships a compiler into production |
Slim (-slim, -slim-trixie) |
41 MB to 350 MB by tag | Almost every Python or Node service | Still needs build-essential in a build stage for some wheels |
Alpine (-alpine) |
17 MB to 100 MB | Pure-Python, Go, or static binaries | musl libc breaks C extension wheels |
| Distroless | Tens of MB | Hardened deployments with a security review | No shell, so docker exec debugging is gone |
Distroless is the one I would not start with: the first time a container crash-loops and you cannot get a shell in, it costs you an afternoon.
How Docker Layer Caching Works, and the Two Mistakes That Break It
Every instruction creates a layer, and Docker reuses a cached layer only if that instruction and everything above it are unchanged. That one rule explains most slow builds you will meet.
Where the cache breaks when you edit one line of code
Same four instructions, different order. The right-hand column reinstalls dependencies only when requirements.txt changes.
Mechanism per Docker's documented build cache behaviour, checked 27 September 2026.
Mistake 4: copying source code before installing dependencies
Fix this first: it costs nothing and nearly everyone has it. Copy the dependency manifest, install, then copy the rest.
Mistake 5: no BuildKit cache mount
Even with correct ordering, changing one line in requirements.txt redownloads every wheel. A cache mount keeps the download cache on the build host without putting it in the image. Add the syntax directive on line one, then mount your package manager's directory: /root/.cache/pip, /root/.npm, /go/pkg/mod, or /root/.m2/repository.
# syntax=docker/dockerfile:1
FROM python:3.13-slim
WORKDIR /app
# dependency manifest first, so this layer survives code edits
COPY requirements.txt .
RUN --mount=type=cache,target=/root/.cache/pip \
pip install --no-cache-dir -r requirements.txt
# source last
COPY . .
CMD ["gunicorn", "--bind", "0.0.0.0:8000", "app:app"]
You add the cache mount, it works locally, and it does nothing in CI. Expected: the cache lives on the build host, and most CI runners are fresh containers with no host to inherit from. You need BuildKit's cache export to a registry, or a persistent runner. And if you forget the # syntax=docker/dockerfile:1 line, older builders reject --mount with a parse error.
Docker Multi-Stage Build: The Fix for Shipping Your Compiler (Mistake 6)
A multi-stage build puts the compiler, headers and build cache in one stage, then copies only the finished artefact into a clean final stage. Everything else is discarded. For Python, build into a virtualenv and copy that. For Go or Rust, copy the binary and the final image can be almost empty.
# syntax=docker/dockerfile:1
# ---- stage 1: build ----
FROM python:3.13-slim AS builder
WORKDIR /app
RUN apt-get update && apt-get install -y --no-install-recommends \
build-essential \
&& rm -rf /var/lib/apt/lists/*
RUN python -m venv /opt/venv
ENV PATH="/opt/venv/bin:$PATH"
COPY requirements.txt .
RUN --mount=type=cache,target=/root/.cache/pip \
pip install --no-cache-dir -r requirements.txt
# ---- stage 2: runtime ----
FROM python:3.13-slim AS runtime
RUN useradd --create-home --uid 10001 appuser
COPY --from=builder /opt/venv /opt/venv
ENV PATH="/opt/venv/bin:$PATH" \
PYTHONDONTWRITEBYTECODE=1 \
PYTHONUNBUFFERED=1
WORKDIR /app
COPY --chown=appuser:appuser . .
USER appuser
EXPOSE 8000
CMD ["gunicorn", "--bind", "0.0.0.0:8000", "app:app"]
Note what stage two does not contain: build-essential, the pip cache, the apt lists, and root as the running user.
Real teams publish the payoff. Docker's own customer story for Siimpl reports a 90% reduction in local build times after moving to Docker Build Cloud remote builders alongside self-hosted GitHub runners, and the Artefact engineering team's write-up of their image and build-time work accounts for 41 hours saved per month. Neither number came from making the application faster. This pipeline shape is also where cloud certification study gets practical: the Azure DevOps and architecture program covering AZ-400 works through it live, and it reappears when you package model-serving containers in an MLOps engineering course.
Mistakes 7 and 8: Secrets Baked Into Layers, and Running as Root
Mistake 7: a secret in a layer you thought you deleted
Copying a token in, using it, then deleting it in a later RUN does not remove it. The layer that held it is still in the image, and anyone who pulls can read it. Prove it to yourself:
# every layer, and the command that created it
docker image history --no-trunc myapp:latest
# pull the layers apart if you want certainty
docker save myapp:latest -o myapp.tar && tar -tf myapp.tar
The fix is BuildKit's secret mount, which exposes the file during one RUN and never writes it to a layer:
# in the Dockerfile
RUN --mount=type=secret,id=pipconf,target=/root/.config/pip/pip.conf \
pip install --no-cache-dir -r requirements.txt
# at build time
docker build --secret id=pipconf,src=$HOME/.config/pip/pip.conf -t myapp:1.4.0 .
Mistake 8: no USER instruction
By default your process runs as root, so one escape or writable mount becomes a much worse afternoon. Two lines fix it, as above, and Kubernetes enforces it with runAsNonRoot: true in the pod security context. Scan before pushing: docker scout quickview myapp:1.4.0 gives a vulnerability summary plus base image update advice, and docker scout cves myapp:1.4.0 --only-severity critical,high --only-fixed narrows it to what you can act on today.
Also read: What Is CI/CD? How the Pipeline Works, CI vs CD, and Your First Workflow in 2026, for where to wire these scans into a GitHub Actions job.
Reduce Docker Image Size: The Two Cleanup Mistakes (9 and 10)
Mistake 9: no .dockerignore
Without one, COPY . . sends your .git directory, local virtualenv, test fixtures and node_modules into the build context and then into the image. On a repository with real history, .git alone is often the largest single thing you ship. A working starting point:
.git
.gitignore
.venv
venv/
__pycache__/
*.pyc
.pytest_cache/
.mypy_cache/
node_modules/
tests/
*.md
.env
.env.*
Dockerfile
docker-compose.yml
Mistake 10: Package Manager Debris, and Three Dockerfile Best Practices in One Line
Three Dockerfile best practices, all in one line. Combine apt-get update and apt-get install in a single RUN so a cached update layer cannot serve stale indexes. Add --no-install-recommends so apt stops pulling in docs and suggested extras. Delete /var/lib/apt/lists/* in that same instruction, because deleting it in the next one leaves it in the previous layer. Same logic for pip install --no-cache-dir and npm ci --omit=dev with npm cache clean --force.
Worked Example: One Flask Image From 1.02 GB to 96 MB
Back to the Pune team. They took their smallest service, a Flask API with eleven dependencies and no C extensions, and applied the changes one at a time, running docker image ls after each. The order matters: it tells you which change actually paid.
| Step | What changed | Approx image size | What it bought |
|---|---|---|---|
| 0 |
FROM python:3.13, COPY . . then pip install |
~1.02 GB | Baseline |
| 1 | FROM python:3.13-slim |
~171 MB | One word, roughly 85% gone |
| 2 | Added .dockerignore
|
~158 MB | Dropped .git, .venv and test fixtures |
| 3 | Reordered layers, added cache mount | ~158 MB | No size change; rebuilds went from minutes to seconds |
| 4 | Multi-stage, copy the venv only | ~112 MB | build-essential and pip cache left behind |
| 5 |
--no-install-recommends, apt lists removed, non-root USER
|
~96 MB | Last of the debris, plus a real security win |
Step 3 changed the size by nothing, and it was the change the team noticed most, because they feel it on every push. Step 1 did most of the work, which is why the advice runs: measure, change the base image, measure again, then get clever. Figures are the ballpark for an app of this shape, not a benchmark.
The caveat this guide owes you. If your image is 3.8 GB because it holds PyTorch and CUDA libraries, none of the above saves you. Those wheels are the image, and the real question is whether GPU libraries belong in the serving container at all. Likewise, on warm Kubernetes nodes that already cache your base layers, shaving 40 MB will not show up in deploy times. Fix build time first.
Go from clean Dockerfiles to a full AWS deployment pipeline, live with a mentor
A 12-week live weekend program that prepares you for both AWS Solutions Architect Associate (SAA-C03) and AWS DevOps Engineer Professional (DOP-C02). Includes hands on projects, mentor support and placement guidance.
Explore the course
Docker Best Practices Checklist: The Four Fixes Worth Doing Today
One hour, nine Dockerfiles, this order. The rest is refinement.
Ranked by payoff per minute of work
Start at card one. Do not start at card four because it is more interesting.
Change the base image tag
Swap the full language image for its slim variant, rebuild, run the tests. Most of your size problem, at two minutes per service.
2 min, biggest winMove COPY . . below the install
Copy the manifest, install, then copy source. Rebuilds stop reinstalling every package. No size change, large time change.
5 min, daily payoffAdd a .dockerignore
Start with .git, .venv, __pycache__, node_modules and tests. Compare the build context size in the first line of build output.
5 minAuthenticate your CI pulls
A docker login step moves your runner from 10 anonymous pulls an hour to 100. Cheapest fix for CI failures you cannot reproduce.
Rate limit figures from Docker Hub usage documentation, checked 27 September 2026.
Docker image optimization: the diagnostic commands to keep to hand
| Command | What it tells you | Run it when |
|---|---|---|
docker image ls |
Size of every local image | Before and after every change |
docker image history --no-trunc myapp:tag |
Which instruction created which layer, and how big it is | You cannot find where the megabytes went |
docker system df |
Disk used by images, containers, volumes and build cache | Your laptop runs out of space |
docker builder prune |
Reclaims build cache | The above says the cache is the problem |
docker scout quickview myapp:tag |
Vulnerability summary and base image update advice | Before every push to a shared registry |
docker buildx build --platform linux/amd64,linux/arm64 . |
Builds for both architectures | You are on Apple Silicon deploying to x86 servers |
That last row wastes whole evenings. You build on an M-series Mac, push, and the container exits instantly with exec format error. You built arm64 and your server is amd64. Build with --platform linux/amd64, or produce a multi-arch image with buildx.
The 10 mistakes at a glance
| # | Symptom you notice | Cause | Fix |
|---|---|---|---|
| 1 | Image near 1 GB for a small app | Full language base image | Use the -slim tag |
| 2 | Two builds of one commit differ |
:latest in FROM
|
Pin the tag, pin the digest if audited |
| 3 | Build slowed to minutes after moving to Alpine | musl libc, no prebuilt wheels | Go back to Debian slim unless measured otherwise |
| 4 | Every push reinstalls all dependencies |
COPY . . above the install step |
Copy the manifest first |
| 5 | Changing one dependency redownloads all of them | No BuildKit cache mount | RUN --mount=type=cache,target=/root/.cache/pip |
| 6 | gcc and headers present in production | Single-stage build | Multi-stage, copy only the artefact |
| 7 | Token visible in docker image history
|
Secret written into a layer |
--mount=type=secret and --secret at build time |
| 8 | Container process runs as root | No USER instruction |
useradd then USER appuser
|
| 9 | Build context far larger than the code | Missing .dockerignore
|
Ignore .git, virtualenvs, caches, tests |
| 10 | Unexplained 40 to 80 MB in one layer | apt lists and pip cache left behind | Clean up inside the same RUN
|
Layer caching and multi-stage builds turn up in the container sections of the full certifications overview across both AWS and Azure DevOps tracks, and containerised deployment is how generative AI work reaches production, which is why a live AI engineering program spends time packaging RAG services rather than only writing them. To watch a 1 GB image become a 96 MB one live before committing to anything, sit in on a free webinar.
Also read: What Is Kubernetes? Plain-English Definition, How It Works and Where You Will Use It in 2026, which picks up where a well-built image ends.
What I Would Do in Your Position
Pick your smallest service, not your most important one. Change the base image to the slim tag, run the tests, note the number. Move COPY . . below the install step and note how long the next rebuild takes. Fifteen minutes on one service will teach you more about your pipeline than any amount of reading.
Then stop. Do not convert nine Dockerfiles to distroless multi-arch builds this week. Ship the two cheap changes everywhere, watch a fortnight of CI runs, and let the remaining pain tell you which of the other eight mistakes you actually have. Bloat is a habit of not measuring, and the habit is what you are replacing.
To build that habit under supervision rather than by trial and error, the AWS Solutions Architect and DevOps Engineer course runs this material live on Saturdays and Sundays, 8 to 11 PM IST across 12 weeks, and you can sit in on a demo class first. Data teams meeting the same problem inside pipelines will find the equivalent work in a live Microsoft Fabric data engineering program.
Related guides
- SAA-C03 vs DOP-C02: Which AWS Certification Should You Take First in 2026? which exam to sit once container work has convinced you to certify.
- DevOps Engineer Salary in India 2026 what this skill set is advertised at by experience band.
- System Administrator to DevOps Engineer in India 2026 the full switch plan if Dockerfiles are your way into DevOps.
- Linux Commands Cheat Sheet for DevOps in 2026 the shell fluency you need when a container misbehaves.
- What Is Infrastructure as Code? the next layer down, once your images are clean.
- DevOps Jobs in Dubai 2026 where this stack pays in the Gulf, and who asks for it by name.
Frequently asked questions
What are the most important Docker best practices in 2026?
In order of payoff: a slim base image instead of the full language image, layers ordered so the dependency install sits above the source copy, a multi-stage build so compilers never reach production, a .dockerignore, a non-root user, and a pinned base tag. The first two account for most of the benefit.
How do I reduce Docker image size without breaking my app?
Change one thing at a time and run your tests after each. Start with the base image tag, add a .dockerignore, then move to a multi-stage build that copies only your virtualenv or binary. Use docker image history --no-trunc to find which layer holds the megabytes.
Should I use Alpine or slim for Python Docker images?
Debian slim, unless you have measured otherwise. Alpine uses musl libc, so packages with C extensions need musllinux wheels; when one is missing pip compiles from source and often forces a compiler into the image anyway. Comparisons published in 2026 note the gap narrows to 20 or 30 MB once real dependencies land.
Does a multi-stage build make my builds slower?
No, and often faster, because the build stage caches independently of your application code while the final stage does almost nothing beyond copying an artefact. What changes is the mental model: anything needed at runtime must be explicitly copied into the final stage.
Why is my Docker build not using the cache?
Almost always because an instruction above it changed. A COPY . . near the top invalidates everything below whenever any file changes, which is why a .dockerignore affects caching as well as size. Other causes: a floating :latest tag, a changed ARG, or a fresh CI container with no cache.
How many Docker Hub pulls do I get for free?
Docker's usage documentation sets unauthenticated pulls at 10 per hour per IP address and a free Docker Personal account at 100 per hour, while Pro, Team and Business accounts get unlimited pulls under fair use. Because the anonymous limit is per IP, an office network shares one allowance.
Do I need to run containers as a non-root user?
Yes, for anything facing a network. Two lines remove a whole class of escalation from any successful exploit. Create the user in the final stage, use COPY --chown for application files, put USER before CMD, and enforce runAsNonRoot: true in your Kubernetes pod security context.
Is a smaller Docker image actually faster in production?
Sometimes, and less often than people assume. On nodes that already cache your base layers, shaving 40 MB changes almost nothing. Smaller images pay off on cold starts, autoscaling onto fresh nodes, serverless platforms, and in registry bandwidth and CI time. The security benefit applies regardless.
About this guide. 360 Digital Transformation is an Authorized Training Partner of Anthropic and Microsoft. Other certification bodies, vendors and employers named here are not affiliated with us. Tools and versions change quickly; commands and figures cited were checked on 27 September 2026.




