ImageLs icon
OutputValues icon
Log icon
If icon
Loop icon
Rm icon
Prune icon
Fail icon
SlackIncomingWebhook icon
Schedule icon

Docker Host Image Retention and Disk Reclaim

Keep the newest N tags per repository and protected tags on a Docker host, remove the rest without force, prune dangling images and verify.

Categories
CoreInfrastructure

Self-hosted CI runners and build hosts fill their disks with images: one per commit, per branch, per service, kept forever. docker image prune -a frees the space but also deletes the base images the next build needs, and the tag someone deployed by hand. So the disk alarm fires every few weeks, and someone deletes images by hand on a production runner.

This blueprint applies a retention policy instead. For every repository matching a filter, it keeps the newest N tags and any image with a protected tag, and removes the other images one by one without --force. An image a container still uses is therefore never deleted. It prunes dangling images, then lists the images again to verify that nothing the policy keeps went missing.

This blueprint was created by zkasuran.

How it works

  1. list_images (io.kestra.plugin.docker.cli.ImageLs, read-only) lists the images matching image_filter.

  2. plan groups their tags by repository and ranks them by creation time:

    • KEEP_NEWEST: one of the newest keep_per_repository of its repository;
    • KEEP_PROTECTED: one of its tags matches protected_tags;
    • otherwise REMOVE.

    summary logs the plan and the space it frees at most.

  3. apply_gate (io.kestra.plugin.core.flow.If) stops there when dry_run is true. Otherwise:

    • remove (io.kestra.plugin.core.flow.Loop over io.kestra.plugin.docker.cli.Rm, force: false) removes each candidate with all its tags. A failure, usually an image in use, does not stop the others.
    • prune (io.kestra.plugin.docker.cli.Prune, IMAGES, dangling: true) removes untagged layers.
    • relist and result compare the images before and after. verify_result fails the run if any kept image is missing.
    • reclaimed logs what was removed. in_use_notice posts to Slack when a candidate was kept because a container uses it, often a forgotten container.
  4. errors alerts Slack when the flow fails, for example on an unreachable daemon.

What you get

  • A retention policy per repository, instead of an all-or-nothing prune.
  • Protected tags and in-use images that are never removed, by design rather than by luck.
  • A before and after comparison on every run, and a notice for forgotten containers holding old images.

Who it's for

  • Teams running self-hosted CI runners (GitHub Actions, GitLab, Buildkite, Jenkins agents) with a local Docker daemon.
  • Platform teams operating shared build hosts or Docker-based edge nodes.
  • Teams evaluating the Docker plugin who want an example of ImageLs, Rm and Prune driven by a plan.

Why orchestrate this with Kestra

A cron job running docker image prune cannot tell a release tag from a commit build, or report what it deleted. Kestra lists the images, computes the plan in one jq expression, gates the change behind dry_run, removes images one task at a time so one failure is isolated, verifies the result, and keeps an execution history of what was removed from which host.

Tested end to end

On Kestra 2.0.5 OSS against an isolated Docker 27 daemon (Docker-in-Docker), with a local HTTP sink standing in for Slack. The daemon held:

  • shop/api tags 1.0.1 to 1.0.7, with 1.0.2 also tagged prod, and 1.0.3 used by a running container;
  • shop/worker tags 2.1 to 2.3;
  • one dangling image, plus the alpine and busybox base images.
Run Inputs Result
1 defaults (dry run) 10 images: 6 newest kept, 1.0.2 kept as protected, 3 to remove (1.0.1, 1.0.3, 1.0.4). Nothing removed
2 dry_run: false 1.0.1 and 1.0.4 removed. 1.0.3 kept because a container uses it, and the Slack notice is sent. Dangling image pruned. Verified: every kept image is still there
3 dry_run: false again Nothing new to remove, 1.0.3 still reported in use
4 any run alpine and busybox untouched: they do not match the filter
5 unreachable daemon errors alert sent

Prerequisites

  • Access to the Docker daemon to clean, through its socket mounted into the Kestra worker or a TCP endpoint.
  • Pick an image_filter that matches only the images your builds produce.
  • Local testing: docker run -d --privileged --name dind -e DOCKER_TLS_CERTDIR= docker:27-dind, then build a few tagged images in it, and point docker_host at tcp://dind:2375.

Secrets

  • SLACK_WEBHOOK_URL: used by in_use_notice and the errors block. In Kestra OSS, provide it as the base64-encoded environment variable SECRET_SLACK_WEBHOOK_URL.

Inputs

Input Default Purpose
docker_host unix:///var/run/docker.sock Daemon to clean
image_filter registry.example.com/shop/* Images considered
keep_per_repository 3 Newest tags kept per repository
protected_tags ^(latest|prod|stable|release-.*)$ Tags always kept
prune_dangling true Also prune untagged images
dry_run true Must be false for anything to be removed

Quick start

  1. Add the SLACK_WEBHOOK_URL secret.
  2. Import the flow and set docker_host and image_filter for one build host.
  3. Run it with the defaults and read the plan in the logs.
  4. Run it with dry_run: false to apply the plan.
  5. Enable the nightly trigger, and change its dry_run input to false once the plans look right.

How to extend

  • Loop over a list of build hosts to clean every runner in one execution.
  • Trigger the flow from your disk usage alert instead of a schedule.
  • Add Prune with BUILD to clear the build cache, and CONTAINERS with until to remove old stopped containers.
  • Keep the tags currently deployed by reading them from your deployment tool and adding them to protected_tags.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.