Skip to content
  • Docker
  • BuildKit

Docker build done. Still exporting cache…

A Docker build could finish its build steps and then spend minutes exporting cache. I wanted to prepare the cache while those steps were running.

Gaurav Tiwari Updated 13 min read

The first cache I built for BoringCache was an archive cache. It was content-addressed, split large archives into reusable pieces, and avoided storing the same content twice within a workspace. The first UX was just a wrapper around the command you were already going to run. For Bundler, it looked like this:

Terminal
$ boringcache run -- bundle install

boringcache run recognised Bundler from the command, restored the archive, ran bundle install, then saved the result if the command succeeded. The next machine could run the same command. Because the archive was content-addressed and chunked, it only uploaded chunks that were not already present. It worked, but after a while I started to question how valuable it was on its own. Storage is cheap. Saving a little space on a Bundler or package-manager cache is useful, but it does not address the build time most people are waiting to reduce. A Docker build can take much longer. So Docker became the second product I looked at.

The registry cache

The first thing I tried was a registry cache. BuildKit already knows how to publish a cache as OCI content, and registries already know how to store and serve those blobs. I built that path and started comparing it with the GitHub Actions cache backend. In my tests it was faster than the GitHub Actions backend, but it was about the same as a good registry cache such as ECR. That told me the storage path worked, but it did not feel like a new product yet.

Then I started reading the build logs more carefully. The Dockerfile steps would finish, sometimes mostly from cache, and the job would keep running:

BuildKit progress
exporting to image
exporting cache
preparing build cache for export

The Docker build was done. BuildKit was still exporting cache.

Why the export happens at the end

BuildKit prepares the external cache at export time. During the solve, BuildKit creates and reuses local refs. An external cache needs a portable graph plus transferable layer blobs. At export time BuildKit walks the eligible cache records, resolves the refs into those blobs, uploads what the destination needs and finally publishes the cache manifest. The upstream BuildKit cache interface makes the division explicit: cache is imported for the solve and exported as one of its outputs. That is a safe default. Materialising and compressing blobs consumes CPU, disk and memory. On a small machine, doing too much of it beside the actual build can make the build slower.

But it leaves all of that work at the end. I wondered whether BuildKit could prepare some of the cache while the build had spare CPU or I/O, instead of waiting for every Dockerfile step to finish. I wanted to prepare enough cache during the solve to reduce the final export time without taking CPU or I/O capacity the build needed.

Cache bodies can be prepared during build steps when spare CPU or I/O is available, before the final manifest is published.
Cache bodies are prepared during the build, leaving less work at the end.

The Dockerfile experiment

Before changing BuildKit, I tried to use the archive system I already had. The expensive work in most Dockerfiles is concentrated in a few steps: installing packages, compiling a frontend, building a binary or running a tool such as Cargo, Turbo or Gradle. If I could put the BoringCache CLI inside the build and cache those smaller units, perhaps the whole Docker build would become faster. It did work. The problem appeared every time I released a new CLI. Changing the BoringCache binary changed the Docker instruction that introduced it. That invalidated the layer and the dependent layers after it. The cache helper had become part of the cache key.

Docker layer caching is not completely all-or-nothing. Earlier and independent work can still be reused. But within one dependent chain, once a layer is invalidated, the following layers rebuild. My workaround could preserve a small inner cache while invalidating a much larger Docker layer cache. It was not a sensible default. But one part of the idea was still useful. Expensive work inside a rebuilt layer can have its own cache. It just should not require including the BoringCache binary in the image graph.

BuildKit state

I then tried moving part of the BuildKit root itself. The same-source build worked: save the portable state, restore it on a fresh builder, then run the same source with every useful step cached. The changed build was the problem. The Rust sidecar could move files, but it did not understand the cache graph the way BuildKit did. I eventually made the restore correct, but a small source change could still make it walk and repack several gigabytes. In paired PostHog runs, the ordinary exporter finished first. The state backend restored correctly but made changed builds slower than the ordinary exporter, so I deleted it. I wrote the full story in The BuildKit backend we deleted after it worked. What I learned was that BuildKit had to stay responsible for whether a cache record was valid. The sidecar should not rebuild a graph BuildKit already understood.

The native BuildKit backend

The next version became the native backend behind boringcache docker. BuildKit still decides which records are valid, which refs belong to the solve, how cache keys match and what the final manifest contains. BoringCache handles the destination: checking which immutable bodies already exist, moving the missing ones, retrying safely and publishing only after the referenced content is durable. That made parallel preparation possible. As immutable refs become available during the solve, the backend can prepare their bodies instead of leaving all of that work for the end. Preparing all blobs concurrently could compete with the build for CPU and I/O. A governor watches CPU and I/O pressure, limits how many blobs it prepares while the build and image output are active, and allows more concurrent preparation once only cache work remains.

Before reading or materialising a layer blob, the backend also asks whether that exact digest and size already exist at the destination. A hit means there is nothing to read, compress or upload again. A miss follows the normal bounded upload path. BuildKit still has to finish the cache graph and publish its manifest. But most of the expensive body work no longer has to begin after the build completes.

Image export

Preparing cache content also reduced the remaining work for image output. Cache blobs and image layers often need the same underlying content. By preparing that content during the solve, the backend also prepares content the image exporter would otherwise process near the end. When the build finishes, both the cache manifest and a requested image output can have much less left to do.

On three changed Formbricks builds, BoringCache's cache export stayed between 2.6 and 2.9 seconds while the GitHub Actions backend took between 179.8 and 274.7 seconds. Across those samples, the average complete measured build path was 2.85 times faster. Those are workload-specific runs, not a promise for every Dockerfile, but they show why export time after the build steps matters. The exact runs are 30089375174, 30090229757 and 30090831887. The broader Docker benchmark results include the workload and product versions beside the numbers. Builds with a large cache graph and a long cache export after the build steps finish benefit most. If most of the time is spent doing uncached build work, there is less export time to remove.

Caches inside the Docker build

The archive experiment identified two useful caches inside a Docker build step. The first is the tool's own remote cache. If a rebuilt RUN invokes sccache, ccache, Turbo, Nx, Gradle, Maven, Bazel or Go, the BoringCache CLI can give that tool its remote-cache configuration for the build session. It is not stored in the image and does not become part of the Docker layer key. If the Docker layer hits, the tool never runs. If the layer misses, the tool can still reuse its own partial work.

The second is a BuildKit cache mount. Cache mounts persist mutable directories such as a Cargo target/ between builds on one builder. That is one reason builds on persistent runners can finish sooner. BoringCache can offload selected mounts through the same content-addressed archive machinery and restore them on a fresh builder.

Terminal
$ boringcache docker

# Add these only when they reduce total build time:
$ boringcache docker --mount-cache --tool-cache sccache

Both stay opt-in because restoring cache also takes time. Saving and restoring a small npm cache can cost more than rebuilding it. A large Cargo target directory produced by a five-minute compile is a different case. In one maintained Rust proof, restoring a 1.1 GB target took about 10 seconds and the fresh Cargo check completed in under half a second. In a PostHog package-manager test, mount work added roughly 43 seconds and did not make the install step faster. The cache has to fit the workload.

A public Proteus build shows the caches composing rather than replacing one another. A source change meant the Rust build steps still had to run. Before they ran, BoringCache restored two BuildKit target mounts in 3.277 and 4.204 seconds. Inside the controller build, sccache served two of its three compiler requests from the remote cache. At the end, BuildKit reported no remaining preparation time for the cache export.

Proteus · linux/amd64
#26 boringcache cache mount hydrate ... status=hit total=3.277s
#29 boringcache cache mount hydrate ... status=hit total=4.204s
#29 Cache hits (Rust)   2
#29 Cache misses (Rust) 1
#34 preparing build cache for export 0.0s done

The build still did real compilation, and new work still missed. What I like about this log is that it shows the layers working together: the Docker layer rebuilt, the target directories kept useful state, sccache reused compiler output, and the native backend completed cache preparation before the build steps finished.

Measure cache export in your build

You do not need BoringCache to find out whether this is your problem. Start by asking Buildx for plain progress:

Terminal
$ export CACHE_REF=registry.example.com/app:buildcache
$ time docker buildx build \
  --progress=plain \
  --cache-from "type=registry,ref=$CACHE_REF" \
  --cache-to "type=registry,ref=$CACHE_REF,mode=max" \
  .

Find the last Dockerfile step, then look at the time spent in exporting to image, exporting cache and preparing build cache for export. The plain and raw JSON progress modes make those phases visible without the updating terminal UI. Then repeat the same commit, machine, builder and image output with cache import left on but --cache-to removed. The difference is a useful approximation of the publication cost. It does not isolate export cost completely because local state and registry conditions still matter, but it tells you whether export deserves further investigation.

On a recent Buildx, docker buildx history trace opens the latest build as an OpenTelemetry trace and can compare it with an earlier record. docker buildx du --verbose shows the size and shape of the selected builder's local cache. The latter does not measure remote export by itself, but it helps explain why a seemingly small Dockerfile can have a large cache graph behind it.

Keep the comparison conditions consistent: same source, same Dockerfile, same runner class, same output and comparable cache history. Record cold, changed and same-source runs separately. A same-source warm build proves that the cache can hit; a changed successor tells you whether it helps during ordinary development. The shorter Docker cache export guide covers mode=min, mode=max and the basic measurement sequence.

What changed

I recorded cache telemetry from the start because build time alone does not tell me whether cached results were reused. For Docker I needed to see import, executed and cached vertices, image-ready time, body preparation, bytes reused, bytes uploaded, final export and the CPU or I/O pressure while those phases overlapped. From those measurements, I decided to remove the state backend even though it worked. I also kept mount cache and tool cache optional. The largest time saving came from preparing cache during the build while BuildKit continued to decide which cache records were valid.

That is how Docker became the second BoringCache product. I started by trying to cache a few expensive instructions. I ended up moving cache preparation into the build itself and keeping useful tool and mount caches available when a layer rebuilt. BuildKit still publishes the final cache manifest, but by then there should be very little expensive work left to do.