I did not want to fork BuildKit. BoringCache already had a CLI that could walk an archive, split it into content-addressed pieces, upload what was missing and restore it on another machine. I thought I could reuse most of that code without modifying BuildKit. I wanted a fresh GitHub Actions runner to restore the BuildKit state saved by the previous runner.
Persistent runners
Most GitHub-hosted jobs start on a new virtual machine. When the job finishes, the local BuildKit state goes away with it. The next job can import an external cache, but it cannot see the previous builder. A persistent runner avoids this. BuildKit keeps its own state on disk between builds, so there is nothing to export and import. It reads the state it wrote. Its cache mounts also remain available between builds unless garbage collection removes them.
But that state stays with the builder. If I build locally, move the job to another CI provider, use an ephemeral runner or deploy from another machine, the new builder cannot access the previous builder’s local state. I wanted to save enough of the BuildKit state after one job to restore it on a fresh builder for the next one.
The first subset of BuildKit state
I used PostHog because it has a large, real Docker build rather than a small cache fixture. In the first useful public run, the BuildKit root was 31.38 GB. About 26.69 GB was expanded snapshot data. The remaining 4.69 GB was the state I tried to move.
BuildKit root 31.38 GB snapshot directories 26.69 GB state without snapshots 4.69 GB save 34s restore to fresh builder 19s same-source build 20s · 68 cached steps changed successor failed · missing parent snapshot bucket
The same-source build initially made me think the saved state was sufficient.
The first changed build failed
The next source change broke it. I had removed the snapshot directories because they looked reconstructible, but the metadata still contained parent relationships that referred to them. The same source could reuse what was already there. A changed successor needed a parent snapshot that I had excluded from the saved state, and BuildKit correctly refused to continue. This showed me the difference between an archive and BuildKit state. A directory archive records which files it contains. BuildKit state is also a graph of cache records, results, parents, snapshots, leases and immutable content. A generic filesystem tool can copy those files, but it does not know which records make a valid portable cache.
Restoring the state in 19 seconds did not prove that the restored builder could build changed source correctly. From then on, I tested three things: restore on a fresh builder, rebuild the same source, and then build a changed successor.
The backend grew
I kept going. In the next design, BuildKit decided which records were eligible. Each generation was published atomically, and a failed build left the previous head in place. Metadata arrived first. Immutable layer bodies could stay remote until BuildKit asked for them. I also tried preparing the state during the build. As immutable refs appeared, the backend could prepare them while later Dockerfile steps were still running. A governor backed off when the build needed CPU or I/O, then used more of the machine when only cache work remained. By this point it was no longer a directory backup. It had lineage checks, signed generations, lazy range reads, cleanup rules, retry boundaries and a fresh-runner canary. The backend worked.
Larger runners
I also wanted to know how much a larger machine would help. The PostHog test ran the exact upstream Dockerfile on 4, 8 and 16-core AMD64 runners:
| Runner | Build | Total job | State overhead | Published state |
|---|---|---|---|---|
| AMD64 · 4 cores | 816s | 848s | 109.9s | 6.916 GB |
| AMD64 · 8 cores | 530s | 574s | 64.2s | 6.917 GB |
| AMD64 · 16 cores | 339s | 378s | 35.6s | 6.917 GB |
The larger runners helped. But four times the cores did not make the build four times faster: 816 seconds became 339. The state overhead also fell from 109.9 seconds to 35.6, but it did not disappear. A larger machine helped, but the state work was still there. The full numbers, including the ARM64 lane and exact source commits, are in the PostHog benchmark pull request.
The benchmark caught stale output
An earlier version composed the state backend with a Turbo remote cache inside the PostHog build. Changed commits added files under frontend/public/services, but PostHog's Turbo build inputs did not include frontend/public. A remote hit replayed the older frontend output. The cached lanes kept reporting 6,103 collected static files while an uncached lane produced 6,108 and then 6,110. The benchmark looked faster because one of its caches had skipped work it should have done. I turned Turbo off in the main benchmark runs and kept it only for diagnosing the issue. The state results then used the upstream Dockerfile without changing its behaviour. The benchmark then included the frontend work that the invalid cache hit had skipped. A fast cache that returns stale output is not a successful cache.
Changed builds were still slower
The final packs-based design could restore a generation lazily and produce 18 and 16-second same-source builds, with 76 cached vertices. This time the cache graph was valid, and a fresh builder could use it. Changed builds were the problem. A small source change could still make the state path process two to six gigabytes locally. In the measured PostHog cadence it recreated about 2.1 GB of pack bodies per realistic changed commit, even when only around 150 MB was new to the remote content store.
| Phase | Total build | State processing | State work after build |
|---|---|---|---|
| Cold PostHog | 610s | 30.7s | about 23s |
| Large endpoint change | 492s | 20.4s | about 22s |
| Same source | 18s / 16s | small | about 2.84s |
The roughly 500-second number was the whole changed build, not a 500-second state save. Much of the state work ran beside the build. Looking only at the final save would miss the CPU, filesystem reads, compression, range fetches and local writes already performed during the solve. The final PostHog state run has both the same-source hits and the slower changed builds. In the paired runs, realistic changed state builds averaged about 513.7 seconds. The simpler layer exporter averaged about 466.0 seconds. The backend was correct and fast when the source was unchanged. But on changed builds, it was still doing more work than the simpler exporter.
Would FastCDC help now?
It might reduce the bytes, but I do not know whether it would change the result. BoringCache archives now use FastCDC boundaries over a canonical tar stream. In one local Cargo experiment, a one-line rebuild changed a 24.7 GB target directory. Fixed chunks reused none of the shifted stream. FastCDC realigned after the changed regions, reused 96.9% of the decoded content and reduced the compressed successor payload from 3.58 GB to 97.6 MB. FastCDC reduced the compression and upload work a lot, but the client still has to walk the directory and produce one canonical tar stream before it can discover those reusable chunks.
BuildKit state adds another problem. It is not just a stable directory archive, and raw OCI or BuildKit layer tar streams can contain representation changes such as fresh timestamps that change chunk boundaries or content. A fresh builder still needs the valid graph reconstructed. Making all of that lazy would bring back the range reader, virtual content provider and generation lifecycle that made the deleted backend complicated.
I tried it again
I did end up running this experiment again. I still wanted to test whole-state caching, mostly because BoringCache has FastCDC now and it is much better at finding the parts of a large archive that have not really changed. The Cargo result made me wonder whether better chunk reuse would also improve the state backend. Maybe the old state backend had been doing too much work simply because it could not see that most of the bytes were still the same.
I did not bring the old backend back. I tried copying the complete state through the existing archive system. I stopped a normal BuildKit builder after the PostHog build, took its complete state exactly as it existed on disk, sent it through the archive system BoringCache already has, and restored it onto a fresh builder. I did not change the Dockerfile or add another state format inside BuildKit. I just wanted to see whether saving and restoring the complete state could provide similar reuse and build times to a persistent builder.
Some measurements suggested the approach might be useful. The BuildKit root was 35.92 GB across 885,800 filesystem entries. Its metadata-preserving archive was 36.91 GB, which FastCDC and compression reduced to 9.88 GB of unique remote content. Once the archive existed, creating its chunk graph took 16.3 seconds, and an exact repeat save finished in 15.4 seconds without uploading the same chunks again. The restored builder was also correct: all 74 Dockerfile steps were cached, the final image digest was identical, and the actual build took 2.6 seconds. Those measurements suggested efficient chunk reuse and a correct same-source restore.
Creating and restoring the archive took much longer than the 2.6-second build. Creating the archive from the stopped BuildKit volume took 606.9 seconds. Restoring it from BoringCache took 588.8 seconds, then putting the reconstructed state into a fresh BuildKit volume took another 225.3 seconds. The complete restored build took 816.7 seconds. Building PostHog cold took 485.6 seconds. The cache worked, but waiting for it was slower than doing the work again.
I also tried avoiding the intermediate archive file and streaming the chunks directly into the fresh BuildKit volume. That saves the extra 36.91 GB file, but it does not remove the work of recreating nearly 886,000 entries, directories, whiteouts and overlay attributes. I stopped that local test after it had run too long for the restored build to finish faster than a cold build. It was a deliberately simple implementation, so I do not use its unfinished time as a benchmark. The completed restore had already shown that this approach took longer than a cold build.
I thought about the other ways of moving the state too. We could copy filesystem blocks, take a volume snapshot, keep the data behind a lazy reader, or attach the same disk somewhere else. Each approach changes which data must be transferred and reconstructed. If we send BuildKit's portable blobs, BuildKit has to create them and later turn them back into local snapshots. If we send the snapshots as they already exist, we have to move and recreate a much larger filesystem. If we keep the same disk attached so that neither of those things happens, then we have made a persistent builder.
I realised these approaches mostly changed when and where state was prepared and reconstructed. FastCDC can make the rolling save smaller, and the exact repeat proved that it would not upload unchanged chunks again. I had written above that the follow-up needed a changed build. I did not get that far, so I still do not have a save-time measurement for a realistic source change. Restore alone had already taken longer than the cold build. A faster rolling save would not reduce that restore time. The experiment answered my question about whole-state restore, but it did not change my decision to remove the state backend.
What I kept
I did not keep the state backend. The current BoringCache BuildKit backend keeps BuildKit's normal cache graph and works with its portable immutable layer descriptors. It asks the destination whether an exact body already exists before opening or materialising it, and it prepares missing bodies during the solve when the machine has spare CPU or I/O capacity. I wrote about the backend I kept in Docker build done. Still exporting cache…. It removes most of the export work at the end. It does not try to recreate another machine's whole builder on a fresh runner.
A persistent runner still has an advantage because keeping local state avoids translation. But that state stays attached to the runner. BoringCache stores portable layer content so the cache can be used locally, in CI and on another machine. I needed a smaller portable unit, not every byte BuildKit kept on disk.
Why I deleted it
I liked the state backend. It had taken a lot of work and, by the end, it restored fresh builders, survived failed publications, loaded bodies lazily and handled the same-source build correctly. But I had built it to save build time, so correctness alone was not enough. On the changed builds that mattered, the simpler backend was faster and much easier to operate. So I deleted the working backend.
I think this is a normal part of building software. You build something, measure it properly, and sometimes find that you learned a lot from it but should not keep it. I used the results to guide the next backend. Disk size and portable content are not the same thing. A same-source hit is not enough to prove correctness. And if cache work moves into the build, I still have to count it. I kept the backend that reused portable layer content without restoring the full builder state.