Make each image an independently maintained component
Our internal images included patched third-party tools, base images, and CI runners. Their build processes varied by repository. Patching required manual investigation, and most images had no ARM build, limiting where they could run.
I built a shared image factory and led the migration of all 48 internal images in four days. The design gave each image its own manifest, build context, ownership, upstream version, and tests.
Use a manifest as the build contract
Adding an image means adding a directory and a configuration file. The factory discovers the component, validates it, and plans the affected builds. Architectures build in parallel, and unrelated images do not need to rebuild for every change.
This example packages a patched third-party tool. Internal identifiers are anonymized; the manifest retains ownership, architecture targets, upstream tracking, release tags, and test stages:
apiVersion: images.example.io/v1alpha1kind: Imagemetadata: name: cloudwatch-exporter owners: [platform-team] # A maintained patch of a third-party image. category: patchedspec: repository: maintained/cloudwatch-exporter defaults: platforms: [linux/amd64, linux/arm64] variants: - name: default build: context: . dockerfile: Dockerfile args: upstream_version: from: spec.upstream.version tags: primary: "2.0.0" aliases: [latest] tests: # The test stage must pass before publishing. buildTargets: [test] # Renovate watches this version for upstream releases. upstream: datasource: docker image: docker.io/prom/cloudwatch-exporter version: "v0.18.0"The primary version identifies our image release, while upstream.version identifies the third-party version it packages. A change to our Dockerfile can require a new primary version even when the upstream version stays the same.
Pull requests validate, build, and run the declared test stages without publishing. The release workflow publishes after merge. Renovate reads upstream versions from the manifest and proposes updates; publishing changed content also requires a new primary version.
Treat publishing as a separate correctness problem
A successful build is insufficient evidence that a registry now contains the intended release. The publisher verifies each architecture’s digest, assembles the multi-platform index, and checks that the published result matches the plan.
The failure policy distinguishes recoverable interruptions from conflicting content:
| Condition | Publisher behavior |
|---|---|
| Version already points to the same digest | Continue without replacing it |
| Version exists with a different digest | Stop; never overwrite the release |
| Index assembly fails transiently | Retry publication using existing verified artifacts |
| An alias update fails | Retry the alias step; retain the verified version |
| An architecture build fails | Require a new build run before release |
Version tags are immutable. Convenience aliases such as latest can move, but only to the verified release. Digest artifacts also carry plan and run identity, allowing the publisher to reject artifacts that belong to a different build plan.
Make migration repeatable for other engineers
I documented the conversion from the old builds into a playbook. Teammates could migrate components independently, which made the four-day rollout possible. The factory entered production in under a month and produced 35 releases in its first 20 days.
The implementation built on existing runners, dependency updates, review checks, and release automation. That foundation let this project concentrate on the image contract and publishing behavior.
What changed
Images gained consistent ownership, versioning, tests, and a traceable release path. Upstream patching could flow through automated pull requests, and ARM builds made those images usable on ARM compute. The migration also removed the need to understand a shared build script before adding or maintaining one image.