Connect image publishing to node replacement
We were building patched Kubernetes node images, but the running nodes were not consistently receiving them. Some replacement budgets were zero; one production pool had only a two-hour window each week. Publishing an image did not complete the patching process.
I designed the connection between image builds and node rotation, with seven colleagues reviewing the policy. Two teammates and I built the event-driven image pipeline. The design had to deliver patches while keeping planned disruption within the team’s support capacity.
Build once, promote deliberately
The pipeline responds to OS security-release notifications and changes to the upstream Kubernetes node image. It builds the patched image and publishes it for each environment. Security updates are applied while the Kubernetes runtime components remain aligned with the upstream node image.
Development and staging select new images by name pattern. Karpenter detects that existing nodes no longer match the desired image and schedules replacements. Production pins exact image IDs; the dependency bot proposes each change as a pull request.
That creates a clear promotion boundary. Lower environments exercise the image first, while production adoption requires a reviewed configuration change.
Schedule disruption around the workload
For ordinary node pools, planned image rotation runs during weekday on-call coverage and stops two hours before coverage ends. The schedules use the overlap between summer and winter coverage because the scheduler evaluates them in UTC.
A reduced budget example shows how the policy combines a concurrency cap with blocked periods:
budgets: - nodes: "15%" reasons: [Drifted] - nodes: "0" reasons: [Drifted] schedule: "0 18 * * 1-5" duration: 13h - nodes: "0" reasons: [Drifted] schedule: "0 0 * * 6" duration: 55hThis blocks drift-based replacement overnight and through the weekend. It controls planned image rotation; it does not prevent unrelated node failures.
Stateful workloads need different policies. Metrics nodes rotate one at a time. Database pools use separate windows by availability zone. Development Kafka brokers use a weekend window to avoid interrupting engineers with rebalances during the week. One streaming workload blocks drift-based replacement because its operator coordinates job migration itself.
Pod disruption budgets constrain which replicas can drain together. An alert catches terminations stuck for more than 12 hours and links to a runbook covering blocked disruption, unhealthy pods, and attached volumes.
What changed
Patching gained a complete path from a vendor release to running nodes. Lower environments can rotate within hours when their windows allow; production adoption follows a reviewed image change and its pool’s schedule. The policy makes both promotion and the permitted pace of disruption visible in code.