Portfolio · Notes · Dotfiles

Search everything

Search case studies, engineering notes, and Dotfiles documentation.

    all case studies

    Case study 08 · featured

    The fleet that patches itself while engineers are watching

    Security-image builds and scheduled node replacement became one workflow: automatic adoption in lower environments, with an explicit image-version review before production rollout.

    My role
    Designed the image and rotation policy, coordinated its review, and co-built the event-driven image pipeline with two teammates.
    Evidence
    Lower environments adopt new images within hours when rotation windows allow; production pins image IDs and promotes them through reviewed pull requests.
    securityreliability

    Connect image publishing to node replacement

    We were building patched Kubernetes node images, but the running nodes were not consistently receiving them. Some replacement budgets were zero; one production pool had only a two-hour window each week. Publishing an image did not complete the patching process.

    I designed the connection between image builds and node rotation, with seven colleagues reviewing the policy. Two teammates and I built the event-driven image pipeline. The design had to deliver patches while keeping planned disruption within the team’s support capacity.

    Build once, promote deliberately

    The pipeline responds to OS security-release notifications and changes to the upstream Kubernetes node image. It builds the patched image and publishes it for each environment. Security updates are applied while the Kubernetes runtime components remain aligned with the upstream node image.

    EXHIBIT — BUILD SIDE · EVENT-DRIVEN
    OS security releasepublic notification feedUpstream EKS image changepublic parameter, polledImage buildsecurity updates onlyruntime, kubelet, kernel untouchedPatched image, datedPublished per environment

    Development and staging select new images by name pattern. Karpenter detects that existing nodes no longer match the desired image and schedules replacements. Production pins exact image IDs; the dependency bot proposes each change as a pull request.

    That creates a clear promotion boundary. Lower environments exercise the image first, while production adoption requires a reviewed configuration change.

    EXHIBIT — CONSUMPTION SIDE · SCHEDULED DRIFT
    dev / stagingname-pattern match → instant driftproductionpinned ID → bot opens pull request→ human mergesBudget-limited rotationcoverage-hours schedule · per-pool exceptionsDisruption-rule-gated drainPatched fleetStuck terminationalert + runbookstuck over 12 h

    Schedule disruption around the workload

    For ordinary node pools, planned image rotation runs during weekday on-call coverage and stops two hours before coverage ends. The schedules use the overlap between summer and winter coverage because the scheduler evaluates them in UTC.

    A reduced budget example shows how the policy combines a concurrency cap with blocked periods:

    nodepool/disruption.yaml
    budgets:
    - nodes: "15%"
    reasons: [Drifted]
    - nodes: "0"
    reasons: [Drifted]
    schedule: "0 18 * * 1-5"
    duration: 13h
    - nodes: "0"
    reasons: [Drifted]
    schedule: "0 0 * * 6"
    duration: 55h

    This blocks drift-based replacement overnight and through the weekend. It controls planned image rotation; it does not prevent unrelated node failures.

    Stateful workloads need different policies. Metrics nodes rotate one at a time. Database pools use separate windows by availability zone. Development Kafka brokers use a weekend window to avoid interrupting engineers with rebalances during the week. One streaming workload blocks drift-based replacement because its operator coordinates job migration itself.

    Pod disruption budgets constrain which replicas can drain together. An alert catches terminations stuck for more than 12 hours and links to a runbook covering blocked disruption, unhealthy pods, and attached volumes.

    What changed

    Patching gained a complete path from a vendor release to running nodes. Lower environments can rotate within hours when their windows allow; production adoption follows a reviewed image change and its pool’s schedule. The policy makes both promotion and the permitted pace of disruption visible in code.