Portfolio · Notes · Dotfiles

Search everything

Search case studies, engineering notes, and Dotfiles documentation.

    all case studies

    Case study 10 · featured

    Environments you can create and destroy with one command

    A complete environment became one Terraform stack, with GitOps bootstrap and optional network attachment. A custom provider verifies cleanup before allowing cluster and VPC destruction to continue.

    My role
    Designed the cell architecture with the platform team and wrote the provider that coordinates and verifies teardown.
    Evidence
    One stack provisions a cell with more than 20 available platform add-ons; teardown records progress and reports surviving resources when verification fails.
    reliabilitycostdelivery

    Make the full environment lifecycle repeatable

    Creating an environment required weeks of work across networking, Kubernetes, identity, secrets, DNS, and delivery. Deleting it was harder than reversing those steps: Kubernetes controllers created cloud resources that Terraform did not directly manage, leaving volumes, load balancers, and instances behind.

    The business driver was expansion into new regions. I designed a cell architecture and built it with the platform team: a repeatable unit containing its network, cluster, and platform services. The same unit could support temporary test environments or dedicated customer environments where needed.

    Give Terraform and GitOps explicit responsibilities

    A cell starts as one Terraform stack. A data-only metadata module validates its inputs, resolves network sizing, and supplies consistent names and access settings. Terraform provisions the VPC, EKS cluster, networking prerequisites, and supporting cloud resources.

    Terraform then installs the Flux operator and declares the Flux runtime. Flux reads the cell’s generated configuration and installs the enabled GitOps-managed add-ons. The catalog includes more than 20 add-ons; a cell selects what it needs rather than receiving every component automatically.

    Metadata module. Validated inputs, naming, sizing, and derived configuration

    Terraform. Cloud resources, cluster prerequisites, Flux bootstrap, and cell metadata

    Flux. Reconciliation of enabled platform add-ons from Git

    Teardown provider. Ordered cleanup and verification of supported controller-created resources

    Cluster metadata crosses the Terraform-to-Flux boundary through a generated ConfigMap. Environment and region variations use reusable Kustomize components. Guardrails reject unresolved substitutions in generated paths, so missing metadata surfaces as a reconciliation failure.

    Network attachment and NAT optimization are explicit options. The network attachment manager also maintains a registry consumed by network and secrets automation. Its key includes account, region, and cluster identity so equal names in different locations cannot overwrite one another.

    Make teardown an enforced dependency

    I wrote a Terraform provider to close the gap between deleting declared infrastructure and cleaning up the resources its controllers had created. Its lifecycle resource runs before the cluster and the infrastructure it references are destroyed.

    This reduced example shows that dependency relationship:

    cell/teardown.tf
    resource "cellops_cell_runtime" "this" {
    cluster_name = module.eks.cluster_name
    region = var.region_name
    vpc_owned = var.create_vpc
    require_clean_slate = true
    guard = {
    vpc = local.vpc_id
    }
    timeouts = {
    delete = "60m"
    }
    }
    1. Preflight. Establish whether the cluster API is reachable
    2. Freeze. Pause reconciliation and remove webhooks that would obstruct deletion
    3. Evacuate. Remove workloads and let live controllers release their resources
    4. Decommission. Remove autoscaled capacity
    5. Reconcile. Find and delete supported AWS resources within the cluster’s ownership scope
    6. Verify. Re-scan, wait for deletions in progress, and report anything still present
    CELL LIFECYCLE · CREATION AND VERIFIED TEARDOWN
    Cell Lifecycle Terraform creates the cluster and bootstraps Flux, whose add-ons create cloud resources through controllers. During destruction the lifecycle provider runs preflight, freeze, evacuate, decommission, reconcile and verify before Terraform can destroy the cluster and its owned infrastructure. If the cluster API is unavailable, cleanup proceeds to AWS reconciliation. Incomplete cleanup stops destruction, reports survivors and can be retried from checkpointed phases. CREATE · DECLARED INFRASTRUCTURE Terraform VPC + EKS + metadata Bootstrap Flux Flux Enabled platform add-ons Controllers Cloud resources: volumes, LBs, nodes DESTROY · LIFECYCLE PROVIDER RUNS FIRST 01 · Preflight Is the cluster API reachable? 02 · Freeze Pause reconciliation Unblock deletion 03 · Evacuate Remove workloads Controllers clean up 04 · Decommission Remove autoscaled capacity 05 · Reconcile AWS cleanup within ownership scope 06 · Verify Re-scan and wait Report survivors API unavailable → AWS cleanup clean incomplete Continue destruction Terraform removes cluster and owned infrastructure Stop + report survivors Retry resumes checkpointed phases; cleanup stays visible
    EXHIBIT 01 — The provider must complete cleanup before Terraform can remove the cluster and its owned infrastructure. The AWS phases still run when the cluster API is unavailable.

    If Kubernetes is unavailable, the provider continues to the AWS cleanup stages. Those final stages determine whether teardown can proceed. A resource still deleting is distinguished from one that is stuck; a failed verification stops destruction and lists the survivors.

    Progress is checkpointed in Terraform’s private state. A retry can resume completed work, but each phase is also safe to re-enter if the checkpoint is unavailable. During normal refresh, the provider reports attached resources outside its tag-based discovery scope, exposing potential cleanup problems before a destroy is requested.

    What changed

    Provisioning gained a reusable interface, and teardown gained explicit completion criteria. New regional environments could use a reviewed stack definition. Failed cleanup remained visible and recoverable instead of being hidden behind a successful-looking infrastructure deletion.