Make the full environment lifecycle repeatable
Creating an environment required weeks of work across networking, Kubernetes, identity, secrets, DNS, and delivery. Deleting it was harder than reversing those steps: Kubernetes controllers created cloud resources that Terraform did not directly manage, leaving volumes, load balancers, and instances behind.
The business driver was expansion into new regions. I designed a cell architecture and built it with the platform team: a repeatable unit containing its network, cluster, and platform services. The same unit could support temporary test environments or dedicated customer environments where needed.
Give Terraform and GitOps explicit responsibilities
A cell starts as one Terraform stack. A data-only metadata module validates its inputs, resolves network sizing, and supplies consistent names and access settings. Terraform provisions the VPC, EKS cluster, networking prerequisites, and supporting cloud resources.
Terraform then installs the Flux operator and declares the Flux runtime. Flux reads the cell’s generated configuration and installs the enabled GitOps-managed add-ons. The catalog includes more than 20 add-ons; a cell selects what it needs rather than receiving every component automatically.
Metadata module. Validated inputs, naming, sizing, and derived configuration
Terraform. Cloud resources, cluster prerequisites, Flux bootstrap, and cell metadata
Flux. Reconciliation of enabled platform add-ons from Git
Teardown provider. Ordered cleanup and verification of supported controller-created resources
Cluster metadata crosses the Terraform-to-Flux boundary through a generated ConfigMap. Environment and region variations use reusable Kustomize components. Guardrails reject unresolved substitutions in generated paths, so missing metadata surfaces as a reconciliation failure.
Network attachment and NAT optimization are explicit options. The network attachment manager also maintains a registry consumed by network and secrets automation. Its key includes account, region, and cluster identity so equal names in different locations cannot overwrite one another.
Make teardown an enforced dependency
I wrote a Terraform provider to close the gap between deleting declared infrastructure and cleaning up the resources its controllers had created. Its lifecycle resource runs before the cluster and the infrastructure it references are destroyed.
This reduced example shows that dependency relationship:
resource "cellops_cell_runtime" "this" { cluster_name = module.eks.cluster_name region = var.region_name vpc_owned = var.create_vpc require_clean_slate = true
guard = { vpc = local.vpc_id }
timeouts = { delete = "60m" }}- Preflight. Establish whether the cluster API is reachable
- Freeze. Pause reconciliation and remove webhooks that would obstruct deletion
- Evacuate. Remove workloads and let live controllers release their resources
- Decommission. Remove autoscaled capacity
- Reconcile. Find and delete supported AWS resources within the cluster’s ownership scope
- Verify. Re-scan, wait for deletions in progress, and report anything still present
If Kubernetes is unavailable, the provider continues to the AWS cleanup stages. Those final stages determine whether teardown can proceed. A resource still deleting is distinguished from one that is stuck; a failed verification stops destruction and lists the survivors.
Progress is checkpointed in Terraform’s private state. A retry can resume completed work, but each phase is also safe to re-enter if the checkpoint is unavailable. During normal refresh, the provider reports attached resources outside its tag-based discovery scope, exposing potential cleanup problems before a destroy is requested.
What changed
Provisioning gained a reusable interface, and teardown gained explicit completion criteria. New regional environments could use a reviewed stack definition. Failed cleanup remained visible and recoverable instead of being hidden behind a successful-looking infrastructure deletion.