In my first weeks I started changing that: I introduced Atlantis for cloud infrastructure and migrated all of Kubernetes into Argo CD, until every change became a reviewable pull request that applies itself. That created the safety to move fast.
Oleksandr Ponomarov
I’m a Platform & Site Reliability Engineer with several years owning the foundation used by a few hundred engineers.
My focus is AWS, Kubernetes, delivery, and reliability: making the right thing the easy thing so changes stay visible and reversible.
01 — Four years at an industrial IoT company
Each layer made the next one possible
For four years, I owned the platform foundation of an industrial IoT company. I joined when infrastructure changes were still applied by hand; by the end, a few hundred engineers were shipping through systems designed to make changes visible, reviewable, and reversible.
The work did not arrive as isolated projects. Each layer unlocked the next: reviewable changes made faster feedback safe; automation turned maintenance into background work; rebuilt foundations made self-service possible; and those foundations supported the newest chapter in secure AI tooling and customer-facing platforms.
This is that four-year story in the order it happened. Each label opens the case study behind the step:
Then the feedback loop got fast: checks that run in milliseconds before you even commit, and fix your mistakes for you instead of just complaining.
With commits standardized, releases became automatic; with releases automatic, dependency updates became automatic; with updates automatic, staying current stopped being a project and became a background process.
On that foundation we could rebuild the hard things: the network, from about 15 module copies into segments defined in one policy, and then entire environments that can be created and destroyed with one command.
The platform proved itself twice over: when the company made an acquisition, one engineer could move the acquired product onto it in a summer; when a vendor raised prices, we could leave in weeks.
And because all of that existed, the newest chapter — AI tooling for every developer, customer code running as self-service functions, and building new platforms in weeks instead of quarters — was possible at all. The container supply chain below went from nothing to production in under a month, precisely because the pull request automation, releases, runners and hooks it depended on were already there.
02
The biggest win
Environments you can create and destroy with one command
- Impact
- A complete environment became one Terraform stack, with GitOps bootstrap and optional network attachment. A custom provider verifies cleanup before allowing cluster and VPC destruction to continue.
- My role
- Designed the cell architecture with the platform team and wrote the provider that coordinates and verifies teardown.
- Evidence
- One stack provisions a cell with more than 20 available platform add-ons; teardown records progress and reports surviving resources when verification fails.
One stack provisions a cell; a custom lifecycle provider checks controller-created cloud resources before cluster and network destruction can proceed.
03
All case studies
- Impact
- Infrastructure changes gained a shared review process: Terraform plans and Kubernetes previews in pull requests, approved applies through Atlantis, and Kubernetes deployment on merge.
- My role
- Introduced Atlantis and Argo CD, built the Kubernetes diff bot, approval workflow, and Atlantis browser controls, and led the staged rollout of automatic deployment.
- Evidence
- Four Kubernetes upgrades in the first six weeks; a diff bot used for three and a half years; infrastructure changes reviewed and applied through shared tooling.
- Impact
- A matched benchmark reduced fifty-file formatting from 356 to 74 milliseconds. Shared checks fix formatting and regenerate documentation before review.
- My role
- Introduced pre-commit checks, built the shared hook library, migrated repositories to
prek, and optimized hooks and CI runners. - Evidence
- A seven-run benchmark of the old and replacement hooks measured 4.8× faster formatting for fifty files and 17.9× for all 1,793 matching handwritten Terraform files.
- Impact
- Repository tool definitions became the shared input for laptop setup and CI. One install command replaced manual setup, with cache isolation between repositories.
- My role
- Introduced
asdf, later evaluated and adoptedmise, and rebuilt the shared pre-commit workflow around repository tool definitions. - Evidence
- Local setup and CI both use
mise.toml; node-local caches are separated by repository, and fork pull requests skip that shared cache path.
- Impact
- Engineering teams provision their collaboration tools, on-call schedules, and alert routing through a reviewed pull request backed by one team definition.
- My role
- Designed the team definition and built the Terraform and Terramate automation across the connected services.
- Evidence
- About 30 team stacks used the workflow; engineers outside the platform team could provision team resources and alert destinations themselves.
- Impact
- The shared Terraform repository grew from dozens of stacks to hundreds with remote state, automated module releases, consistent structure, and an inventory engineers could navigate.
- My role
- Led the repository improvements, designed the Terramate layout and release workflow, built the stack explorer, and supported adoption through a Terraform community channel.
- Evidence
- Migrated every stack to remote state; introduced versioned module releases and a terminal inventory spanning accounts, regions, and environments.
- Impact
- Dependency updates became a continuous flow of small pull requests, supported by automated releases, deployment previews, and a supervised review playbook.
- My role
- Deployed Renovate, established the module release automation, reviewed infrastructure updates, and encoded recurring review decisions in a supervised agent skill.
- Evidence
- Restored clean planning for about a dozen infrastructure stacks; cleared a months-old update backlog in weeks with recorded evidence for each merge.
- Impact
- About 550 live Kafka topics gained versioned definitions and named owners without recreation. Admission policies then constrained changes that could bypass the reviewed workflow.
- My role
- Built the migration generator, designed the environment layout, ran the adoption, and introduced the admission policies.
- Evidence
- About 550 topics adopted without recreation; replica changes and partition reductions rejected at admission; production topic ownership synchronized to the service catalog.
- Impact
- Security-image builds and scheduled node replacement became one workflow: automatic adoption in lower environments, with an explicit image-version review before production rollout.
- My role
- Designed the image and rotation policy, coordinated its review, and co-built the event-driven image pipeline with two teammates.
- Evidence
- Lower environments adopt new images within hours when rotation windows allow; production pins image IDs and promotes them through reviewed pull requests.
- Impact
- A fragile route model became safe to change through explicit resource identities. A separate Cloud WAN design then expressed network segmentation and attachment rules in one policy.
- My role
- Rebuilt the Terraform route model, proved the state migration with a no-op plan, and proposed the Cloud WAN design adopted by the team.
- Evidence
- A no-op plan verified the routing refactor; the Cloud WAN policy defines four segments and sends unmatched attachments to quarantine.
- Impact
- A complete environment became one Terraform stack, with GitOps bootstrap and optional network attachment. A custom provider verifies cleanup before allowing cluster and VPC destruction to continue.
- My role
- Designed the cell architecture with the platform team and wrote the provider that coordinates and verifies teardown.
- Evidence
- One stack provisions a cell with more than 20 available platform add-ons; teardown records progress and reports surviving resources when verification fails.
- Impact
- NAT instances became the default egress path with managed-gateway fallback. I contributed the deployment changes needed to run the failover function without a container build.
- My role
- Adopted alterNAT, integrated it into the cell platform, and contributed Zip packaging and dependency removal upstream.
- Evidence
- Two merged upstream pull requests; the production configuration uses the Zip deployment path and checks connectivity every minute.
- Impact
- I migrated six services from Heroku to the shared AWS platform in about two months while the acquired team continued shipping, with a rehearsed database move and reversible DNS cutover.
- My role
- Led and implemented the migration, including service deployment, supporting infrastructure, database rehearsals, and cutover.
- Evidence
- Six services migrated in about two months; the old hosting account closed; the acquired team had shared monitoring and its own alert channel at handover.
- Impact
- Registry mirrors and admission-time image rewrites let Kubernetes workloads move to our own registry before teams changed their manifests. The Docker Hub subscription was then retired.
- My role
- Proposed and implemented the migration, including pull-through caches, staged rewrite policies, manifest updates, and subscription retirement.
- Evidence
- The mirror served hundreds of images within weeks; the subscription was cancelled, and the rewrite rule remained as a fallback for old references.
- Impact
- I fixed a Terraform provider that was blocking team automation, then built a registry to distribute the fork through normal version pins and signed releases.
- My role
- Forked and fixed the FireHydrant provider, reported the issues upstream, and established the internal registry and publishing workflow.
- Evidence
- Provider fixes completed in four days, with the registry built in parallel in three; about 30 team stacks used the fork.
- Impact
- I built a shared image pipeline and led the migration of 48 images in four days. Releases build for x86 and ARM, run tests, and verify published artifacts against their digests.
- My role
- Designed and built the image factory, established its publishing contract, and created the migration playbook used by teammates.
- Evidence
- 48 images migrated in four days; 35 releases in the first 20 days; version conflicts, architecture mismatches, and publishing retries covered by tests.
- Impact
- Customer extensions move from upload to a running Lambda through automated builds and a small provisioning API, with function-specific roles, authenticated invocation, and logs.
- My role
- Implemented the selected AWS architecture, introduced Crossplane, and built the infrastructure, build workflow, and provisioning controls.
- Evidence
- The SDK service provisions functions through a Kubernetes claim; the composition supplies a per-function IAM role, IAM-authenticated function URL, and log group.
- Impact
- AI assistants gained a shared route to operational tools using developers’ existing AWS identities. Backend restrictions and repository playbooks keep access and infrastructure changes within defined controls.
- My role
- Built the MCP gateway and signing proxy, configured restricted backends, and wrote repository guidance and supervised automation playbooks.
- Evidence
- Gateway deployed across four environments; no additional per-tool credentials for developers; production deployment access is read-only and infrastructure changes use existing review gates.
04
How I work
Decisions get written down.
I introduced architecture decision records to the organization and authored a large share of them — including the standard for how to write them. When we chose the tool that now manages all our infrastructure code, I didn’t write an opinion piece: I built the same infrastructure three ways in working prototypes and let the comparison decide.
I teach what I build.
Standards arrived with a community channel, recorded walkthroughs, worked examples, and patient answers to beginner questions.
Measure, don't assert.
Benchmarks before rewrites, no-op-plan proofs before risky migrations, restore drills before trusting backups, rehearsals before destroys.
Automation over heroics.
Large mechanical migrations run as scripted, repeatable campaigns — and in the last year, as supervised AI-agent campaigns with safety rules encoded in — turning fleet-wide changes from quarter-long projects into focused weeks.