Portfolio · Notes · Dotfiles

Search everything

Search case studies, engineering notes, and Dotfiles documentation.

    Oleksandr Ponomarov

    I’m a Platform & Site Reliability Engineer with several years owning the foundation used by a few hundred engineers.

    My focus is AWS, Kubernetes, delivery, and reliability: making the right thing the easy thing so changes stay visible and reversible.

    01 — Four years at an industrial IoT company

    Each layer made the next one possible

    For four years, I owned the platform foundation of an industrial IoT company. I joined when infrastructure changes were still applied by hand; by the end, a few hundred engineers were shipping through systems designed to make changes visible, reviewable, and reversible.

    The work did not arrive as isolated projects. Each layer unlocked the next: reviewable changes made faster feedback safe; automation turned maintenance into background work; rebuilt foundations made self-service possible; and those foundations supported the newest chapter in secure AI tooling and customer-facing platforms.

    This is that four-year story in the order it happened. Each label opens the case study behind the step:

    01

    In my first weeks I started changing that: I introduced Atlantis for cloud infrastructure and migrated all of Kubernetes into Argo CD, until every change became a reviewable pull request that applies itself. That created the safety to move fast.

    02

    Then the feedback loop got fast: checks that run in milliseconds before you even commit, and fix your mistakes for you instead of just complaining.

    04

    On that foundation we could rebuild the hard things: the network, from about 15 module copies into segments defined in one policy, and then entire environments that can be created and destroyed with one command.

    05

    The platform proved itself twice over: when the company made an acquisition, one engineer could move the acquired product onto it in a summer; when a vendor raised prices, we could leave in weeks.

    06

    And because all of that existed, the newest chapter — AI tooling for every developer, customer code running as self-service functions, and building new platforms in weeks instead of quarters — was possible at all. The container supply chain below went from nothing to production in under a month, precisely because the pull request automation, releases, runners and hooks it depended on were already there.

    03

    All case studies

    01Making infrastructure changes boring
    Impact
    Infrastructure changes gained a shared review process: Terraform plans and Kubernetes previews in pull requests, approved applies through Atlantis, and Kubernetes deployment on merge.
    My role
    Introduced Atlantis and Argo CD, built the Kubernetes diff bot, approval workflow, and Atlantis browser controls, and led the staged rollout of automatic deployment.
    Evidence
    Four Kubernetes upgrades in the first six weeks; a diff bot used for three and a half years; infrastructure changes reviewed and applied through shared tooling.
    deliverydevexreliabilitysecurityread →
    02A feedback loop measured in milliseconds
    Impact
    A matched benchmark reduced fifty-file formatting from 356 to 74 milliseconds. Shared checks fix formatting and regenerate documentation before review.
    My role
    Introduced pre-commit checks, built the shared hook library, migrated repositories to prek, and optimized hooks and CI runners.
    Evidence
    A seven-run benchmark of the old and replacement hooks measured 4.8× faster formatting for fifty files and 17.9× for all 1,793 matching handwritten Terraform files.
    devexdeliveryread →
    03 · sequelOne tool version everywhere
    Impact
    Repository tool definitions became the shared input for laptop setup and CI. One install command replaced manual setup, with cache isolation between repositories.
    My role
    Introduced asdf, later evaluated and adopted mise, and rebuilt the shared pre-commit workflow around repository tool definitions.
    Evidence
    Local setup and CI both use mise.toml; node-local caches are separated by repository, and fork pull requests skip that shared cache path.
    devexdeliveryread →
    04Teams that create themselves
    Impact
    Engineering teams provision their collaboration tools, on-call schedules, and alert routing through a reviewed pull request backed by one team definition.
    My role
    Designed the team definition and built the Terraform and Terramate automation across the connected services.
    Evidence
    About 30 team stacks used the workflow; engineers outside the platform team could provision team resources and alert destinations themselves.
    devexdeliveryread →
    05Turning a Terraform repository into a product
    Impact
    The shared Terraform repository grew from dozens of stacks to hundreds with remote state, automated module releases, consistent structure, and an inventory engineers could navigate.
    My role
    Led the repository improvements, designed the Terramate layout and release workflow, built the stack explorer, and supported adoption through a Terraform community channel.
    Evidence
    Migrated every stack to remote state; introduced versioned module releases and a terminal inventory spanning accounts, regions, and environments.
    devexdeliveryread →
    06Dependency updates: from quarterly panic to background noise
    Impact
    Dependency updates became a continuous flow of small pull requests, supported by automated releases, deployment previews, and a supervised review playbook.
    My role
    Deployed Renovate, established the module release automation, reviewed infrastructure updates, and encoded recurring review decisions in a supervised agent skill.
    Evidence
    Restored clean planning for about a dozen infrastructure stacks; cleared a months-old update backlog in weeks with recorded evidence for each merge.
    securityreliabilityairead →
    07Kafka topics as code: adopting 550 live topics
    Impact
    About 550 live Kafka topics gained versioned definitions and named owners without recreation. Admission policies then constrained changes that could bypass the reviewed workflow.
    My role
    Built the migration generator, designed the environment layout, ran the adoption, and introduced the admission policies.
    Evidence
    About 550 topics adopted without recreation; replica changes and partition reductions rejected at admission; production topic ownership synchronized to the service catalog.
    reliabilitydeliverysecurityread →
    08 · featuredThe fleet that patches itself while engineers are watching
    Impact
    Security-image builds and scheduled node replacement became one workflow: automatic adoption in lower environments, with an explicit image-version review before production rollout.
    My role
    Designed the image and rotation policy, coordinated its review, and co-built the event-driven image pipeline with two teammates.
    Evidence
    Lower environments adopt new images within hours when rotation windows allow; production pins image IDs and promotes them through reviewed pull requests.
    securityreliabilityread →
    09 · featuredRebuilding the network as one reviewable policy
    Impact
    A fragile route model became safe to change through explicit resource identities. A separate Cloud WAN design then expressed network segmentation and attachment rules in one policy.
    My role
    Rebuilt the Terraform route model, proved the state migration with a no-op plan, and proposed the Cloud WAN design adopted by the team.
    Evidence
    A no-op plan verified the routing refactor; the Cloud WAN policy defines four segments and sends unmatched attachments to quarantine.
    networkingreliabilityread →
    10 · featuredEnvironments you can create and destroy with one command
    Impact
    A complete environment became one Terraform stack, with GitOps bootstrap and optional network attachment. A custom provider verifies cleanup before allowing cluster and VPC destruction to continue.
    My role
    Designed the cell architecture with the platform team and wrote the provider that coordinates and verifies teardown.
    Evidence
    One stack provisions a cell with more than 20 available platform add-ons; teardown records progress and reports surviving resources when verification fails.
    reliabilitycostdeliveryfull story ↑
    11The NAT bill and the fix that went upstream
    Impact
    NAT instances became the default egress path with managed-gateway fallback. I contributed the deployment changes needed to run the failover function without a container build.
    My role
    Adopted alterNAT, integrated it into the cell platform, and contributed Zip packaging and dependency removal upstream.
    Evidence
    Two merged upstream pull requests; the production configuration uses the Zip deployment path and checks connectivity every minute.
    costnetworkingread →
    12Absorbing an acquisition
    Impact
    I migrated six services from Heroku to the shared AWS platform in about two months while the acquired team continued shipping, with a rehearsed database move and reversible DNS cutover.
    My role
    Led and implemented the migration, including service deployment, supporting infrastructure, database rehearsals, and cutover.
    Evidence
    Six services migrated in about two months; the old hosting account closed; the acquired team had shared monitoring and its own alert channel at handover.
    deliveryreliabilitycostread →
    13Leaving Docker Hub without a flag day
    Impact
    Registry mirrors and admission-time image rewrites let Kubernetes workloads move to our own registry before teams changed their manifests. The Docker Hub subscription was then retired.
    My role
    Proposed and implemented the migration, including pull-through caches, staged rewrite policies, manifest updates, and subscription retirement.
    Evidence
    The mirror served hundreds of images within weeks; the subscription was cancelled, and the rewrite rule remained as a fallback for old references.
    costreliabilitysecurityread →
    14The fork that needed a home
    Impact
    I fixed a Terraform provider that was blocking team automation, then built a registry to distribute the fork through normal version pins and signed releases.
    My role
    Forked and fixed the FireHydrant provider, reported the issues upstream, and established the internal registry and publishing workflow.
    Evidence
    Provider fixes completed in four days, with the registry built in parallel in three; about 30 team stacks used the fork.
    reliabilitydeliveryread →
    15 · featuredTurning container images from a liability into a supply chain
    Impact
    I built a shared image pipeline and led the migration of 48 images in four days. Releases build for x86 and ARM, run tests, and verify published artifacts against their digests.
    My role
    Designed and built the image factory, established its publishing contract, and created the migration playbook used by teammates.
    Evidence
    48 images migrated in four days; 35 releases in the first 20 days; version conflicts, architecture mismatches, and publishing retries covered by tests.
    securitycostdeliveryread →
    16Self-service cloud functions for customer code
    Impact
    Customer extensions move from upload to a running Lambda through automated builds and a small provisioning API, with function-specific roles, authenticated invocation, and logs.
    My role
    Implemented the selected AWS architecture, introduced Crossplane, and built the infrastructure, build workflow, and provisioning controls.
    Evidence
    The SDK service provisions functions through a Kubernetes claim; the composition supplies a per-function IAM role, IAM-authenticated function URL, and log group.
    securitydeliveryread →
    17 · featuredSafe AI tooling for every developer
    Impact
    AI assistants gained a shared route to operational tools using developers’ existing AWS identities. Backend restrictions and repository playbooks keep access and infrastructure changes within defined controls.
    My role
    Built the MCP gateway and signing proxy, configured restricted backends, and wrote repository guidance and supervised automation playbooks.
    Evidence
    Gateway deployed across four environments; no additional per-tool credentials for developers; production deployment access is read-only and infrastructure changes use existing review gates.
    aisecuritydevexread →

    04

    How I work

    01

    Decisions get written down.

    I introduced architecture decision records to the organization and authored a large share of them — including the standard for how to write them. When we chose the tool that now manages all our infrastructure code, I didn’t write an opinion piece: I built the same infrastructure three ways in working prototypes and let the comparison decide.

    02

    I teach what I build.

    Standards arrived with a community channel, recorded walkthroughs, worked examples, and patient answers to beginner questions.

    03

    Measure, don't assert.

    Benchmarks before rewrites, no-op-plan proofs before risky migrations, restore drills before trusting backups, rehearsals before destroys.

    04

    Automation over heroics.

    Large mechanical migrations run as scripted, repeatable campaigns — and in the last year, as supervised AI-agent campaigns with safety rules encoded in — turning fleet-wide changes from quarter-long projects into focused weeks.

    Case study reader

    Choose a case study to read.