Portfolio · Notes · Dotfiles

Search everything

Search case studies, engineering notes, and Dotfiles documentation.

    all case studies

    Case study 04

    Teams that create themselves

    Engineering teams provision their collaboration tools, on-call schedules, and alert routing through a reviewed pull request backed by one team definition.

    My role
    Designed the team definition and built the Terraform and Terramate automation across the connected services.
    Evidence
    About 30 team stacks used the workflow; engineers outside the platform team could provision team resources and alert destinations themselves.
    devexdelivery

    Turn team setup into one reviewed change

    Creating an engineering team required separate requests for GitHub membership, Slack channels, on-call schedules, a service-catalog entry, and alert routing. Different administrators handled each system, so setup took several days and produced inconsistent results.

    I built a self-service workflow around one directory and one JSON definition per team. Terramate generates the Terraform configuration, and the existing plan-and-approval process provisions the requested resources.

    One engineering team can own several GitHub teams

    The directory identifies the engineering team, and its configuration selects the integrations it needs. The GitHub teams are a list: one engineering team can have a default group and separate frontend and backend groups, each with its own membership and repository access.

    This fictional team uses all three GitHub groups, a shared Slack channel, a service-catalog entry, alert channels, and on-call schedules in two time zones:

    teams/payments/team.json
    {
    "team_name_readable": "Payments",
    "portfolio": "Commerce",
    "github": {
    "teams": [
    { "type": "default", "description": "All Payments engineers" },
    { "type": "frontend", "description": "Payments web client" },
    { "type": "backend", "description": "Payments APIs and workers" }
    ]
    },
    "slack": {
    "create": true,
    "channel": "team-payments",
    "topic": "Payments services: questions, on-call and alerts"
    },
    "opslevel": {
    "create": true
    },
    "alerts": {
    "create_slack_channels": true
    },
    "firehydrant": {
    "create": true,
    "team_oncall_enabled": true,
    "schedules": [
    {
    "timezone": "Europe/London",
    "daily_start_time": "09:00:00",
    "daily_end_time": "17:00:00"
    },
    {
    "timezone": "America/New_York",
    "daily_start_time": "09:00:00",
    "daily_end_time": "17:00:00"
    }
    ]
    }
    }
    ONE TEAM DEFINITION · SEVERAL CONNECTED SERVICES
    Team Provisioning One Payments engineering team defines team.json. Terramate generates Terraform, which is planned and applied after approval. The definition provisions three GitHub teams: default, frontend and backend. It also provisions a Slack channel, a service-catalog entry, two on-call schedules in London and New York, and alert destinations selected by owner, environment and urgency. Payments One engineering team · team.json Terramate → Terraform Generate → plan → approved apply GitHub teams Three groups, one owner default All engineers frontend Web client backend APIs and workers Slack #team-payments Service catalog Payments ownership entry On-call schedules London · 09:00–17:00 New York · 09:00–17:00 Alert destinations Owner + environment + urgency
    EXHIBIT 01 — The Payments example creates three GitHub teams and two on-call schedules. The engineering team remains the common owner across the connected services.

    A schema documents the accepted attributes. The generated Terraform is separated by integration, so the plan exposes what will change in GitHub, Slack, the service catalog, and the on-call system. Identity-provider synchronization can manage membership for the GitHub groups. The two schedule entries let the engineering team define coverage across time zones within the same team definition.

    Include alert routing in the definition

    Provisioning channels alone would still leave a manual handoff. I connected team setup to alert routing so alerts for the team’s services reach its own destinations once the configuration is applied.

    The routing contract uses a team owner label together with environment and urgency. Shared routing rules consume those labels, while the team module creates the schedules, escalation policies, and channels they resolve to. That keeps new teams from requiring another bespoke set of routing rules.

    The generator walkthrough covers the configuration flow. As adoption grew, the FireHydrant provider’s rate limiting and state behavior became a separate reliability problem, addressed in the provider fork.

    What changed

    About 30 team stacks adopted the model. Engineers outside the platform team could request their own setup through a pull request, with a visible plan and a versioned definition of the connected resources. Alert routing became part of that setup instead of a follow-up ticket.