Make operational access consistent across AI clients
AI assistants became more useful for operations when they could inspect deployments, dashboards, and metrics. Connecting each developer’s assistant directly to every service would have scattered credentials and access configuration across laptops.
I built an MCP gateway with one endpoint per environment. MCP is the protocol assistants use to discover and call tools. Developers authenticate with their existing AWS identities, while the gateway handles the separate authentication needed to reach backend services.
The deployment covers four environments. Developers do not manage additional per-tool credentials; backend service credentials still exist and are managed centrally.
The committed client configuration makes that access practical. This OpenCode example uses an anonymized gateway hostname and AWS profile:
{ "mcp": { "platform-tools-dev": { "type": "local", "command": [ "uvx", "mcp-proxy-for-aws", "https://gateway.dev.example.com/mcp", "--service", "execute-api", "--region", "eu-west-1", "--profile", "development" ] } }}OpenCode launches the local proxy, which signs requests using the developer’s existing AWS profile. The same command works in Cursor’s MCP configuration. Connecting another environment changes the endpoint and profile; backend credentials remain managed by the gateway.
Enforce access in the gateway and backends
The client signs its request for API Gateway using AWS SigV4. A front-door Lambda removes the incoming signing headers and signs a new request for AgentCore. AgentCore then routes tool calls through OAuth-protected target APIs and VPC proxy functions to the internal services.
Re-signing is necessary because the client signature is bound to the public host and service. Passing it unchanged to AgentCore would fail authentication.
The backend configuration limits what a successfully authenticated caller can do:
Production deployment backend. Read-only tool access
Dashboard backend. Selected tool groups exposed; writes disabled
Developer access. Granted through the existing identity platform
Committed client configuration. Development enabled by default; production requires explicit opt-in
One upstream upgrade exposed a useful integration edge: the dashboard MCP server rejected requests carrying the gateway hostname. I adjusted its host check for the internal proxy path while retaining origin validation and gateway authentication. The distinction between these checks mattered more than simply making the error disappear.
Give agents the repository context they need
I wrote AGENTS.md guidance for the core platform repositories, focusing on information an agent could not infer from a single file: contracts consumed by other repositories, generated paths, access conventions, and changes that require coordination.
For example, a Vault role declared in the Terraform repository is consumed by external-secret resources in the Kubernetes repository. Renaming it can break a consumer without changing that consumer’s files. The guidance names that contract and points to both sides, so an agent can identify the coordination required before editing.
Nested guidance keeps component-specific rules close to the code. Committed client configurations point assistants at the gateway, making the configured tool path available when work starts.
These instructions guide agent behavior. The backend permissions and deployment controls enforce the access boundaries.
Encode recurring work in supervised playbooks
I also wrote playbooks for dependency updates, stack migrations, and image maintenance. Their limits follow the failure modes of each task:
Dependency update. Inspect the actual plan or rendered diff; stop after limited repair attempts
Stack migration. Copy state to the new backend; retain the old object; accept only the expected tag changes
Image maintenance. Publish changed content under a new immutable version
The stack-migration runbook makes the stopping condition concrete. Copy the state before opening the pull request, because Atlantis plans automatically and an empty destination would appear to need every resource created. Preserve module and provider versions during the move. The acceptance check permits only the expected tags and tags_all changes: zero creates, replacements, destroys, or other attribute changes. An unexpected diff starts an investigation into provider drift, aliases, module paths, or backend selection.
Campaigns run as reviewable pull requests through the existing infrastructure workflow. Plans, diffs, and CI results provide the evidence. Runbooks live beside the code, and autofix commits carry a marker to prevent the checks from repeatedly triggering themselves.
What changed
Developers gained a consistent way to connect assistants to operational data using identities they already had. Repeated engineering tasks gained versioned instructions and explicit stopping conditions, while production tool restrictions and the established review process continued to govern what could change.