Platform engineering field guide

Platform engineering on AWS: from infrastructure to a usable product

A practical framework for building cloud foundations that reduce delivery friction without hiding operational reality from engineering teams.

Author
Adrian Magarola
Updated
Reviewed September 4, 2026
Reading time
10 minute read

01

What platform engineering should solve

Platform engineering is the discipline of designing and operating an internal product that helps developers deliver software safely. The platform is not merely a Kubernetes cluster or a collection of Terraform modules: it is the supported path from a repository to a reliable production service.

A useful platform removes repeated decisions, preserves escape hatches and makes the secure path the easiest path. Its success is measured in developer outcomes and operational quality, not in the number of tools installed.

  • Reduce the time and cognitive load required to create, deploy and operate a service.
  • Encode security, identity, networking and observability defaults once.
  • Give teams clear ownership, feedback and recovery paths when production changes.

02

A five-layer operating model

Treating the platform as a stack of responsibilities keeps tool choices subordinate to the experience being delivered.

  1. 01

    Cloud foundation

    Accounts, identity, networks, encryption, budgets and audit controls form the boundary every workload inherits.

    AWS Organizations · IAM · VPC · KMS
  2. 02

    Provisioning

    Versioned Terraform modules expose deliberate inputs and produce environments that can be reviewed and reproduced.

    Terraform · policy checks · remote state
  3. 03

    Delivery

    A standard pipeline builds once, proves provenance and promotes an immutable artifact instead of rebuilding per environment.

    GitHub Actions · OIDC · ECR
  4. 04

    Runtime

    The runtime matches workload needs. Kubernetes is valuable when its operational model earns its cost; ECS or managed services may be the better default.

    EKS or K3s · ECS · managed data services
  5. 05

    Developer interface

    Templates, documentation, service metadata and self-service actions turn infrastructure capabilities into a product teams can actually use.

    Golden paths · runbooks · scorecards

03

An AWS reference path

A small platform can begin with one paved road and grow only where real demand appears.

  1. 1

    Repository

    A service starts from a maintained template with ownership, health checks and deployment metadata.

  2. 2

    Identity

    GitHub Actions exchanges OIDC claims for a narrowly scoped IAM role; long-lived AWS keys stay out of CI.

  3. 3

    Artifact

    The pipeline tests, scans and publishes one immutable image to ECR, tagged with a release and commit SHA.

  4. 4

    Desired state

    A separate configuration change records the image version, resource policy and environment-specific values.

  5. 5

    Reconciliation

    Argo CD compares Git with the cluster, applies the declared state and exposes drift or failed health checks.

  6. 6

    Operations

    Metrics, logs, alerts, backups and a tested rollback path close the loop after deployment.

04

The GitOps delivery loop

GitOps is useful when it creates a legible control plane, not when it simply adds another deployment tool.

01Commit
02Checks
03Immutable image
04Git state
05Argo CD
06Runtime signals
  • Git records the intended version and the review that approved it.
  • The reconciler reports drift instead of allowing silent manual changes to become permanent.
  • Rollback means reverting desired state to a known image, while database recovery remains a separate and tested procedure.
  • Secrets are referenced from an external store and materialized with explicit ownership and rotation boundaries.

05

Golden path checklist

A golden path should be opinionated enough to be useful and small enough for one team to maintain.

  • A repository template with build, test, ownership and dependency-update conventions.
  • OIDC-based cloud access with one role per deployment responsibility.
  • Reusable infrastructure modules with examples, versioning and policy validation.
  • Health, readiness and graceful-shutdown behavior defined before production.
  • Default metrics, structured logs, dashboards and actionable alerts.
  • Resource requests, limits, disruption behavior and scaling expectations.
  • Automated backups plus a restore exercise, not only a successful upload job.
  • A rollback runbook that distinguishes application, configuration and data recovery.

06

Signals that matter

Platform adoption is not a vanity metric. Combine delivery, reliability and developer-experience signals.

Lead time
Time from an approved change to a healthy production deployment.
Recovery
Time to restore service after a failed release or infrastructure incident.
Adoption
Percentage of eligible services using the supported path without forced migration.
Friction
Manual steps, support requests and repeated exceptions required to ship a normal change.

07

Common failure modes

Most platform problems are product and ownership problems expressed through infrastructure.

Starting with a portal

A polished catalog cannot compensate for unreliable provisioning and unclear operational ownership.

Making Kubernetes the goal

A cluster is one runtime choice. It does not create self-service, standards or a support model by itself.

Hiding every detail

Abstractions should reduce repetition while preserving enough visibility for teams to diagnose production.

Ignoring the migration path

A technically elegant platform with no incremental adoption strategy becomes shelfware.

Measuring tool activity

Pipeline runs and module counts say little about whether delivery is faster or safer.

08

Practical questions

Does every company need a platform team?

No. Start with shared conventions and a clear owner. A dedicated team makes sense when repeated platform work is materially slowing several product teams.

Should the default runtime be Kubernetes?

Only when workload diversity, scheduling needs and organizational capability justify operating it. A good platform can expose simpler managed runtimes alongside Kubernetes.

Where should a small team begin?

Choose one frequent journey, usually deploying a web service, and make it reliable end to end. Add capabilities after observing real friction.

What makes a platform production-ready?

Defined ownership, secure identity, observable delivery, capacity boundaries, backups, restore tests and recovery procedures that work under pressure.

Adrian Magarola, Platform Engineer

Adrian Magarola · Platform Engineer

Working through a platform engineering decision?

I can help turn infrastructure sprawl or delivery friction into a smaller, operable platform roadmap.

Explore freelance DevOps and cloud consulting services

Start a conversation
v2.60