01
What platform engineering should solve
Platform engineering is the discipline of designing and operating an internal product that helps developers deliver software safely. The platform is not merely a Kubernetes cluster or a collection of Terraform modules: it is the supported path from a repository to a reliable production service.
A useful platform removes repeated decisions, preserves escape hatches and makes the secure path the easiest path. Its success is measured in developer outcomes and operational quality, not in the number of tools installed.
- Reduce the time and cognitive load required to create, deploy and operate a service.
- Encode security, identity, networking and observability defaults once.
- Give teams clear ownership, feedback and recovery paths when production changes.
02
A five-layer operating model
Treating the platform as a stack of responsibilities keeps tool choices subordinate to the experience being delivered.
-
01
AWS Organizations · IAM · VPC · KMS
Cloud foundation
Accounts, identity, networks, encryption, budgets and audit controls form the boundary every workload inherits.
-
02
Terraform · policy checks · remote state
Provisioning
Versioned Terraform modules expose deliberate inputs and produce environments that can be reviewed and reproduced.
-
03
GitHub Actions · OIDC · ECR
Delivery
A standard pipeline builds once, proves provenance and promotes an immutable artifact instead of rebuilding per environment.
-
04
EKS or K3s · ECS · managed data services
Runtime
The runtime matches workload needs. Kubernetes is valuable when its operational model earns its cost; ECS or managed services may be the better default.
-
05
Golden paths · runbooks · scorecards
Developer interface
Templates, documentation, service metadata and self-service actions turn infrastructure capabilities into a product teams can actually use.
03
An AWS reference path
A small platform can begin with one paved road and grow only where real demand appears.
- 1
Repository
A service starts from a maintained template with ownership, health checks and deployment metadata.
- 2
Identity
GitHub Actions exchanges OIDC claims for a narrowly scoped IAM role; long-lived AWS keys stay out of CI.
- 3
Artifact
The pipeline tests, scans and publishes one immutable image to ECR, tagged with a release and commit SHA.
- 4
Desired state
A separate configuration change records the image version, resource policy and environment-specific values.
- 5
Reconciliation
Argo CD compares Git with the cluster, applies the declared state and exposes drift or failed health checks.
- 6
Operations
Metrics, logs, alerts, backups and a tested rollback path close the loop after deployment.
04
The GitOps delivery loop
GitOps is useful when it creates a legible control plane, not when it simply adds another deployment tool.
- Git records the intended version and the review that approved it.
- The reconciler reports drift instead of allowing silent manual changes to become permanent.
- Rollback means reverting desired state to a known image, while database recovery remains a separate and tested procedure.
- Secrets are referenced from an external store and materialized with explicit ownership and rotation boundaries.
05
Golden path checklist
A golden path should be opinionated enough to be useful and small enough for one team to maintain.
- A repository template with build, test, ownership and dependency-update conventions.
- OIDC-based cloud access with one role per deployment responsibility.
- Reusable infrastructure modules with examples, versioning and policy validation.
- Health, readiness and graceful-shutdown behavior defined before production.
- Default metrics, structured logs, dashboards and actionable alerts.
- Resource requests, limits, disruption behavior and scaling expectations.
- Automated backups plus a restore exercise, not only a successful upload job.
- A rollback runbook that distinguishes application, configuration and data recovery.
06
Signals that matter
Platform adoption is not a vanity metric. Combine delivery, reliability and developer-experience signals.
- Lead time
- Time from an approved change to a healthy production deployment.
- Recovery
- Time to restore service after a failed release or infrastructure incident.
- Adoption
- Percentage of eligible services using the supported path without forced migration.
- Friction
- Manual steps, support requests and repeated exceptions required to ship a normal change.
07
Common failure modes
Most platform problems are product and ownership problems expressed through infrastructure.
Starting with a portal
A polished catalog cannot compensate for unreliable provisioning and unclear operational ownership.
Making Kubernetes the goal
A cluster is one runtime choice. It does not create self-service, standards or a support model by itself.
Hiding every detail
Abstractions should reduce repetition while preserving enough visibility for teams to diagnose production.
Ignoring the migration path
A technically elegant platform with no incremental adoption strategy becomes shelfware.
Measuring tool activity
Pipeline runs and module counts say little about whether delivery is faster or safer.
08
Practical questions
Does every company need a platform team?
No. Start with shared conventions and a clear owner. A dedicated team makes sense when repeated platform work is materially slowing several product teams.
Should the default runtime be Kubernetes?
Only when workload diversity, scheduling needs and organizational capability justify operating it. A good platform can expose simpler managed runtimes alongside Kubernetes.
Where should a small team begin?
Choose one frequent journey, usually deploying a web service, and make it reliable end to end. Add capabilities after observing real friction.
What makes a platform production-ready?
Defined ownership, secure identity, observable delivery, capacity boundaries, backups, restore tests and recovery procedures that work under pressure.