Production Engineering¶
Getting a workload to run on Kubernetes is the easy part. This section is the discipline that separates a cluster that survives a bad Tuesday from one that doesn't: sizing it correctly, controlling its cost, deciding whether you need more than one, planning for the day it fails outright, and proving — with a checklist, not a feeling — that a workload is actually ready for production traffic.
Read in this order¶
- Cluster Sizing and Capacity Planning — the math behind node sizing, reserved overhead, and etcd's hard limits
- Cost Optimization — right-sizing, spot capacity, autoscalers, and where the money actually goes
- Multi-Cluster and Multi-Region — why teams split clusters, and the patterns that actually get used
- Disaster Recovery — RTO/RPO, backups, failover, and what a real runbook contains
- Production Readiness Checklist — the gate a workload should pass before it takes real traffic
This section assumes the fundamentals
If you haven't yet covered workloads, networking, storage, and security individually, work through those sections first — this one is about combining them under real operational pressure, not teaching them from scratch.
Next¶
Continue to OpenShift to see how a major enterprise Kubernetes distribution packages many of these same production concerns into the platform itself.