Case Study: Debugging a CrashLoopBackOff Incident¶
Scenario¶
At 14:02, a routine deploy of order-service:1.4.0 goes out through the normal pipeline. By 14:05, the on-call engineer gets paged: error rate on order-service has spiked, and kubectl get pods shows most replicas in CrashLoopBackOff. The previous version, 1.3.2, is still running fine on the pods that haven't been replaced yet. There's no obvious signal in the deploy pipeline itself — it reported success because the pods it created did technically start.
Requirements¶
- Identify the actual root cause using only
kubectl— no re-deploying blind and hoping, no guessing from the changelog - Distinguish between the real cause and any misleading secondary signals along the way
- Fix it with the smallest possible change, and confirm the fix before considering the incident resolved
- Come out the other side with a concrete change that would have caught this before it reached production (the fix: the pipeline now renders manifests and runs the new image once with the target environment's ConfigMap in a pre-production namespace, so a missing key fails CI instead of production)
Solution Walkthrough¶
Step 1: Confirm the blast radius¶
NAME READY STATUS RESTARTS AGE
order-service-7d9f8c6b5d-2kx9p 0/1 CrashLoopBackOff 6 4m12s
order-service-7d9f8c6b5d-8mz3q 0/1 CrashLoopBackOff 6 4m10s
order-service-7d9f8c6b5d-vq7lh 0/1 CrashLoopBackOff 5 3m58s
order-service-5b6c9d7f4-nk2wp 1/1 Running 0 3h20m
All three new-ReplicaSet pods (7d9f8c6b5d) are crash-looping; the one old pod (5b6c9d7f4) still on 1.3.2 is fine, waiting on maxUnavailable to let it go. This confirms it's the new release, not infrastructure or a coincidence — every pod running 1.4.0 fails the same way.
Step 2: describe pod — the first real evidence¶
...
Containers:
order-service:
Image: registry.example.com/order-service:1.4.0
State: Waiting
Reason: CrashLoopBackOff
Last State: Terminated
Reason: Error
Exit Code: 2
Started: Mon, 24 Aug 2026 14:03:41 +0000
Finished: Mon, 24 Aug 2026 14:03:42 +0000
Ready: False
Restart Count: 6
Limits:
cpu: 500m
memory: 256Mi
Requests:
cpu: 250m
memory: 128Mi
Events:
Type Reason Age From Message
---- ------ ---- ---- -------
Normal Pulled 4m (x6 over 4m12s) kubelet Successfully pulled image "registry.example.com/order-service:1.4.0"
Normal Created 4m (x6 over 4m12s) kubelet Created container order-service
Normal Started 4m (x6 over 4m12s) kubelet Started container order-service
Warning Unhealthy 4m (x9 over 4m10s) kubelet Readiness probe failed: Get "http://10.244.2.17:8080/ready": dial tcp 10.244.2.17:8080: connect: connection refused
Warning BackOff 2m (x14 over 3m50s) kubelet Back-off restarting failed container
The loudest signal is the stream of Readiness probe failed ... connection refused warnings, and the obvious read is "1.4.0 starts slower, the probe gives up too early, raise initialDelaySeconds." This is the red herring. A readiness failure never restarts a container; it only removes the Pod from Service endpoints. Something else is killing the process, and the probe is simply finding nothing listening because the process has already died.
The evidence that matters is in Last State:
| Field | Value | What it tells you |
|---|---|---|
Reason |
Error |
The process exited by itself. OOMKilled would mean the kernel killed it for exceeding its memory limit |
Exit Code |
2 |
Non-zero, chosen by the application (a Go panic exits with 2). 137 would be SIGKILL (OOM or a liveness kill), 143 SIGTERM |
Started → Finished |
One second | It dies during startup, before it ever serves a request |
An application that exits on its own within a second of starting is almost always refusing its configuration. The logs will say which part.
Step 3: logs --previous — what actually happened before the kill¶
2026-08-24T14:03:41.812Z INFO Starting order-service v1.4.0
2026-08-24T14:03:41.834Z INFO Loading configuration from environment
2026-08-24T14:03:41.836Z FATAL Required config key PAYMENT_GATEWAY_URL is not set (found PAYMENT_GATEWAY_ENDPOINT)
panic: fatal configuration error: PAYMENT_GATEWAY_URL is not set
goroutine 1 [running]:
main.loadConfig(...)
/app/config.go:47 +0x1c5
main.main()
/app/main.go:22 +0x89
This is the actual root cause: 1.4.0 renamed the expected environment variable from PAYMENT_GATEWAY_ENDPOINT to PAYMENT_GATEWAY_URL, but the ConfigMap backing it wasn't updated to match, so the new binary panics immediately on startup with a fatal config error. Nothing about probes or timing — the probe failures in step 2 were a symptom of a process that was already dead.
Step 4: Confirm against the actual ConfigMap¶
apiVersion: v1
kind: ConfigMap
metadata:
name: order-service-config
data:
PAYMENT_GATEWAY_ENDPOINT: "https://payments.internal.example.com"
LOG_LEVEL: "info"
Confirmed: the ConfigMap still has the old key name. The deploy shipped an application change and a required config change as two separate, uncoordinated artifacts — the image got the new key name, the ConfigMap didn't.
Step 5: Fix and roll out¶
kubectl patch configmap order-service-config --type merge -p \
'{"data":{"PAYMENT_GATEWAY_URL":"https://payments.internal.example.com"}}'
The old key is deliberately left in place rather than removed in the same patch, in case anything else still references it — cleaning it up is a separate, lower-urgency follow-up once the incident itself is resolved.
ConfigMap changes don't automatically restart pods that consume them as env vars (the same behavior covered in Secrets Rotation), so the crash-looping pods need an explicit restart to pick up the new key:
Waiting for deployment "order-service" rollout to finish: 1 old replicas are pending termination...
deployment "order-service" successfully rolled out
Verification¶
NAME READY STATUS RESTARTS AGE
order-service-7d9f8c6b5d-4h8kd 1/1 Running 0 45s
order-service-7d9f8c6b5d-9nzp2 1/1 Running 0 40s
order-service-7d9f8c6b5d-fx3lw 1/1 Running 0 38s
RESTARTS: 0 and READY: 1/1 across all pods confirms they're staying up, not just briefly passing a check before crashing again. Tail the logs to confirm no more fatal config panics, and specifically watch for a few minutes rather than declaring victory the instant pods go Running — a CrashLoopBackOff pod does show Running briefly during every restart attempt:
No output after several minutes, alongside a stable restart count, is what actually closes the incident.
What Could Go Wrong¶
- Tuning the probe instead of reading
logs --previous— this incident's biggest trap. RaisinginitialDelaySecondsorfailureThresholdin response to the readiness warnings would not have fixed anything; the process would keep exiting on startup no matter how patient the probe was. ReadLast State(reason and exit code) and the previous container's logs before changing any setting. - Misreading exit codes —
137(SIGKILL, oftenOOMKilled),143(SIGTERM), and small application codes like1or2point to completely different causes. The CrashLoopBackOff troubleshooting page has the full table. kubectl logswithout--previouson a crash-looping pod — the plainkubectl logscommand shows the current container attempt, which for a pod that's already back in its backoff-wait state may show nothing at all, or only a fragment.--previousis what shows the fatal error from the container instance that actually just crashed.- Deploying an application change and its required config change as separate, uncoordinated steps — the actual root cause here. A safer pattern is to have the application accept both the old and new key names for one release (with a deprecation warning on the old one), decoupling the code rollout from the config rollout by a full release cycle.
- Restarting the Deployment before actually fixing the ConfigMap — this "fixes" the immediate symptom of stale pods for about as long as it takes the new pods to hit the same fatal panic again, and burns time that should have gone into finding the real cause.
- Not checking whether the previous ReplicaSet was still around for instant rollback — in this incident, fixing the ConfigMap and restarting was faster and safer than rolling back, since the ConfigMap fix also unblocks
1.4.0's actual intended change. But when the root cause is inside the application code itself rather than an environment mismatch,kubectl rollout undo deployment/order-serviceback to1.3.2is almost always the faster way to stop the bleeding while the real fix is developed properly.
For a reusable, step-by-step method beyond this incident, see Fix Kubernetes CrashLoopBackOff.
Next¶
Return to Case Studies for the full set, or continue to Troubleshooting for a broader diagnostic playbook beyond this one incident.