Read previous logs, last termination details and events together. Distinguish application exits, probe-triggered restarts and memory kills before choosing a fix.
Capture the failing attempt
Set the actual namespace, Pod and container names below. A Pod can contain several containers, so specify -c. Previous logs help when the current attempt has not yet reached the failure. They may be unavailable if no previous instance is retained.
Record Last State, Reason, Exit Code, restart count and event timestamps. Check that the Pod belongs to the expected Deployment and release. Avoid deleting it first: a replacement can erase the evidence you need. The backoff delay can vary with cluster version and kubelet configuration; do not diagnose from an assumed universal timer.
NS=demo
POD=api-7d8c9f-example
CONTAINER=api
kubectl get pod "$POD" -n "$NS" -o wide
kubectl describe pod "$POD" -n "$NS"
kubectl logs "$POD" -n "$NS" -c "$CONTAINER" --previous --tail=100
kubectl logs "$POD" -n "$NS" -c "$CONTAINER" --tail=100
kubectl get events -n "$NS" --field-selector "involvedObject.name=$POD" --sort-by=.metadata.creationTimestampChoose a cause from evidence
Imagine a new API release repeatedly logs booting, then disappears without its own exception. That alone does not establish an application crash. Compare its termination with kubelet events. A failed liveness probe can kill an otherwise progressing process; a missing required setting can make the process exit itself.
Exit code 137 means a SIGKILL-related exit convention, not proof of a memory-limit breach. Look for OOMKilled and supporting resource evidence. ImagePullBackOff and an unscheduled Pending Pod need different investigations because the application may not have started at all.
| Evidence | Investigate |
|---|---|
| Application exception | Config, command, dependency |
| Unhealthy + Killing events | Probe endpoint and timing |
| OOMKilled | Memory use and limit |
| Exit 0, repeated restart | Wrong workload or command |
Worked case: startup takes 45 seconds
Our example API loads an index for about 45 seconds before exposing health endpoints. Its liveness probe starts immediately, checks every 10 seconds, and restarts it after three failures. Events show probe failures followed by Killing while logs repeatedly stop during loading. This explains why it never finishes startup.
Use this fragment under the API container, after implementing the three endpoints. /startedz succeeds after initialization; /livez checks whether this process can keep working; /readyz checks whether it can serve requests. The startup settings provide roughly a 90-second failure budget, chosen for this example rather than as a universal recommendation.
startupProbe:
httpGet:
path: /startedz
port: 8080
periodSeconds: 5
timeoutSeconds: 2
failureThreshold: 18
livenessProbe:
httpGet:
path: /livez
port: 8080
periodSeconds: 10
timeoutSeconds: 2
failureThreshold: 3
readinessProbe:
httpGet:
path: /readyz
port: 8080
periodSeconds: 5
timeoutSeconds: 2
failureThreshold: 2Keep startup, liveness and readiness distinct
A configured startup probe holds back liveness and readiness probing until it succeeds. Repeated startup or liveness failures can trigger a restart; readiness failure removes eligibility for normal Service traffic without restarting the container. Making readiness more lenient cannot repair a liveness-driven loop.
Do not make a liveness endpoint fail whenever the database is briefly unavailable. That can restart every API replica during a database incident and slow recovery. Express temporary inability to serve through readiness where appropriate, and keep liveness focused on conditions a restart can improve. Probe checks should be cheap and have clear timeouts.
Fix exits and OOM with their own evidence
For an immediate nonzero exit, reproduce the image's command and inspect required environment names, mounted paths, permissions and dependency errors. Compare the deployed specification with the last working revision. A long-running Deployment whose command finishes successfully will also restart under its normal Always policy; a one-off task usually belongs in a Job.
For OOMKilled, examine memory peaks, concurrency and the configured limit. Live metrics can miss the short spike that killed the process. A larger limit may be a temporary mitigation when node capacity permits, but investigate a leak or unexpectedly large input. A memory request affects scheduling and does not replace the limit.
Prove recovery across the rollout
Deploy the correction through your usual process. Watch all replacement Pods, readiness, restarts and events. Check the request path beyond the previous failure window. A Running phase alone does not prove readiness or stability.
For our case, expect completed initialization, successful startup probing and a stable restart count. Save the cause and verification with the incident. If the release remains unhealthy, consider rollback under your deployment policy while preserving logs.
Quick answers
Frequently asked questions
Is CrashLoopBackOff a Pod phase?
No. It describes restart backoff for a failing container and often appears in kubectl's status display. Inspect container state and Pod events to find the cause.
Why are current logs empty?
The new container attempt may not have logged yet. Request --previous for the named container to inspect its retained previous attempt; log retention and Pod replacement can limit what remains.
Can a readiness probe restart a container?
Readiness failure alone does not restart it. Startup and liveness failures can do so after their configured thresholds. Check the events before attributing a restart to a probe.
Does deleting the Pod fix the loop?
A controller usually creates another Pod with the same failing configuration. Deletion can remove useful evidence. Fix or roll back the underlying workload and verify the replacement.
Source notes
References and review policy
Information checked on October 4, 2026. Section links identify sources for factual claims and technical explanations. Interpretations, practice scenarios and preparation recommendations are RecallDeck’s editorial work.
From reading to recall
Practice the full interview loop.
RecallDeck schedules the concepts you miss and keeps coding, design, and behavioral fundamentals available when the interviewer changes direction.