SRE: Debugging & Recovery
Debugging 5XX Errors — The SRE Way
When your application starts returning 5XX errors, there is a systematic investigation order. Jumping straight to application logs often wastes time — start from the outside (load balancer, pod status) and work inward.
The Debugging Funnel
graph TD
classDef blue fill:#3498db,stroke:#2980b9,color:#fff
classDef green fill:#2ecc71,stroke:#27ae60,color:#fff
classDef red fill:#e74c3c,stroke:#c0392b,color:#fff
classDef orange fill:#e67e22,stroke:#d35400,color:#fff
classDef purple fill:#9b59b6,stroke:#8e44ad,color:#fff
classDef teal fill:#1abc9c,stroke:#16a085,color:#fff
classDef dark fill:#2c3e50,stroke:#1a252f,color:#fff
classDef yellow fill:#f39c12,stroke:#d68910,color:#000
classDef k8s fill:#326ce5,stroke:#254ea8,color:#fff
classDef aws fill:#ff9900,stroke:#cc7a00,color:#000
ALERT["🚨 Alert fires:<br/>5XX rate above SLO burn-rate threshold"]:::red --> L1
subgraph LAYER1["Layer 1 — outermost: Load Balancer"]
L1["Step 1: Check load balancer metrics first"]:::blue --> LB_CHECK{"LB healthy?"}
LB_CHECK -->|"ALB 5XX but<br/>target group healthy"| LB_ISSUE["ALB issue: listener rules,<br/>health check config,<br/>or expired SSL cert"]:::orange
end
LB_CHECK -->|"5XX from targets"| L2
subgraph LAYER2["Layer 2 — Pod Status"]
L2["Step 2: kubectl get pods<br/>-n ns -l app=name"]:::blue --> POD_CHECK{"All pods Running?"}
POD_CHECK -->|CrashLoopBackOff| CRASH["App crashing on startup —<br/>kubectl logs pod --previous"]:::red
POD_CHECK -->|OOMKilled| OOM["Memory limit exceeded —<br/>check limits, heap profile"]:::purple
POD_CHECK -->|Pending| PENDING["Scheduling issue —<br/>kubectl describe pod, check Events"]:::yellow
end
POD_CHECK -->|"Running but 5XX"| L3
subgraph LAYER3["Layer 3 — Pod Events"]
L3["Step 3: kubectl describe pod name"]:::blue --> DESC_CHECK{"Events clean?"}
DESC_CHECK -->|"Readiness probe failing"| PROBE["App unhealthy internally —<br/>check startup errors in logs"]:::green
DESC_CHECK -->|"Resource pressure warnings"| RESOURCES["Node under pressure —<br/>kubectl top nodes / kubectl top pods"]:::yellow
end
DESC_CHECK -->|"Clean events"| L4
subgraph LAYER4["Layer 4 — Application Logs"]
L4["Step 4: kubectl logs pod<br/>-f --tail=200"]:::blue --> LOG_CHECK{"Errors in logs?"}
LOG_CHECK -->|"DB connection errors"| DB_ISSUE["Database issue —<br/>check RDS, connection pool exhaustion"]:::teal
LOG_CHECK -->|"Timeout errors to upstream"| UPSTREAM["Upstream dependency degraded —<br/>check circuit breaker metrics"]:::teal
LOG_CHECK -->|"panic / nil deref"| PANIC["Application bug —<br/>capture goroutine dump, fix and redeploy"]:::red
end
LOG_CHECK -->|"Clean logs"| L5
subgraph LAYER5["Layer 5 — Resource Metrics"]
L5["Step 5: kubectl top pods —<br/>check Prometheus/Grafana"]:::blue --> METRICS_CHECK{"Resource exhaustion?"}
METRICS_CHECK -->|"CPU throttled"| CPU_ISSUE["CPU limit too low —<br/>throttled pod causes slow responses and timeouts"]:::yellow
METRICS_CHECK -->|"Memory near limit"| MEM_ISSUE["About to OOM —<br/>increase memory limit or fix leak"]:::purple
end
METRICS_CHECK -->|"Normal resources"| L6
subgraph LAYER6["Layer 6 — innermost: Network"]
L6["Step 6: kubectl exec pod -- curl upstream<br/>and nslookup service"]:::blue --> NET_CHECK{"Connectivity OK?"}
NET_CHECK -->|"DNS fails"| DNS_ISSUE["CoreDNS issue —<br/>kubectl get pods -n kube-system -l k8s-app=kube-dns"]:::dark
NET_CHECK -->|"Connection refused"| NET_POL["NetworkPolicy blocking —<br/>kubectl get networkpolicy -n ns"]:::red
NET_CHECK -->|"Timeouts"| UPSTREAM2["Upstream too slow —<br/>check service latency p99"]:::teal
end
subgraph LEGEND["Legend — what each color means"]
LG_BLUE["Command to run next"]:::blue
LG_RED["Critical failure —<br/>app or policy bug"]:::red
LG_PURPLE["Memory pressure /<br/>OOM territory"]:::purple
LG_YELLOW["Scheduling or<br/>resource warning"]:::yellow
LG_TEAL["Downstream dependency<br/>is the real culprit"]:::teal
LG_DARK["Cluster infra —<br/>DNS / control plane"]:::dark
LG_GREEN["App-level misconfig,<br/>not infra"]:::green
end
kubectl get pods -n ns -l app=name. CrashLoopBackOff means the app is crashing on startup (go straight to kubectl logs --previous). OOMKilled means the memory limit was exceeded. Pending means a scheduling issue. Only if pods are Running but still serving 5XX do you move inward.
kubectl describe pod name. A failing readiness probe means the app is unhealthy internally — check its startup errors. Resource pressure warnings mean the node itself is under strain, not the app.
kubectl logs pod -f --tail=200. DB connection errors point at the database; upstream timeouts point at a degraded dependency; a panic or nil dereference is an application bug that needs a fix and redeploy — not a config change.
kubectl top pods, cross-checked against Prometheus/Grafana. CPU throttling causes slow responses and timeouts without ever triggering an OOM kill — it's a silent cause easy to miss if you only watch for OOMKilled events. Memory near the limit means you're about to OOM.
kubectl exec pod -- curl upstream and nslookup service. DNS failures point at CoreDNS; connection refused often means a NetworkPolicy is blocking the traffic; timeouts mean the upstream itself is slow. By the time you're here, every outer layer has already been ruled out.
kubectl Runbook
# --- Step 1: Pod overview ---
kubectl get pods -n <namespace> -l app=<name> -o wide
# Look for: STATUS (CrashLoopBackOff, OOMKilled, Pending), RESTARTS count, NODE assignment
# --- Step 2: Describe pod — most important first step ---
kubectl describe pod <pod-name> -n <namespace>
# Key sections to scan:
# Events: at the bottom — "Back-off restarting failed container", "OOMKilled", "FailedScheduling"
# Conditions: Ready=False, reason
# Containers → State: Waiting/Running/Terminated, LastState: exit code
# Exit code reference:
# 137 = OOMKilled (128 + signal 9 SIGKILL)
# 1 = Application error / unhandled exception
# 2 = Misuse of shell command
# 143 = SIGTERM (graceful shutdown, 128 + signal 15)
# --- Step 3: Logs ---
kubectl logs <pod-name> -n <namespace> --tail=200
kubectl logs <pod-name> -n <namespace> --previous # logs from last crashed container
kubectl logs <pod-name> -n <namespace> -c <container> # specific container in multi-container pod
kubectl logs -l app=<name> -n <namespace> --tail=50 # logs from ALL pods with this label
# --- Step 4: Resource consumption ---
kubectl top pods -n <namespace> --sort-by=memory
kubectl top nodes
kubectl describe node <node-name> | grep -A5 "Allocated resources"
# --- Step 5: Events (cluster-wide, sorted by time) ---
kubectl get events -n <namespace> --sort-by='.lastTimestamp'
kubectl get events -n <namespace> --field-selector reason=OOMKilling
# --- Step 6: Exec into pod for debugging ---
kubectl exec -it <pod-name> -n <namespace> -- sh
# Inside: curl, wget, nslookup, cat /proc/meminfo, env
# --- Step 7: Check endpoints (is service pointing to healthy pods?) ---
kubectl get endpoints <service-name> -n <namespace>
# If empty or missing IPs → pods not matching service selector, or pods not Ready
# --- Step 8: Port-forward to test pod directly (bypass LB/ingress) ---
kubectl port-forward pod/<pod-name> 8080:8080 -n <namespace>
curl -v localhost:8080/healthz
128 + 9. The kernel's OOM killer sent SIGKILL because the container exceeded its cgroup memory limit. No graceful shutdown ran — the process was killed mid-instruction. This is the one to page on immediately.
kubectl logs --previous first.
128 + 15. A SIGTERM was sent and the process was given a chance to shut down gracefully — a rolling update, a node drain, a manual kubectl delete pod. This is expected, routine termination, not a crash.
A container exits with code 143. Was it OOMKilled?
Common 5XX Root Causes
| Symptom | Likely cause | Fix |
|---|---|---|
CrashLoopBackOff, exit code 1 |
App panics on startup — missing env var, bad config, failed DB migration | Check logs from previous container |
OOMKilled, exit code 137 |
Memory limit too low, or memory leak | Increase limit, profile heap, check for goroutine leaks |
| Readiness probe fails, pod not Ready | App takes too long to start, or /readyz endpoint broken | Tune initialDelaySeconds, fix readiness logic |
| All pods Running but 5XX | Upstream dependency down (DB, cache, external API) | Check dependency health, circuit breaker open? |
| 5XX only from some pods | Node-level issue (disk pressure, kernel bug) | kubectl cordon <node>, drain, investigate node |
CPU throttling (kubectl top shows 100% but no OOM) |
CPU limit too restrictive | Increase CPU limit, or remove limit entirely (requests only) |
Pending pods |
Insufficient cluster capacity, PodAffinity mismatch, PV stuck | Check events: FailedScheduling, check kubectl describe pod |
Prevention: Alert on SLO burn rate (not raw error count) so you're paged before users notice. Set progressDeadlineSeconds on all Deployments — failed rollouts self-report. Add a post-deploy smoke test in CI: kubectl rollout status && curl /healthz. Use preStop: sleep 5 on all pods to prevent connection reset on rolling updates.
Why does this Prevention rule call for alerting on SLO burn rate rather than raw error count?
Debugging 5XX Errors — EKS-Specific
When running on EKS, you have additional AWS-native tooling on top of the kubectl workflow above.
AWS Load Balancer Controller — Target Group Health
The AWS Load Balancer Controller creates ALBs/NLBs in response to Ingress/Service objects. If pods are healthy but ALB returns 503, the target group may not have registered the pods yet (or the health check is misconfigured).
# Find the ALB created for your ingress
kubectl get ingress -n <namespace>
# ANNOTATION: kubernetes.io/ingress.class: alb shows it's managed by LBC
# Get the ALB ARN from the ingress status
kubectl describe ingress <name> -n <namespace>
# Look for: Address: <alb-dns-name>
# Check target group health via AWS CLI
aws elbv2 describe-target-health \
--target-group-arn arn:aws:elasticloadbalancing:us-east-1:123:targetgroup/k8s-xxx/xxx \
--region us-east-1
# Unhealthy targets show: State.Reason = "Target.FailedHealthChecks"
# Common cause: security group on the node/pod doesn't allow
# health check traffic from the ALB security group on the health check port
kubectl get ingress -n namespace. The kubernetes.io/ingress.class: alb annotation confirms this Ingress is managed by the AWS Load Balancer Controller, not a generic ingress-nginx setup.
kubectl describe ingress name -n namespace and read the Address: field — that's the ALB's DNS name, which you'll need to find the matching resource in the AWS console or CLI.
aws elbv2 describe-target-health --target-group-arn .... Unhealthy targets report State.Reason = "Target.FailedHealthChecks" — this confirms the pods themselves aren't the problem, the ALB just can't reach them on the health check path/port.
Target group health shows State.Reason = "Target.FailedHealthChecks", but kubectl get pods shows every pod Running and passing its own readiness probe. What's the most common cause?
CloudWatch Container Insights
When Container Insights is enabled (via aws-node add-on or ADOT), metrics flow to CloudWatch:
# Query pod OOM events via CloudWatch Logs Insights
# Log group: /aws/containerinsights/<cluster>/performance
fields @timestamp, PodName, reason
| filter Type = "Pod" and reason = "OOMKilling"
| sort @timestamp desc
| limit 50
# Application logs from pods (if using Fluent Bit DaemonSet)
# Log group: /aws/containerinsights/<cluster>/application
fields @timestamp, kubernetes.pod_name, log
| filter kubernetes.namespace_name = "production"
| filter log like /ERROR|PANIC|fatal/
| sort @timestamp desc
| limit 100
You want to find every pod OOM event from the last hour across the cluster. Which CloudWatch log group do you query, and why not the application log group?
/aws/containerinsights/<cluster>/performance, filtering on Type = "Pod" and reason = "OOMKilling". The performance log group carries the pod/node-level metrics and lifecycle events shipped by the aws-node/ADOT add-on — OOM kills live there, not in application log lines. The application log group (shipped separately by the Fluent Bit DaemonSet) only has what your app itself printed to stdout/stderr; a kernel-level OOM kill happens below the application, so it never appears there.EKS Control Plane Logs for Debugging
# Scheduler logs — why is my pod Pending?
aws logs filter-log-events \
--log-group-name /aws/eks/my-cluster/cluster \
--log-stream-name-prefix kube-scheduler \
--filter-pattern '"my-pod-name"' \
--region us-east-1
# API Server audit log — who deleted/modified a resource?
aws logs filter-log-events \
--log-group-name /aws/eks/my-cluster/cluster \
--log-stream-name-prefix kube-apiserver-audit \
--filter-pattern '{ $.requestURI = "/apis/apps/v1/namespaces/prod/deployments/my-app" }' \
--region us-east-1
# Authenticator logs — auth failures (403, unauthorized)
aws logs filter-log-events \
--log-group-name /aws/eks/my-cluster/cluster \
--log-stream-name-prefix authenticator \
--filter-pattern '"error"' \
--region us-east-1
X-Ray / AWS Distro for OpenTelemetry (ADOT)
If your app instruments with OpenTelemetry and sends traces to X-Ray via ADOT Collector:
# X-Ray service map shows 5XX at which service hop
aws xray get-service-graph \
--start-time $(date -u -v-1H +%s) \
--end-time $(date -u +%s) \
--region us-east-1
# Get traces with 5XX status
aws xray get-trace-summaries \
--start-time $(date -u -v-1H +%s) \
--end-time $(date -u +%s) \
--filter-expression 'responsetime > 5 AND http.status = 500' \
--region us-east-1
Four AWS-native tools, four different jobs — pick based on what you're trying to answer:
aws-node/ADOT) and application logs (via Fluent Bit) into CloudWatch Logs Insights, queryable with the same syntax shown above.
Pending pods, the API server audit log answers "who changed this resource," and authenticator logs surface 403/unauthorized auth failures — none of this is visible from inside the cluster with kubectl alone.
Prevention: Enable Container Insights from cluster creation, not after an incident. Set ALB deregistration_delay to 30s (default 300s causes slow deployments and lingering 502s). Use IRSA for all pod-level AWS API access — eliminates the 401 Unauthorized class of 5XX. Enable X-Ray tracing before you need it; retrofitting is painful.
A pod calling S3 gets intermittent 401 Unauthorized errors that show up as 5XX to callers. Per this section's Prevention rule, what eliminates this entire class of error?
OOM-Killed Recovery
Why OOM Kill Happens
The Linux kernel's OOM Killer is invoked when a container exceeds its memory limit (cgroup memory limit set from spec.containers[].resources.limits.memory). The kernel sends SIGKILL (signal 9) to the process — not SIGTERM. There is no graceful shutdown. The process is immediately killed.
Kubernetes detects the exit code 137 (128 + 9) and records the reason as OOMKilled in the pod's lastState.
sequenceDiagram
participant APP as Container process
participant CG as cgroup memory limit
participant KERNEL as Linux OOM Killer
participant KUBELET as kubelet
participant API as Kubernetes API
rect rgb(52, 73, 94)
Note over APP,CG: Phase 1 — growth, still recoverable
activate APP
APP->>CG: Memory usage grows past limits.memory
CG->>KERNEL: cgroup limit exceeded, invoke OOM killer
end
rect rgb(192, 57, 43)
Note over KERNEL,APP: Phase 2 — the kill, zero warning
KERNEL->>APP: SIGKILL (signal 9) — no graceful shutdown
deactivate APP
Note over APP: Process dies immediately, no SIGTERM handler runs
end
rect rgb(41, 128, 185)
Note over KUBELET,API: Phase 3 — kubelet reacts and restarts
KUBELET->>APP: Detect container exited with code 137
KUBELET->>API: Record lastState.reason = OOMKilled
activate KUBELET
KUBELET->>KUBELET: Pull image, if not already cached on the node
KUBELET->>APP: Start new container, restartPolicy Always
deactivate KUBELET
Note over KUBELET,API: Zero capacity during image pull + startup —<br/>this is exactly the risk the Prevention<br/>section below is written to close
end
Containers:
app:
Last State: Terminated
Reason: OOMKilled
Exit Code: 137
Started: Sat, 06 Jun 2026 10:00:00
Finished: Sat, 06 Jun 2026 10:15:32
Requests vs Limits for memory:
requests.memory: the amount the scheduler reserves on the node. Used for placement. Guaranteed to the container.limits.memory: the hard ceiling the kernel enforces. Exceeding this = OOMKill.- Best practice: set requests = your p95 steady-state memory, limits = your p99.9 + buffer. Never set limits to 10x requests "just in case" — this causes node over-commitment and cascading OOM kills during memory pressure.
SIGKILL — nothing graceful about it. This is the number that actually determines whether a memory spike becomes an outage.
Someone sets limits.memory to 10x requests.memory "just in case," reasoning that a generous limit gives the app plenty of headroom before ever risking an OOM kill. What does this best-practice note say actually goes wrong?
requests, so the node happily packs in far more pods than it could actually support if they all grew toward their generous limits at once. When several pods spike memory simultaneously, the node runs out of real memory and the kernel starts OOM-killing pods — potentially cascading across several pods on that node, not just the one that grew. The fix is realistic limits (p99.9 + buffer), not maximally generous ones.Diagnosing the Memory Issue
# 1. Confirm OOM kill and see historical container memory usage
kubectl describe pod <pod-name>
# 2. Check current memory usage
kubectl top pods -n <namespace> --containers
# 3. If app is still running (not yet OOM killed), capture heap profile
# (requires pprof endpoint in the app)
kubectl port-forward pod/<pod-name> 6060:6060
go tool pprof http://localhost:6060/debug/pprof/heap
# In pprof: top20, list <func>, web (opens flame graph)
# 4. Check for goroutine leaks (goroutines hold stack memory)
curl http://localhost:6060/debug/pprof/goroutine?debug=2 | head -100
kubectl describe pod pod-name shows lastState.reason = OOMKilled and the historical container memory usage leading up to the kill — your starting evidence that this really was a memory limit, not something else.
kubectl top pods -n namespace --containers shows whether the replacement container (or other pods in the same workload) are trending toward the same limit right now.
/debug/pprof/heap — top20 and list func in the pprof shell point at exactly which allocations are dominating the heap.
/debug/pprof/goroutine?debug=2 dumps every goroutine's stack so you can spot ones that never should have stayed alive.
Memory usage climbs slowly and steadily over days rather than spiking suddenly before an OOM kill. Which of the two profiling steps above points at this, and why?
/debug/pprof/goroutine?debug=2), not the heap profile on its own. Goroutines hold stack memory even while idle, so a leak — goroutines started and never cleaned up — shows up as gradual, steady growth rather than a sudden allocation spike. A one-off heap profile snapshot can miss this pattern entirely if you only look at what's dominating the heap at a single point in time instead of tracking the goroutine count trend.Recovery: Singleton Pod
A singleton is a single-replica deployment — typically a controller, cron job, leader-elected worker, or stateful singleton service.
The risk: When OOM-killed, the pod restarts (kubelet's restartPolicy: Always). During the restart window (image pull + startup time), there is zero capacity serving requests. For a stateful singleton, in-flight operations are lost.
Mitigation strategies:
apiVersion: apps/v1
kind: Deployment
metadata:
name: my-singleton
spec:
replicas: 1 # singleton
template:
spec:
containers:
- name: app
image: my-org/my-app:v1
resources:
requests:
memory: "256Mi" # what the scheduler reserves
cpu: "100m"
limits:
memory: "512Mi" # OOM if exceeded — set realistically
cpu: "500m"
# Give the process time to finish in-flight work before SIGKILL
lifecycle:
preStop:
exec:
command: ["/bin/sh", "-c", "sleep 5"] # drain connections
# Readiness probe prevents traffic during restart
readinessProbe:
httpGet:
path: /readyz
port: 8080
initialDelaySeconds: 5
periodSeconds: 5
failureThreshold: 3
# Give the pod up to 30s to shutdown gracefully after SIGTERM
terminationGracePeriodSeconds: 30
For a true singleton, also consider Vertical Pod Autoscaler (VPA) to automatically right-size memory based on historical usage:
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: my-singleton-vpa
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: my-singleton
updatePolicy:
updateMode: "Auto" # or "Off" to just see recommendations without applying
resourcePolicy:
containerPolicies:
- containerName: app
minAllowed:
memory: "128Mi"
maxAllowed:
memory: "2Gi"
In the singleton mitigation manifest, the pod has both a preStop hook and a readinessProbe. During an OOM kill specifically, which of these two actually gets a chance to run, and why does the other one matter anyway?
preStop nor graceful shutdown logic runs during an OOM kill — the kernel sends SIGKILL directly, with no warning and no lifecycle hook invoked. preStop only helps during voluntary terminations (rolling updates, node drains, scale-downs), not an OOM kill. The readinessProbe is what actually matters here: after the pod restarts, it keeps the pod out of Service endpoints until /readyz passes, so the restart window's zero capacity doesn't turn into requests being routed to a not-yet-ready container.Recovery: Distributed/Replicated Pod
A distributed workload runs replicas: N > 1. When one pod OOM-kills, others continue serving. The key is ensuring:
- Enough replicas so one death doesn't cause capacity collapse
- PodDisruptionBudget so rolling restarts/node drains don't kill too many at once
- Anti-affinity so replicas aren't all on the same node
apiVersion: apps/v1
kind: Deployment
metadata:
name: my-service
spec:
replicas: 3
strategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 1 # at most 1 pod down during update/restart
maxSurge: 1
template:
spec:
# Spread pods across nodes — don't put all replicas on same node
affinity:
podAntiAffinity:
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 100
podAffinityTerm:
labelSelector:
matchLabels:
app: my-service
topologyKey: kubernetes.io/hostname
containers:
- name: app
resources:
requests:
memory: "256Mi"
cpu: "200m"
limits:
memory: "512Mi"
cpu: "1000m"
---
# PodDisruptionBudget — prevent too many simultaneous disruptions
# (node drains, rolling deployments, voluntary disruptions)
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: my-service-pdb
spec:
selector:
matchLabels:
app: my-service
minAvailable: 2 # always keep at least 2 pods running
# OR: maxUnavailable: 1
With PDB in place: When a node is drained (for upgrade, scaling down), the drain will block if removing a pod would violate minAvailable. The node drain waits until a replacement pod is healthy before proceeding. This prevents rolling OOM-kills from cascading into a full outage.
Prevention: Set memory requests == limits (Guaranteed QoS) for critical services — prevents the OOM killer from targeting them during node pressure. Set GOMEMLIMIT in Go services to ~90% of the K8s limit so GC reclaims memory before the kernel kills the process. Use VPA in Off mode first to get right-sizing recommendations before enabling Auto. Alert on container_memory_working_set_bytes / container_spec_memory_limit_bytes > 0.85.
A node is under memory pressure and the kubelet has to evict something. Why does setting requests == limits (Guaranteed QoS) protect a critical pod from being the one chosen?
Debugging Services Without SSH or SSM Access
In production EKS/GKE environments you often have no direct shell access to nodes. This is the full toolkit ordered from least to most invasive.
--previous), endpoint health, and the namespace-wide events timeline. This alone answers most "why is this pod broken" questions without touching a shell at all.
hostPID/hostNetwork and a chroot into the host filesystem. The most invasive kubectl-based option.
kubectl access is unavailable: CloudWatch Logs Insights for application logs, ALB target health for traffic-layer visibility, EKS control plane logs for scheduler/auth/audit questions, and X-Ray for end-to-end trace visibility.
Layer 1 — kubectl (no exec required)
# Pod state and events — always start here
kubectl get pods -n <ns> -l app=<name> -o wide
kubectl describe pod <pod> -n <ns>
# Read the Events section at the bottom first:
# "Back-off restarting failed container" → CrashLoopBackOff with logs available
# "OOMKilling" → memory limit hit
# "Liveness probe failed" → readiness/liveness misconfigured
# "FailedScheduling" → no node can fit the pod
# Logs — current and previous container
kubectl logs <pod> -n <ns> --tail=200 -f
kubectl logs <pod> -n <ns> --previous # last crashed container's logs
kubectl logs -l app=<name> -n <ns> --tail=50 # all pods in selector simultaneously
# Is the Service actually backed by healthy pods?
kubectl get endpoints <svc-name> -n <ns>
# Empty = pods not matching selector labels OR no pods in Ready state
# Events timeline across namespace
kubectl get events -n <ns> --sort-by='.lastTimestamp' | tail -30
Layer 2 — Port-forward to isolate the problem
Port-forward bypasses the entire LB → Ingress → Service → kube-proxy chain. Use it to test the pod in isolation:
# Test pod directly (eliminates LB, Ingress, Service, kube-proxy as suspects)
kubectl port-forward pod/<pod-name> 8080:8080 -n <ns>
curl -v localhost:8080/healthz
# Test via Service (validates kube-proxy rules and endpoint selection)
kubectl port-forward svc/<svc-name> 8080:80 -n <ns>
curl -v localhost:8080/healthz
# Decision tree:
# pod PF works + svc PF works → problem is at Ingress or LB layer
# pod PF works + svc PF fails → kube-proxy or endpoint selector issue
# pod PF fails → problem is in the application itself
flowchart TD
classDef test fill:#3498db,stroke:#2471a3,color:#fff
classDef good fill:#27ae60,stroke:#1e8449,color:#fff
classDef bad fill:#e74c3c,stroke:#c0392b,color:#fff
classDef verdict fill:#8e44ad,stroke:#6c3483,color:#fff
START(["5XX reported,<br/>layer unknown"]):::verdict --> PODPF
subgraph TEST1["Test 1 — bypass everything, hit the pod directly"]
PODPF["kubectl port-forward pod/name"]:::test --> PODRESULT{"Pod responds<br/>directly?"}
end
PODRESULT -->|No| APPBUG["Problem is in the application itself —<br/>Service/Ingress/LB are not the cause"]:::bad
PODRESULT -->|Yes| SVCPF
subgraph TEST2["Test 2 — bring the Service back into the path"]
SVCPF["kubectl port-forward svc/name"]:::test --> SVCRESULT{"Service<br/>responds?"}
end
SVCRESULT -->|No| PROXY["kube-proxy rules or<br/>endpoint selector issue"]:::bad
SVCRESULT -->|Yes| LBISSUE["Both layers work in isolation —<br/>problem is at Ingress or LB layer"]:::good
Pod port-forward responds fine, but Service port-forward hangs. Which two layers does this rule out, and which one is now implicated?
Layer 3 — Ephemeral debug containers (K8s 1.23+)
Inject a debug container into a running pod. It shares the pod's namespaces without modifying the original container or requiring a pod restart:
# Inject busybox into a running pod
kubectl debug -it <pod> -n <ns> \
--image=busybox:latest \
--target=<container-name>
# Inject netshoot (full network tools)
kubectl debug -it <pod> -n <ns> \
--image=nicolaka/netshoot \
--target=<container-name>
# Now you can: curl, tcpdump, ss, nslookup, traceroute, iperf3
# The --target flag shares the target container's process namespace
# so you can see the app's processes and file descriptors
# Ephemeral containers are not restarted and cannot be removed until pod dies
kubectl describe pod <pod> -n <ns> # shows ephemeral containers section
Layer 4 — Temporary debug pod in the same namespace
When you need network tools but the target pod is CrashLoopBackOff (no exec possible):
# Run netshoot as a temporary pod in the problem namespace
kubectl run debug-pod --rm -it \
--image=nicolaka/netshoot \
--restart=Never \
-n <ns> \
-- bash
# Now inside netshoot — same namespace as the broken service:
# DNS resolution
nslookup payments-svc.payments.svc.cluster.local
nslookup payments-svc # short name, relies on search domains
# Connectivity test
curl -v http://payments-svc:8080/healthz
curl -v http://10.96.45.20:8080/healthz # direct ClusterIP (bypasses DNS)
# Port scan (is the app even listening?)
nc -zv payments-svc 8080
# Trace route to pod (shows where packets are dropped)
traceroute payments-svc
# Capture traffic (if you know which pod IP)
tcpdump -i eth0 host <pod-ip> and port 8080
CrashLoopBackOff. One-way door: ephemeral containers are never restarted and can't be removed until the pod itself dies.
CrashLoopBackOff and has nothing running to attach to, because it doesn't depend on the broken pod having a live container. Gives full network tooling (DNS, connectivity, port scan, packet capture) from the same network vantage point as the broken service.
A pod is stuck in CrashLoopBackOff and you need netshoot's tooling to debug DNS and connectivity. Why won't kubectl debug --target (the ephemeral container approach) work here, and what's the alternative?
CrashLoopBackOff pod has no live container to attach to, since it's repeatedly starting and immediately dying. The alternative is Layer 4: run netshoot as its own standalone pod (kubectl run debug-pod --image=nicolaka/netshoot ...) in the same namespace. It's not attached to the broken pod at all, so it doesn't need the broken pod to be running — it just needs to sit in the same namespace to reach the same Services and test DNS/connectivity from a comparable vantage point.Layer 5 — Debug a node problem (via privileged DaemonSet)
When the issue is at the node level (disk pressure, kernel issue, iptables corruption) and you have no SSH:
# Create a privileged pod on a specific node
kubectl run node-debug \
--image=busybox \
--restart=Never \
--rm -it \
--overrides='{
"spec": {
"nodeName": "<node-name>",
"hostPID": true,
"hostNetwork": true,
"containers": [{
"name": "node-debug",
"image": "busybox",
"stdin": true,
"tty": true,
"securityContext": {"privileged": true},
"volumeMounts": [{"name": "host-root","mountPath": "/host"}]
}],
"volumes": [{"name": "host-root","hostPath": {"path": "/"}}]
}
}' -- sh
# Inside: chroot to host filesystem
chroot /host bash
# Now you have full access to the node's filesystem and processes
# Check iptables, ss, top, dmesg, journalctl
iptables -t nat -L KUBE-SERVICES | head -50
ss -tlnp
journalctl -u kubelet --tail=100
This debug pod sets nodeName, hostPID: true, hostNetwork: true, privileged: true, and mounts / from the host. Why do you need all of these together, instead of just execing in with elevated privileges?
nodeName pins the pod to the specific node you actually need to inspect (a regular pod could land anywhere). hostPID and hostNetwork share the node's process and network namespaces so tools like ss and process listings see the real node, not the pod's own isolated namespace. privileged: true grants the syscall capabilities that commands like iptables need. And mounting / as host-root, then chroot-ing into it, is what makes the node's actual binaries, config files, and journalctl logs accessible as if you'd SSH'd in directly. This is also the most invasive tool in the whole toolkit precisely because it grants this much — it's the last resort, not the first thing to reach for.Layer 6 — AWS-specific (CloudWatch, X-Ray)
# Search application logs via CloudWatch Logs Insights
# (assumes Fluent Bit DaemonSet shipping to CloudWatch Container Insights)
aws logs start-query \
--log-group-name /aws/containerinsights/<cluster>/application \
--start-time $(date -u -v-1H +%s) \
--end-time $(date -u +%s) \
--query-string '
fields @timestamp, kubernetes.pod_name, log
| filter kubernetes.namespace_name = "payments"
| filter log like /ERROR|PANIC|fatal/
| sort @timestamp desc
| limit 50
'
# Get query ID from response, then:
aws logs get-query-results --query-id <id>
# ALB target health (why are targets unhealthy?)
aws elbv2 describe-target-health \
--target-group-arn <arn> --region <region>
# "Reason": "Target.FailedHealthChecks" = app not responding on health check port
# "Reason": "Target.DeregistrationInProgress" = pod draining
# EKS control plane logs (scheduler, authenticator, API server)
# Enable first: EKS Console → Cluster → Logging → enable scheduler + api
aws logs filter-log-events \
--log-group-name /aws/eks/<cluster>/cluster \
--log-stream-name-prefix kube-scheduler \
--filter-pattern '"<pod-name>"'
aws logs filter-log-events \
--log-group-name /aws/eks/<cluster>/cluster \
--log-stream-name-prefix authenticator \
--filter-pattern '"Unauthorized"'
# X-Ray — trace 5XX errors end-to-end
aws xray get-trace-summaries \
--start-time $(date -u -v-1H +%s) \
--end-time $(date -u +%s) \
--filter-expression 'http.status = 500' \
--region <region>
Decision tree — which tool to use
flowchart TD
classDef state fill:#2c3e50,stroke:#1a252f,color:#fff
classDef cmd fill:#3498db,stroke:#2471a3,color:#fff
classDef cause fill:#e67e22,stroke:#ba6018,color:#fff
classDef fix fill:#27ae60,stroke:#1e8449,color:#fff
classDef bad fill:#e74c3c,stroke:#c0392b,color:#fff
START{"What state is<br/>the pod in?"}:::state
START -->|CrashLoopBackOff| CRASH1
subgraph CRASHLOOP["Pod is CrashLoopBackOff"]
CRASH1["kubectl logs --previous<br/>(always start here)"]:::cmd --> CRASH_EMPTY{"Logs empty?"}
CRASH_EMPTY -->|Yes| CRASH_DESC["kubectl describe pod —<br/>exit code 137=OOM, 1=app error"]:::cmd
end
START -->|"Running but 5XX"| RUN1
subgraph RUNNING5XX["Pod is Running but 5XX"]
RUN1["kubectl port-forward pod —<br/>does pod respond directly?"]:::cmd --> RUN_DIRECT{"Direct<br/>response?"}
RUN_DIRECT -->|Yes| RUN_SVC["kubectl port-forward svc —<br/>check Service/endpoints"]:::cmd
RUN_DIRECT -->|No| RUN_APP["Application bug —<br/>check kubectl logs -f"]:::bad
RUN2["kubectl get endpoints —<br/>is Service backed by any pod?"]:::cmd --> RUN_EMPTY{"Endpoints<br/>empty?"}
RUN_EMPTY -->|Yes| RUN_LABEL["Label mismatch on selector"]:::cause
RUN3["kubectl exec OR kubectl debug —<br/>test internal connectivity"]:::cmd --> RUN_CURL["curl postgres-svc —<br/>DNS + connectivity in one shot"]:::fix
end
START -->|Pending| PEND1
subgraph PENDING_G["Pod is Pending"]
PEND1["kubectl describe pod —<br/>Events: FailedScheduling + reason"]:::cmd --> PEND_CAUSE{"Reason?"}
PEND_CAUSE -->|"Insufficient<br/>memory/cpu"| PEND_TOP["kubectl top nodes"]:::cause
PEND_CAUSE -->|"No nodes<br/>match affinity"| PEND_AFFINITY["Check nodeSelector/affinity"]:::cause
PEND_CAUSE -->|"PVC<br/>unbound"| PEND_PVC["kubectl describe pvc"]:::cause
end
START -->|"No kubectl access<br/>(pure AWS)"| AWS_G
subgraph AWS_G["No kubectl access (pure AWS)"]
AWS1["CloudWatch Logs Insights →<br/>application logs"]:::fix
AWS2["ALB target health →<br/>is pod receiving traffic?"]:::fix
AWS3["EKS control plane logs →<br/>auth failures, scheduling issues"]:::fix
end
subgraph LEGEND["Legend"]
LG_STATE["Pod state — the branch point"]:::state
LG_CMD["Command to run"]:::cmd
LG_CAUSE["Root cause identified"]:::cause
LG_FIX["Resolves or confirms healthy"]:::fix
LG_BAD["Confirmed broken — needs a code/config fix"]:::bad
end
Useful netshoot commands cheatsheet
# DNS
nslookup svc-name.namespace.svc.cluster.local
dig svc-name.namespace.svc.cluster.local
# Check /etc/resolv.conf for search domains
cat /etc/resolv.conf
# Connectivity
curl -sv http://svc:port/path 2>&1 | head -50
nc -zv svc-name port # TCP reachability without curl
wget -qO- http://svc:port/healthz
# Network state
ss -tlnp # listening ports inside pod
ss -tnp state established # active connections
# Packet capture
tcpdump -i eth0 -nn port 8080 -w /tmp/capture.pcap
tcpdump -i eth0 -nn 'host 10.0.1.5'
# Routing
ip route show
ip addr
# TLS
openssl s_client -connect svc:443 -servername hostname
curl -kv https://svc:443/ # ignore cert errors
From inside netshoot, curl -v http://payments-svc:8080/healthz hangs, but curl -v http://10.96.45.20:8080/healthz (the same Service's ClusterIP) responds instantly. What does that split result isolate?
payments-svc to that IP, which points at CoreDNS or /etc/resolv.conf's search domains rather than at the application or the network.