Debugging Production Pods Without SSH or Hot Patching
There are two tempting shortcuts when something breaks in production. One is to get a shell in the container and patch a file. The other is to run production with hot reload, so a code change goes live without a deploy. Both produce running code that is not in Git or in any image, that nobody reviewed, and that disappears on the next restart or differs between replicas.
Hot reload tools like nodemon, air or webpack's HMR are development tools. Production should run immutable images that went through the pipeline. If the pipeline is too slow for an urgent fix, the fix is a faster pipeline, not a side door.
What you do need in production is a way to look inside a running container without changing it. Kubernetes has good tools for that.
Start from the outside
Most problems can be diagnosed without entering the container:
kubectl -n my-app get pods -l app=my-app
kubectl -n my-app describe pod my-app-7d9c8b6f5-abcde
kubectl -n my-app logs my-app-7d9c8b6f5-abcde --previous
kubectl -n my-app get events --sort-by=.lastTimestamp
describe shows restart counts, the last termination reason (OOMKilled, exit code 137, failed probes) and scheduling problems. --previous gives the logs of the container instance that crashed, which is usually the one you care about.
Ephemeral debug containers
Many production images are distroless or minimal and have no shell. kubectl debug adds an ephemeral container with your tools to the running pod. Ephemeral containers are stable since Kubernetes 1.25.
kubectl -n my-app debug -it my-app-7d9c8b6f5-abcde \
--image=nicolaka/netshoot --target=my-app -- bash
--target puts the debug container in the process namespace of the my-app container, so ps shows the application process. Because the pod's network namespace is shared, ss, curl localhost:8080, dig and tcpdump see exactly what the application sees.
Things to know:
- The default profile is
general. If you need more privileges, for example to capture packets,--profile=netadminor--profile=sysadminadd them, as far as your Pod Security settings allow. - An ephemeral container cannot be removed from the pod. It stays in the pod spec until the pod is deleted.
- Creating one needs RBAC on the
pods/ephemeralcontainerssubresource. Grant it deliberately, because it gives access to everything the pod can reach.
Debug a copy, not the live pod
When you need to change something, such as the command, an environment variable or the image, do it on a copy:
kubectl -n my-app debug my-app-7d9c8b6f5-abcde -it \
--copy-to=my-app-debug --container=my-app -- sh
By default the copy does not keep the original labels, so Services do not send it traffic and the ReplicaSet does not manage it. Liveness, readiness and startup probes are removed too, so the copy is not restarted while you work. --set-image=my-app=registry.example.com/my-app:1.2.3-debug swaps in a debug build. Delete the copy when you are done.
For problems below the pod, kubectl debug node/my-node -it --image=busybox starts a pod on the node with the host filesystem mounted at /host.
Change behavior safely at runtime
Some things should be adjustable without a deploy. Make them explicit features of the application instead of ad hoc patches.
Log level. Turning on debug logging for one pod is often all you need. In Go, slog.LevelVar makes the level adjustable:
package main
import (
"log/slog"
"net/http"
"os"
)
func main() {
level := new(slog.LevelVar) // Info by default
slog.SetDefault(slog.New(slog.NewJSONHandler(os.Stdout, &slog.HandlerOptions{Level: level})))
admin := http.NewServeMux()
admin.HandleFunc("PUT /loglevel", func(w http.ResponseWriter, r *http.Request) {
if err := level.UnmarshalText([]byte(r.URL.Query().Get("level"))); err != nil {
http.Error(w, err.Error(), http.StatusBadRequest)
return
}
slog.Info("log level changed", "level", level.Level())
})
go func() { slog.Error("admin server", "err", http.ListenAndServe("127.0.0.1:6060", admin)) }()
app := http.NewServeMux()
app.HandleFunc("GET /", func(w http.ResponseWriter, r *http.Request) {
slog.Debug("request", "path", r.URL.Path)
w.Write([]byte("ok\n"))
})
slog.Error("server stopped", "err", http.ListenAndServe(":8080", app))
}
The admin port listens on localhost only, so it is not reachable through the Service. kubectl port-forward can still reach it:
kubectl -n my-app port-forward pod/my-app-7d9c8b6f5-abcde 6060:6060
curl -X PUT 'http://localhost:6060/loglevel?level=debug'
The change affects one pod and is gone after a restart. For a debug switch, that is what you want. Spring Boot's actuator loggers endpoint does the same for Java applications.
Configuration. ConfigMaps mounted as volumes are updated in the pod after a delay. Mounts that use subPath and environment variables are not updated. The application has to watch the file or reload on a signal. Make the change in Git, so the next deploy does not undo it.
Profiling. Go's net/http/pprof, Java Flight Recorder or a continuous profiler such as Pyroscope or Parca show where CPU and memory go without attaching a debugger to a production process.
Feature flags and kill switches turn behavior on and off without a deploy, through a tool with an audit log.
Where hot reload belongs
Use it in local and development environments: Skaffold's file sync, Tilt's live update or docker compose watch give a fast loop against a dev cluster or local containers. To test local code against real dependencies, tools like mirrord or Telepresence connect a local process to a staging cluster. None of these should point at production.
Checklist
- Images are immutable, fixes go through the pipeline.
- Diagnose with
describe,logs --previousand events first. - Use
kubectl debugephemeral containers instead of shells in the app container. - Use
--copy-tofor anything that changes the pod. pods/ephemeralcontainerspermission is granted deliberately.- Log level, config reload and profiling are built-in, local-only features.
- Hot reload stays in development.
