Go Gotcha #6: Your Go Service Drops Requests Every Deployment
Meta: "Go services drop requests on every deploy without proper SIGTERM handling. Here's the golang graceful shutdown Kubernetes pattern that actually works."
E
very deployment is a small outage. Most teams accept this as normal. It isn't.When Kubernetes terminates a pod during a rolling update, a scale-down, or a node drain it sends a signal to your process and expects it to clean up. If your Go service doesn't handle that signal correctly, requests in flight at the moment of termination get dropped. Connections reset. Clients see errors. For a service handling payments, order submissions, or any write operation, that's not a minor inconvenience.
The golang graceful shutdown Kubernetes pattern that fixes this is about 30 lines of code. But there's a race condition in how Kubernetes terminates pods that most implementations miss and it's the reason services still drop requests even after teams think they've fixed it.
What Kubernetes actually does when it terminates a pod
The termination sequence matters. When Kubernetes decides to kill a pod, two things happen concurrently:
- The pod is removed from the Service endpoints list
- SIGTERM is sent to the container
Kubernetes doesn't do these in sequence. Both happen at roughly the same time. The endpoint removal then has to propagate through kube-proxy or your CNI plugin, updating iptables or ipvs rules on every node. That propagation takes time typically 2 to 10 seconds depending on your cluster size and configuration.
So: your process receives SIGTERM. It starts shutting down. But traffic is still routing to it for several more seconds while the endpoint removal propagates through the cluster.
If your shutdown handler stops accepting connections immediately on SIGTERM, those requests arriving during the propagation window get connection-reset errors. You've handled SIGTERM correctly and still dropped requests.
The naive implementation that looks right
Most teams reach for something like this first:
Untitled1quit := make(chan os.Signal, 1)2signal.Notify(quit, syscall.SIGTERM)3<-quit4os.Exit(0)
That's a hard stop. Any request in flight at the moment of SIGTERM is killed. No drain, no clean close. Kubernetes's terminationGracePeriodSeconds (default: 30 seconds) means nothing to your process you've already exited.
The next iteration adds http.Server.Shutdown:
Untitled1<-quit2ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)3defer cancel()4srv.Shutdown(ctx)
This is closer. Shutdown stops the listener, waits for in-flight requests to complete, then returns. But the endpoint propagation race is still there. Traffic arrives after SIGTERM. Shutdown has already closed the listener. Those connections get refused.
The preStop hook: the missing piece
The fix for the propagation race is a preStop lifecycle hook in your pod spec. This hook runs before Kubernetes sends SIGTERM giving you a window to let the endpoint removal propagate before your process starts shutting down.
Untitled1spec:2 containers:3 - name: api4 lifecycle:5 preStop:6 exec:7 command: ["/bin/sh", "-c", "sleep 10"]
With this in place, the sequence becomes:
- Kubernetes decides to terminate the pod
preStophook runs sleeps 10 seconds- During those 10 seconds, endpoint removal propagates through the cluster
- After
preStopcompletes, Kubernetes sends SIGTERM - Your process receives SIGTERM with no new traffic arriving
- Shutdown drains in-flight requests cleanly
Ten seconds is a conservative estimate for propagation. Five seconds works in most clusters. Use 10 if you're on a large cluster or you've seen propagation lag in your metrics.
One important note: preStop time counts against terminationGracePeriodSeconds. If your grace period is 30 seconds and your preStop sleeps for 10, your process has 20 seconds to drain. Size your shutdown timeout accordingly.
The complete pattern
Here's a production-ready graceful shutdown implementation for a Go HTTP service on Kubernetes:
Untitled1package main23import (4 "context"5 "log"6 "net/http"7 "os"8 "os/signal"9 "syscall"10 "time"11)1213func main() {14 mux := http.NewServeMux()15 mux.HandleFunc("/health", func(w http.ResponseWriter, r *http.Request) {16 w.WriteHeader(http.StatusOK)17 })18 // ... register your handlers1920 srv := &http.Server{21 Addr: ":8080",22 Handler: mux,23 }2425 // Start server in background26 go func() {27 log.Println("server starting on :8080")28 if err := srv.ListenAndServe(); err != nil && err != http.ErrServerClosed {29 log.Fatalf("server error: %v", err)30 }31 }()3233 // Block until signal34 quit := make(chan os.Signal, 1)35 signal.Notify(quit, syscall.SIGTERM, syscall.SIGINT)36 sig := <-quit37 log.Printf("received signal %s, shutting down", sig)3839 // Drain with timeout less than terminationGracePeriodSeconds40 ctx, cancel := context.WithTimeout(context.Background(), 20*time.Second)41 defer cancel()4243 if err := srv.Shutdown(ctx); err != nil {44 log.Fatalf("shutdown error: %v", err)45 }4647 log.Println("server stopped cleanly")48}
Pair this with the preStop sleep in your pod spec and a terminationGracePeriodSeconds of at least 35 seconds (10 for preStop + 20 for drain + 5 buffer before SIGKILL):
Untitled1terminationGracePeriodSeconds: 352containers:3 - name: api4 lifecycle:5 preStop:6 exec:7 command: ["/bin/sh", "-c", "sleep 10"]
The Part Most People Get Wrong
There are two common mistakes after teams implement the pattern above.
The first: setting terminationGracePeriodSeconds too low. If you set it to 15 seconds and your preStop sleeps for 10, your service has 5 seconds to drain. A request that takes 6 seconds a long database query, a downstream API call gets killed by SIGKILL. Set your grace period to comfortably exceed preStop duration + expected max request duration + a few seconds buffer.
The second: not closing downstream connections. srv.Shutdown stops the HTTP listener and drains HTTP connections. It doesn't close your database pool, your Redis client, or your message queue consumer. If you have long-running background goroutines Kafka consumers, scheduled jobs, cache warmers those need their own shutdown logic tied into the same signal handler.
Untitled1sig := <-quit2log.Printf("received signal %s", sig)34// Stop HTTP first5ctx, cancel := context.WithTimeout(context.Background(), 20*time.Second)6defer cancel()7srv.Shutdown(ctx)89// Then close downstream10dbPool.Close()11redisClient.Close()12kafkaConsumer.Close()
Order matters. Stop accepting new work first, then clean up the resources that existing work depends on.
The rewriting a running system problem is fundamentally the same as this one: the tricky part isn't the shutdown itself, it's the transition window where old and new coexist.
Real-World Example
A Gojek-adjacent payments middleware team had a Go service that processed transaction callbacks from payment gateways. Every deployment caused a handful of callback errors payment gateways received connection resets, retried, and sometimes double-processed transactions. The service was deployed 8-12 times per day. The error rate was low enough to accept as "deployment noise" for months.
When they instrumented the shutdown sequence they found: no SIGTERM handler at all. The Go process received SIGTERM, did nothing with it, and Kubernetes killed it with SIGKILL after 30 seconds of inactivity. Those 30 seconds of in-flight callbacks dropped.
They added the full pattern: SIGTERM handler, srv.Shutdown with a 25-second timeout, preStop sleep of 5 seconds, and bumped terminationGracePeriodSeconds to 35.
Deployment errors went to zero. The fix was one afternoon of work. The cost of not fixing it duplicate transaction risk, gateway retry traffic, reconciliation overhead had been real but invisible against the noise of normal operations.
FAQ
Q: Should I handle both SIGTERM and SIGINT? A: Yes. SIGTERM is what Kubernetes sends. SIGINT is what you get from Ctrl+C in local development. Handling both means your local shutdown behavior mirrors production, which makes it easier to test that the shutdown path actually works.
Q: What's the right terminationGracePeriodSeconds for my service?
A: Start with: preStop sleep time + your p99 request duration + 10 seconds buffer. For most API services this lands between 30 and 60 seconds. If your p99 is 500ms, 35 seconds is more than enough. If you're running long batch operations through HTTP, size it to cover those. The default of 30 seconds is too short once you add a preStop hook.
Q: Does http.Server.Shutdown wait for WebSocket connections too?
A: No. Shutdown only waits for HTTP/1.1 and HTTP/2 connections that are in the middle of a request. Long-lived connections like WebSockets need explicit handling listen for the context cancellation inside your WebSocket handler and close the connection cleanly when it fires.
Q: How do I test graceful shutdown locally?
A: Run your service, send a SIGINT with Ctrl+C mid-request, and verify the in-flight request completes before the process exits. With curl, you can test with a slow endpoint: curl localhost:8080/slow & then immediately kill -TERM <pid>. The curl should complete. If it gets a connection reset, your handler isn't working.
Q: Do I need a preStop hook if I'm using an Ingress controller or a service mesh like Istio? A: Yes. The propagation delay exists at the iptables/ipvs layer regardless of what sits in front of your service. Istio and similar proxies add their own connection draining, but they don't eliminate the endpoint propagation race for direct pod-to-pod traffic. Keep the preStop hook.
Zero-downtime deployments aren't free. They require you to think about shutdown as carefully as startup. The pattern is straightforward once you understand the race and once it's in place, deployments stop being something you schedule around peak traffic.
If your team treats request drops on deploy as normal, they're paying a small continuous cost that compounds. SpectreDev builds these patterns in from the start, not as a retrofit.
Internal links used:
- rewriting a running system the same "transition window" problem at a larger scale
External links used:
- None needed standard library and Kubernetes spec are self-contained references
Word count: ~1,510