Bulkhead Pattern: How to Stop One Service from Killing Everything
Meta: "Bulkhead pattern isolates failures in microservices so one slow dependency can't drain your entire system. Real Go implementation and when it actually matters. (~156 chars)"
S
hips don't sink the way most people imagine. It's not one catastrophic breach that floods everything at once. It's a single compartment filling with water and if the hull is divided into isolated sections, that's the only compartment that floods. The rest of the ship keeps sailing.Naval engineers figured this out centuries ago. They called those isolated sections bulkheads.
Gojek's infrastructure team figured out the software equivalent around 2017, the hard way: a slow internal service was silently consuming every available connection in a shared thread pool. Not crashing. Not erroring. Just... holding connections open, waiting for responses that were taking too long. Meanwhile, the rest of the system fast services, healthy dependencies, unrelated features couldn't get a connection at all. The shared pool was full.
One slow thing had rendered everything else unavailable. That's the failure mode the bulkhead pattern exists to prevent.
The Shared Pool Problem
Most backend services share resources. Thread pools. Connection pools. Worker queues. HTTP client instances. Sharing is efficient you're not spinning up dedicated infrastructure for every dependency you call.
The problem is what happens when one dependency goes slow.
Say your Go service calls three downstream APIs: a payment processor, a KYC verification service, and a notification service. You've allocated a pool of 100 goroutines for handling outbound requests. Under normal conditions, all three services are fast. Your 100 goroutines cycle through work quickly. Everything's fine.
Now the KYC service has a bad hour. Response times climb from 200ms to 8 seconds. Your goroutines calling the KYC service stop returning they're blocked waiting on a slow response. Within minutes, all 100 goroutines are occupied waiting on KYC. Payment processing requests come in. Notification calls come in. There's nothing left to handle them. Your payment flow which has nothing to do with KYC is now broken because it's sharing resources with something that's struggling.
This is resource exhaustion through shared pools. It's more common than cascade failure from hard crashes, and it's sneakier because everything looks "up" in your health checks while your users are getting timeouts.
What the Bulkhead Pattern Does
The fix is isolation. You divide your resources goroutines, connections, workers into separate pools, one per dependency or per criticality tier. If KYC's pool fills up, it fills up alone. The payment processor's pool is unaffected. Your notification service's pool keeps moving.
In Go, this translates to bounded goroutine pools with a semaphore pattern. The simplest version:
Untitled1type BulkheadClient struct {2 semaphore chan struct{}3 client *http.Client4 name string5}67func NewBulkheadClient(name string, maxConcurrent int, timeout time.Duration) *BulkheadClient {8 return &BulkheadClient{9 semaphore: make(chan struct{}, maxConcurrent),10 client: &http.Client{Timeout: timeout},11 name: name,12 }13}1415func (b *BulkheadClient) Do(ctx context.Context, req *http.Request) (*http.Response, error) {16 // Try to acquire a slot17 select {18 case b.semaphore <- struct{}{}:19 defer func() { <-b.semaphore }()20 case <-ctx.Done():21 return nil, fmt.Errorf("bulkhead %s: context cancelled waiting for slot", b.name)22 default:23 // Pool is full fail fast instead of waiting24 return nil, fmt.Errorf("bulkhead %s: at capacity (%d/%d)",25 b.name, len(b.semaphore), cap(b.semaphore))26 }2728 return b.client.Do(req.WithContext(ctx))29}
Now you create a separate BulkheadClient per dependency:
Untitled1var (2 paymentClient = NewBulkheadClient("payment", 20, 3*time.Second)3 kycClient = NewBulkheadClient("kyc", 10, 10*time.Second)4 notifyClient = NewBulkheadClient("notification", 30, 2*time.Second)5)
The KYC client gets 10 concurrent slots with a generous 10-second timeout, because KYC calls are legitimately slow and that's acceptable. Payment gets 20 slots with a strict 3-second timeout, because payment latency directly affects user-visible response times. Notification gets 30 slots with the tightest timeout, because notifications are high-volume and you'd rather drop one than block.
When KYC's 10 slots are full because it's having a bad hour incoming KYC calls get a fast rejection instead of queuing up and draining shared capacity. Payment and notification keep running completely unaffected.
That's the bulkhead. The water stays in its compartment.
Sizing the Bulkheads: The Math Most Teams Skip
The thing nobody tells you is that picking the wrong pool size is almost as bad as having no bulkhead at all. Too small and you're rejecting valid requests unnecessarily. Too large and a slow dependency can still exhaust the pool just more slowly.
The formula that actually works starts from your SLO, not from guessing:
pool size = (requests per second) × (expected p99 latency in seconds)
That's Little's Law. If your KYC service handles 5 calls per second and your p99 latency is 4 seconds, you need a minimum pool of 20 slots to serve that load without rejections. Add 20-30% headroom for spikes, and you're at 25-26. Round up to 30.
Now the safety ceiling: this number should never exceed what the downstream service can actually handle. If your KYC provider's API is documented to support 50 concurrent connections before they start throttling you, don't set your pool to 80 and assume it'll be fine.
In practice, most teams start with a rough estimate, deploy with metrics on pool utilization, and tune from there. The metric you want is pool saturation the percentage of time the semaphore is at capacity. If it's above 10-15% under normal load, your pool is too small. If it's never above 1%, it might be too generous (though being too generous is a much better problem to have).
The Part Most People Get Wrong
The bulkhead pattern is often described as just "limit concurrency per dependency." That's true but incomplete, and the incomplete version leads to a subtle mistake I see repeatedly.
Teams implement per-dependency pools, feel good about it, and skip the failure path. What happens when the pool is full? The code above fails fast it returns an error immediately. But the calling code often isn't ready for that. The service receives a "bulkhead at capacity" error and doesn't know what to do with it, so it bubbles it up as a 500. The user sees an error for a request that had nothing to do with KYC being slow.
The bulkhead needs a deliberate fallback, not just an error.
For non-critical paths marketing notifications, recommendation engines, activity feeds the fallback is usually "skip and move on." Log it, increment a counter, return a degraded response. The user doesn't need to know.
For critical paths payments, order confirmations, auth the fallback is harder. You can't silently skip a payment. Here, the bulkhead should pair with a circuit breaker that detects the pool is consistently saturated and starts rejecting at the circuit level while surfacing a proper error to the user: "payment processing is temporarily unavailable, please try again."
This is why circuit breaker in Go and the bulkhead pattern are almost always deployed together in serious production systems. They solve adjacent problems. The circuit breaker detects downstream failure. The bulkhead contains resource exhaustion. You want both.
How Gojek Stopped One Service from Killing Everything
Around 2017-2018, as Gojek was scaling from a ride-hailing app into a super app, their engineering team hit the shared pool problem in a particularly painful way. GoFood the food delivery product was growing fast. The order service was calling a merchant validation API that, under certain conditions during peak lunch hour, became extremely slow.
The merchant validation service wasn't down. It was just slow. Slow enough that connections to it piled up in the shared HTTP client pool that the order service used for all its downstream calls. Driver assignment calls, pricing calls, fraud checks none of which had anything to do with merchant validation were timing out because there were no connections left.
At lunch. On weekdays. Consistently.
The fix wasn't to make merchant validation faster (they did that too, eventually). The immediate fix was isolation a dedicated connection pool for merchant validation, sized to handle the worst-case concurrency with a timeout that matched the service's actual behavior. When merchant validation slowed down during peak hours, it filled its own pool and started rejecting new calls. The fallback for a rejected merchant validation call was to accept the order provisionally and validate asynchronously a tradeoff they'd consciously designed for.
The order flow stopped breaking. Lunch rush became survivable. The merchant validation service was still slow, but it was now slow in its own lane.
This is the thing about the bulkhead pattern: it doesn't fix the slow service. It contains the damage. You still have a slow service to fix. But you've bought time your system stays up while you fix it, instead of having a total outage that forces you to fix it right now under pressure.
For teams building with the saga pattern for distributed transactions, bulkheads matter at each step. Every saga step that calls an external dependency should have its own pool. If one step's dependency is slow, you want it to fail fast and trigger compensation, not sit blocking resources while half your saga steps are stuck waiting.
FAQ
Q: Is the bulkhead pattern only relevant for microservices? A: No it's useful anywhere you're sharing resources across multiple callers with different criticality levels. Even in a monolith, if you have a single database connection pool, a slow background job consuming connections can starve your user-facing request handlers. Separate pools for background work vs. user-facing work is a bulkhead, even in a monolith.
Q: How is the bulkhead pattern different from rate limiting? A: Rate limiting controls how many requests a caller can send per time window it protects the downstream service from overload. Bulkheads control how many concurrent resources a caller allocates for a particular dependency they protect the caller from being drained by a slow downstream. They operate in opposite directions and solve different problems. You often want both.
Q: What should I do when the bulkhead rejects a call? A: It depends on the call's criticality. For non-critical paths, return a degraded response or skip the call entirely. For critical paths, return a clear error to the user (not a 500 something meaningful like "temporarily unavailable") and alert on-call. Never silently drop a critical operation without logging it somewhere with enough context to reconstruct what happened.
Q: Does Go's standard HTTP client share connections across all calls by default?
A: Yes. The default http.DefaultTransport is a shared global. If you're using it for all your outbound calls which many teams do by default you have an implicit shared pool and no bulkhead at all. Create separate http.Client instances per dependency with explicitly configured Transport.MaxIdleConnsPerHost and Transport.MaxConnsPerHost. That's your starting point.
Q: When is the bulkhead pattern overkill? A: When you have one or two downstream dependencies and the cost of a single dependency failing is acceptable. Early-stage services calling one database and one third-party API don't need elaborate pool isolation a timeout and a circuit breaker are enough. Bulkheads pay for themselves when you have many dependencies with different latency profiles and criticality levels, and when the cost of cross-contamination one slow service killing an unrelated feature is genuinely unacceptable.
One slow service should never have the power to take down everything else. The bulkhead pattern is how you take that power away.
If you're building or scaling a Go-based service platform in Indonesia and want this kind of resilience designed in from the start, SpectreDev's team has shipped exactly this architecture across fintech, logistics, and super-app products in the Jakarta market.
Internal links used:
- circuit breaker in Go natural pair; breaker detects failure, bulkhead contains resource exhaustion
- the saga pattern for distributed transactions bulkhead per saga step
External links used:
- None
Word count: ~1,820