Spectre
// PUBLISHED03.10.26
// TIME7 MINS
// TAGS
#GOLANG#KUBERNETES#PERFORMANCE#CONTAINERS
// AUTHOR
Spectre Command

Go Gotcha #5: GOMAXPROCS and the Kubernetes CPU Trap

Meta: "Go sets GOMAXPROCS from node CPU count, not your pod limit. This GOMAXPROCS Kubernetes CPU trap causes latency spikes. Fix it in two lines of code."


D

eploy a Go service to Kubernetes and it might be quietly running wrong. Not crashing. Not throwing errors. Just running with far more OS threads than it should, burning CPU time on context switching instead of actual work.

The golang GOMAXPROCS Kubernetes CPU mismatch is one of those problems that's invisible until you're debugging latency at scale. Your service looks healthy. p50 latency is fine. Then load increases, p99 goes wide, and nothing in your code explains why.

Here's the mechanism and the two-line fix.


What GOMAXPROCS actually does

GOMAXPROCS controls how many OS threads Go uses to run goroutines simultaneously. Set it to 4, and Go's scheduler uses 4 threads.

By default, Go sets GOMAXPROCS to runtime.NumCPU() at startup the number of logical CPUs available on the machine.

That's the right behavior on a bare metal server or a VM you own fully. If the machine has 8 cores, use 8 threads.

Untitled
1fmt.Println(runtime.GOMAXPROCS(0)) // 0 = query current value, don't change it

Run that on your local machine with 8 cores: prints 8. Makes sense.

Run it in a Kubernetes pod on a 64-core node with a CPU limit of 2: also prints 64.


What Kubernetes CPU limits actually do

Kubernetes CPU limits don't remove CPUs from the machine. They throttle CPU time using Linux cgroup quotas.

Set resources.limits.cpu: "2" on a pod, and Kubernetes tells the kernel: this process can use at most 2 CPU-seconds per second. If it tries to consume more, the kernel throttles it.

The physical CPUs on the node are still there. The Go runtime can see all of them. runtime.NumCPU() returns the node's CPU count because that's what the Linux kernel reports at the hardware level. The cgroup quota is enforced separately, at a different layer.

So Go starts 64 OS threads on a 64-core node. Your pod's cgroup limit is 2 CPU-seconds per second. Sixty-four threads compete for the equivalent of 2 cores worth of time. The kernel throttles them aggressively. Context switching overhead spikes. Goroutines that should run in microseconds wait in the scheduler queue.


Why the symptoms are misleading

CPU throttling in Kubernetes shows up as throttled_time in cgroup metrics but most teams don't have that wired into their dashboards. What they see instead is:

  • p99 and p999 latency that's disproportionately high compared to p50
  • Request queue lengths growing under load
  • CPU usage that looks low (because the process is throttled, not because it's idle)
  • GC pauses that seem longer than they should be

The CPU usage reading is the one that throws people. Your pod's CPU usage shows 1.8 cores out of a 2-core limit. Looks almost maxed out. But the underlying problem is thread contention many threads fighting for a small CPU budget not actual computation overload.

This is why teams add more replicas and see diminishing returns. More pods, each with 64 threads competing for 2 cores. The per-pod overhead multiplies.


The Part Most People Get Wrong

The usual response when someone first hears about this: "My service is I/O bound, not CPU bound, so GOMAXPROCS doesn't matter."

It's not quite right.

GOMAXPROCS affects goroutine scheduling even for I/O-heavy work. When goroutines block on I/O, Go parks them and runs something else. The more OS threads you have, the more the kernel has to schedule. Excessive threads mean more preemption events, more cache misses, more time in the scheduler rather than in your code.

For network services (simultaneously I/O bound on upstream calls and database queries, and CPU active on serialising responses and running business logic), thread count matters. The sweet spot is roughly matching your CPU allowance.

There's also a GC interaction. Go's garbage collector uses goroutines that run on the OS threads controlled by GOMAXPROCS. Too many threads, and GC work gets interleaved across more cores than your quota allows, which can cause GC to take longer wall-clock time than expected even if CPU time is the same.

You can verify this by checking observability metrics before and after the fix specifically goroutine count, GC pause duration, and scheduler latency from GODEBUG=schedtrace=1000.


The fix: two lines

Uber open-sourced a package called automaxprocs that reads the Linux cgroup CPU quota and sets GOMAXPROCS to match. Import it with a blank identifier so it runs its init() function:

Untitled
1import (
2 _ "go.uber.org/automaxprocs"
3)

That's it. On a pod with a 2-core CPU limit, GOMAXPROCS becomes 2. On bare metal with 64 cores and no limit, it stays 64. The package handles both cases and falls back to runtime.NumCPU() when cgroup information isn't available.

You can also set it manually:

Untitled
1runtime.GOMAXPROCS(2)

But that hardcodes the value, which breaks the moment you change the pod's resource limits. Automaxprocs reads the limit at startup, so it stays in sync automatically.

Add it to main.go alongside your other initialisation imports:

Untitled
1package main
2
3import (
4 "log"
5 "runtime"
6 _ "go.uber.org/automaxprocs"
7)
8
9func main() {
10 log.Printf("GOMAXPROCS: %d", runtime.GOMAXPROCS(0))
11 // startup proceeds normally
12}

Real-World Example

A logistics platform running Go microservices on AWS Jakarta (ap-southeast-3) was seeing p99 latency of 800ms on a route that benchmarks showed should run in under 100ms. p50 was 60ms. The gap was suspicious.

Their nodes were c5.4xlarge 16 vCPUs. Each pod had resources.limits.cpu: "1". Go was starting 16 threads per pod, competing for 1 CPU-second per second.

They added automaxprocs, redeployed, and watched the metrics. GOMAXPROCS dropped from 16 to 1. CPU throttling events dropped to near zero. p99 went from 800ms to 95ms over the next 30 minutes.

The fix took 10 minutes to implement. The investigation took two days.

The two-day investigation happened because the team was looking at the right metrics latency, CPU usage, error rates but missing the one that points directly to throttling: container_cpu_cfs_throttled_seconds_total in Prometheus. That metric shows cumulative CPU throttling per container. If you see it growing under load, check GOMAXPROCS next.


FAQ

Q: Does this affect all Go versions, or was it fixed at some point? A: All versions as of this writing. Go's runtime.NumCPU() reads hardware CPU count, not cgroup quota a deliberate decision. The runtime doesn't assume it's running in a container. automaxprocs was created specifically to bridge this gap. There's no automatic cgroup-aware GOMAXPROCS in the standard library yet.

Q: What if my pod has no CPU limit set? A: automaxprocs falls back to runtime.NumCPU(), same as default behavior. No limit means you can use the full node, so using all CPU threads is correct. The problem only appears when limits are set and the node has significantly more CPUs than your quota.

Q: Should I set GOMAXPROCS higher than my CPU limit to compensate for I/O blocking? A: Generally no. Go's goroutine scheduler handles I/O efficiently without excess threads blocked goroutines don't hold an OS thread. Setting GOMAXPROCS higher than your CPU quota means deliberately increasing throttling exposure. Match your CPU limit and let the scheduler do its job.

Q: How do I verify the fix is working? A: Log runtime.GOMAXPROCS(0) at startup. In Prometheus, check container_cpu_cfs_throttled_seconds_total before and after deployment. If throttling drops significantly, you've confirmed the cause. p99 latency improvement under sustained load is the more meaningful business signal.

Q: Does this apply to non-Kubernetes container runtimes like Docker or ECS? A: Yes. Any environment that uses Linux cgroups to limit CPU Docker, ECS, containerd has this behavior. automaxprocs works by reading /sys/fs/cgroup/cpu/cpu.cfs_quota_us and cpu.cfs_period_us directly, so it works wherever cgroup v1 or v2 CPU limits are enforced.


Kubernetes resource limits and Go's runtime make assumptions about the environment that don't always align. This is one of the most impactful, least-diagnosed performance issues in containerised Go services not because it's obscure, but because the symptoms look like a dozen other things first.

Add automaxprocs to every Go service you deploy to Kubernetes. One import, no configuration. The kind of thing you want in place before you're debugging latency at 2am.


Internal links used:

External links used:

Word count: ~1,450

Note: Graceful shutdown post (I-6) not yet published left as HTML comment for future linking.

// END_OF_LOGSPECTRE_SYSTEMS_V1

Is your current architecture slowing you down?

Stop guessing where the bottlenecks are. We partner with founders and CTOs to audit technical debt and execute zero-downtime system rewrites.

Book an Architecture Audit