October 5, 2026

How Rightsizing Kubernetes Resource Requests Cut Cloud Costs 40% and Alert Fatigue 80%

The problem: EKS nodes scaling on the wrong signal

One EKS cluster I worked on had grown node count steadily for months, and nobody could explain why. Traffic hadn't grown nearly as fast as the bill had. At the same time, the on-call channel for that cluster had become background noise: CPU and memory alerts on individual nodes fired constantly, enough that engineers had stopped reacting to them in real time and just checked in during business hours.

Both problems traced back to the same mix-up: the cluster was making decisions, both scaling decisions and alerting decisions, off resource requests (what pods reserve) instead of actual utilization (what pods use). Those are two different numbers, and conflating them had been quietly expensive in both directions.

Why cost and alert fatigue were the same problem

Kubernetes scheduling runs on requests. The cluster autoscaler (Karpenter, in this case) adds a node when pending pods can't fit into the allocatable capacity of existing nodes, and "allocatable capacity" is computed from what's requested, not what's actually consumed. Pod requests across this workload had been set defensively months earlier, well above real usage, to avoid any risk of throttling or an OOM kill under load. Nobody had revisited them since.

Requests are a reservation, not a forecast of real usage. A pod that requests 2 vCPUs and actually uses 200m of it still blocks 2 vCPUs of scheduling room on its node. Scale enough pods like that across enough nodes, and the cluster looks full on paper while sitting mostly idle in practice. A cluster autoscaler will dutifully keep adding nodes to satisfy that paper fullness.

That inflated-requests problem explains the cost side directly: the autoscaler kept provisioning nodes to satisfy reserved capacity that was never actually used, so the cluster carried far more EC2 capacity than the workload needed. The alerting side came from the opposite direction. Node-level CPU and memory alerts were built on raw utilization thresholds, the kind you'd use on a plain EC2 host, without accounting for how densely (or loosely) a node was packed by scheduling. A node running near its requested capacity but low actual usage could sit for hours looking "fine" by one metric and "full" by another, and the alert logic couldn't tell real pressure (a node actually throttling pods, or close to an eviction) from a scheduling artifact that was never going to cause an incident. Same root confusion, cost on one side, noise on the other.

Checking it yourself: reserved vs. actual, side by side

You don't need a dashboard to see this gap. Two commands against any node show both sides of it at once:

$ kubectl describe node ip-10-0-12-34.ec2.internal
...
Allocatable:
  cpu:                3920m
  memory:             15841428Ki
Allocated resources:
  Resource           Requests      Limits
  --------           --------      ------
  cpu                3200m (81%)   4000m (102%)
  memory             10Gi (66%)    12Gi (79%)

$ kubectl top node ip-10-0-12-34.ec2.internal
NAME                          CPU(cores)   CPU%   MEMORY(bytes)   MEMORY%
ip-10-0-12-34.ec2.internal    420m         10%    3812Mi          25%

kubectl describe node's "Allocated resources" table is the reservation side: this node is committed to 81% of its CPU by pod requests. kubectl top node (needs metrics-server running in the cluster) is the actual side: only 420m, about 10%, is really in use. Line the two up across enough nodes and the mismatch driving the extra node count becomes obvious instead of theoretical.

The fix: rightsizing requests and splitting the metrics

Fixing this meant two things, done together. First, rightsizing: I pulled several weeks of actual per-pod CPU and memory usage from CloudWatch Container Insights and compared it against each workload's configured requests. Most services were requesting two to four times their observed p95 usage. I brought requests down to match real usage, not the defensive guess baked in originally, using the Vertical Pod Autoscaler's recommendation mode to validate the new numbers against live traffic before committing to them.

Requests and limits got sized off different numbers on purpose. A request near p95 is fine for CPU, since CPU is compressible: going over it just throttles the pod, it doesn't kill it. Memory is a different story. Exceeding a memory limit gets the pod OOM-killed outright, so that number needs to come from closer to the observed max, not p95, where by definition 5% of samples already sit above it. A sustained traffic spike is handled by the Horizontal Pod Autoscaler adding replicas, not by inflating one pod's request to cover a case that mostly won't happen.

# before: a defensive guess, set once and never revisited
resources:
  requests:
    cpu: "2000m"
    memory: "4Gi"
  limits:
    cpu: "2000m"
    memory: "4Gi"

# after: request near observed p95, memory limit near observed max (CPU limit left
# generous since throttling is recoverable; OOM kills are not)
resources:
  requests:
    cpu: "500m"
    memory: "1Gi"
  limits:
    cpu: "1000m"
    memory: "2Gi"

That request-equals-limit pattern in the "before" block wasn't incidental. It's the textbook definition of Kubernetes' Guaranteed QoS class: the kubelet treats requests matching limits on every container as a promise worth protecting above all else, which is also why those pods were last in line to ever get evicted under pressure. The "after" config, with requests below limits, is Burstable instead: still protected ahead of BestEffort pods, but no longer hoarding guaranteed capacity nobody was using. Giving up eviction priority you were never actually at risk of needing is a reasonable trade for the capacity you get back today.

That one change, repeated across the handful of services carrying the worst mismatch, is what let the cluster autoscaler's node math start matching reality instead of the original guess.

Second, splitting the signals that scaling and alerting each depend on. The cluster autoscaler kept doing what it does well, scaling nodes off scheduling pressure and requested capacity, but now that requests actually meant something close to real usage, node count tracked real demand instead of a reservation that was mostly fiction. Alerting moved to a different axis entirely: actual node and pod CPU/memory utilization, kubelet eviction signals, and throttling metrics, which tell you when a node is genuinely under pressure, separated from whatever the scheduler-facing capacity numbers say. Two different questions ("can we schedule more work" and "is a node actually struggling") finally had two different metrics answering them, instead of one noisy threshold trying to answer both.

# scheduling-facing: what's reserved (drives the cluster autoscaler)
sum(kube_pod_container_resource_requests{resource="cpu"}) by (node)

# pressure-facing: what's actually being used (drives alerting)
sum(rate(container_cpu_usage_seconds_total[5m])) by (node)

Alerting rules got pointed at the second query and its memory equivalent, not the first. The first query is still useful, just for a different job: capacity planning, not paging anyone.

The result: 40% lower costs, 80% fewer noisy alerts

The cluster's EC2 spend dropped 40%, in line with the average cost reduction I see on engagements like this one, once node count actually reflected real demand instead of inflated reservations. Alert volume on that cluster dropped 80%, because alerts now corresponded to conditions that were either real or close to it, instead of firing on the gap between what was reserved and what was used.

The alerting improvement mattered beyond comfort. A clean signal means the handful of alerts that do fire get taken seriously and get a fast response, instead of getting lost in a channel everyone has learned to tune out.

It shows up in MTTR too. When a real incident hits, the on-call engineer isn't starting from a channel that's cried wolf eighty times that week, deciding whether this alert is the real thing or more of the same noise. The alerts that do fire are credible by default, so triage starts immediately instead of a few minutes late.

Signs your cluster has the same problem

A few things worth checking if this sounds familiar:

  • Pod resource requests were set once, early on, "to be safe," and nobody has compared them to actual usage since.
  • Your cluster autoscaler keeps adding nodes while CloudWatch Container Insights (or Prometheus) shows real CPU and memory utilization staying low.
  • Node-level alerts fire often enough that people have started ignoring the channel, even though nothing seems to actually be wrong when someone checks.
  • Nobody on the team can say, with confidence, whether a given dashboard metric reflects what's reserved or what's actually being used.

One footnote, if you're migrating workloads over from ECS: Kubernetes CPU is decimal, not ECS's 1024-units-per-vCPU convention. cpu: "1" is one full vCPU, and 500m (500 millicpu) is half of one. There's no 1024 conversion to do, just move the decimal. Memory does the opposite of what you'd expect by default: Gi/Mi are the binary (1024-based) suffixes, while the plain G/M are decimal and slightly smaller for the same-looking number, worth checking carefully when porting an ECS task definition's MiB values into a Kubernetes manifest.

If any of that sounds familiar, it's a fixable mismatch, not a fundamentally hard problem. Finding it, and rebuilding both the autoscaling and the alerting around the right metric, is exactly the kind of AWS Cost & Security Audit and Observability & Monitoring work I can help with at Cloud with Gus.