When a client tells us their Kubernetes bill is out of control, we don't open the AWS console first. We open their deployment manifests.
Nine times out of ten, the bill is the symptom. The cause is three architectural choices made in week one that cascade.
Choice one: unbounded resource requests. Pods request 2 CPUs they never use. The cluster pre-provisions for the request, not the actual load. Right-sizing those manifests typically reclaims 30–45% of cluster capacity without touching a single line of application code.
Choice two: no cluster autoscaler tuning. Karpenter or Cluster Autoscaler with defaults will keep nodes warm long after the workload has scaled down. Tighten the cooldown, set scaledown utilisation thresholds, and your bill bends the moment traffic does.
Choice three: dev/staging on production-grade nodes. We've seen entire dev clusters running on m5.4xlarge instances because someone copied prod's nodegroup template. Right-sized t-class spot instances handle the workload at a fraction of the cost.
Combined, these three got Helios Energy's cluster bill down 38% in eight weeks — with zero impact on uptime.