Topic 9.6
Running Kubernetes Cost-Efficiently
In one line
Kubernetes costs come mostly from nodes, and nodes are sized by the requests pods make, not by what they use. Most clusters waste a large share of capacity through oversized requests, poor bin-packing, idle environments, and expensive instance choices. The levers: right-size requests from real usage, let autoscalers pack and remove nodes, use spot and ARM instances where safe, scale idle environments to zero, and make cost visible per team with labels and tools like OpenCost.
Think of it like this
Renting hotel rooms for a conference. If every attendee books a suite 'just in case' (high requests), you pay for empty space. Booking what people actually need, sharing rooms well (bin-packing), and cancelling rooms nobody uses (idle environments) cuts the bill without making anyone uncomfortable.
Words you'll meet
New words in this topic, in plain English. Come back here whenever one feels fuzzy.
- Right-sizing
- Setting resource requests close to what workloads actually use.
- Bin-packing
- Placing pods onto nodes so that node capacity is used efficiently.
- Spot instances
- Spare cloud capacity at a large discount that can be reclaimed with short notice.
- Consolidation
- An autoscaler moving pods off underused nodes and removing those nodes.
- OpenCost
- An open-source tool that allocates Kubernetes costs to namespaces, labels, and workloads.
- Idle cost
- Money spent on capacity that's reserved or running but not doing useful work.
Step by step
01Where the money goes
Tiffin's production cluster costs about ₹9 lakh a month. OpenCost shows CPU requests at three times usage across most namespaces, and the dev namespaces running 24/7.
02Three changes
The team right-sizes the biggest Deployments from VPA recommendations, lets Karpenter use spot and Graviton instances for stateless workloads, and scales dev to zero outside 9:00-21:00 India time.
apiVersion: karpenter.sh/v1
kind: NodePool
metadata: { name: general }
spec:
template:
spec:
requirements:
- { key: karpenter.sh/capacity-type, operator: In, values: [spot, on-demand] }
- { key: kubernetes.io/arch, operator: In, values: [arm64, amd64] }
- { key: karpenter.k8s.aws/instance-category, operator: In, values: [c, m, r] }
nodeClassRef: { group: karpenter.k8s.aws, kind: EC2NodeClass, name: default }
disruption:
consolidationPolicy: WhenEmptyOrUnderutilized
consolidateAfter: 2m03Scaling dev to zero at night
A KEDA cron trigger keeps dev services at one replica during working hours and zero otherwise. Karpenter then removes the empty nodes.
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata: { name: web, namespace: dev }
spec:
scaleTargetRef: { name: web }
minReplicaCount: 0
triggers:
- type: cron
metadata:
timezone: Asia/Kolkata
start: 0 9 * * 1-5
end: 0 21 * * 1-5
desiredReplicas: "1"Break it on purpose
Errors are the best teachers. Make each change, read the error, guess what went wrong, then reveal the answer.
Break #1
Spot for everything
To save money, every workload, including the single-replica payments API and Postgres for staging, runs on spot nodes.
Myth vs fact
Myth
Low CPU usage on nodes means the cluster is efficient.
Fact
Nodes are paid for by size. Low usage with high requests means you're paying for reserved but idle capacity.
Pro corner
Extra depth for experienced readers. New to this? Skip it for now and come back later.
- ▸
Make cost a regular engineering metric: a monthly review of the biggest namespaces' efficiency, with right-sizing changes as ordinary pull requests. Small, steady changes are safer than a big cost-cutting push.
Remember this
- 1
Requests drive cost: the scheduler reserves requested CPU and memory, the autoscaler adds nodes for pending requests, so oversized requests mean more nodes even when CPU sits at 10%.
- 2
Right-sizing: compare requests with real usage (p95 over weeks) using Prometheus, VPA recommendations, Goldilocks, or Kubecost/OpenCost, and adjust gradually.
- 3
Bin-packing: Karpenter consolidation and Cluster Autoscaler scale-down remove underused nodes. Large pods on small nodes, or many node groups, leave gaps.
- 4
Cheaper capacity: spot instances for stateless and batch workloads (with PDBs and diversified instance types), ARM instances (Graviton) with multi-arch images, savings plans for the steady baseline.
- 5
Idle environments: scale dev and preview namespaces to zero at night (KEDA cron, or a scheduler), and delete preview environments when pull requests close.
- 6
Visibility: label workloads with team and app, install OpenCost/Kubecost, and show each team its cost. Watch hidden costs too: cross-zone traffic, NAT gateways, load balancers, and storage.
Explain it without notes
Why do requests, not usage, drive Kubernetes cost?
Which workloads suit spot instances?
Practice
Find the 5 workloads with the biggest gap between requests and usage.
Plan a schedule that scales a dev namespace to zero at night and on weekends.
Trade-offs
- ↔
Aggressive right-sizing and consolidation save money but leave less headroom and cause more pod moves. Spot is much cheaper but interruptible. Visibility tools cost effort but make savings targeted rather than guessed.
Done when you can
I size requests from real usage.
I use spot and ARM capacity where it's safe.
Idle environments scale down and costs are visible per team.