Command Palette

Search for a command to run...

Hectal
PHASE 9Advanced ~15 min· topic 6 of 6

Topic 9.6

Running Kubernetes Cost-Efficiently

In one line

Kubernetes costs come mostly from nodes, and nodes are sized by the requests pods make, not by what they use. Most clusters waste a large share of capacity through oversized requests, poor bin-packing, idle environments, and expensive instance choices. The levers: right-size requests from real usage, let autoscalers pack and remove nodes, use spot and ARM instances where safe, scale idle environments to zero, and make cost visible per team with labels and tools like OpenCost.

0/6 · 0%

Think of it like this

Renting hotel rooms for a conference. If every attendee books a suite 'just in case' (high requests), you pay for empty space. Booking what people actually need, sharing rooms well (bin-packing), and cancelling rooms nobody uses (idle environments) cuts the bill without making anyone uncomfortable.

Words you'll meet

New words in this topic, in plain English. Come back here whenever one feels fuzzy.

Right-sizing
Setting resource requests close to what workloads actually use.
Bin-packing
Placing pods onto nodes so that node capacity is used efficiently.
Spot instances
Spare cloud capacity at a large discount that can be reclaimed with short notice.
Consolidation
An autoscaler moving pods off underused nodes and removing those nodes.
OpenCost
An open-source tool that allocates Kubernetes costs to namespaces, labels, and workloads.
Idle cost
Money spent on capacity that's reserved or running but not doing useful work.

Step by step

01Where the money goes

Tiffin's production cluster costs about ₹9 lakh a month. OpenCost shows CPU requests at three times usage across most namespaces, and the dev namespaces running 24/7.

terminal
$ kubectl cost namespace --window 30d --show-efficiency | head -6
── expected output ──
+-----------+------------+--------------+------------+
| NAMESPACE | MONTHLY | CPU EFFIC. | MEM EFFIC. |
+-----------+------------+--------------+------------+
| tiffin | ₹3,42,000 | 31% | 58% |
| dispatch | ₹1,96,000 | 22% | 47% |
| dev | ₹1,48,000 | 6% | 19% |
Output from the kubectl-cost plugin (Kubecost/OpenCost). Efficiency = usage ÷ requests.
Where the money goesdiagram
Rendering diagram…

02Three changes

The team right-sizes the biggest Deployments from VPA recommendations, lets Karpenter use spot and Graviton instances for stateless workloads, and scales dev to zero outside 9:00-21:00 India time.

karpenter/nodepool-general.yaml (excerpt)whole fileyaml
apiVersion: karpenter.sh/v1
kind: NodePool
metadata: { name: general }
spec:
  template:
    spec:
      requirements:
        - { key: karpenter.sh/capacity-type, operator: In, values: [spot, on-demand] }
        - { key: kubernetes.io/arch, operator: In, values: [arm64, amd64] }
        - { key: karpenter.k8s.aws/instance-category, operator: In, values: [c, m, r] }
      nodeClassRef: { group: karpenter.k8s.aws, kind: EC2NodeClass, name: default }
  disruption:
    consolidationPolicy: WhenEmptyOrUnderutilized
    consolidateAfter: 2m

03Scaling dev to zero at night

A KEDA cron trigger keeps dev services at one replica during working hours and zero otherwise. Karpenter then removes the empty nodes.

dev/web-scaledobject.yamlwhole fileyaml
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata: { name: web, namespace: dev }
spec:
  scaleTargetRef: { name: web }
  minReplicaCount: 0
  triggers:
    - type: cron
      metadata:
        timezone: Asia/Kolkata
        start: 0 9 * * 1-5
        end: 0 21 * * 1-5
        desiredReplicas: "1"
terminal
$ kubectl cost namespace --window 30d | grep -E 'NAMESPACE|tiffin|dispatch|dev' # a month later
── expected output ──
| NAMESPACE | MONTHLY |
| tiffin | ₹2,05,000 |
| dispatch | ₹1,12,000 |
| dev | ₹41,000 |

Break it on purpose

Errors are the best teachers. Make each change, read the error, guess what went wrong, then reveal the answer.

Break #1

Spot for everything

To save money, every workload, including the single-replica payments API and Postgres for staging, runs on spot nodes.

terminal
$ kubectl get events -A --field-selector reason=SpotInterrupted | tail -2
── what you'll see ──
kube-system Warning SpotInterrupted node/ip-10-0-22-37 Spot instance interruption notice received
payments Warning FailedScheduling pod/payments-api-7d9c-x2l4k 0/5 nodes are available: ...
# payments was down for 4 minutes during the lunch peak

Myth vs fact

Myth

Low CPU usage on nodes means the cluster is efficient.

Fact

Nodes are paid for by size. Low usage with high requests means you're paying for reserved but idle capacity.

Pro corner

Extra depth for experienced readers. New to this? Skip it for now and come back later.

  • ▸

    Make cost a regular engineering metric: a monthly review of the biggest namespaces' efficiency, with right-sizing changes as ordinary pull requests. Small, steady changes are safer than a big cost-cutting push.

Remember this

  1. 1

    Requests drive cost: the scheduler reserves requested CPU and memory, the autoscaler adds nodes for pending requests, so oversized requests mean more nodes even when CPU sits at 10%.

  2. 2

    Right-sizing: compare requests with real usage (p95 over weeks) using Prometheus, VPA recommendations, Goldilocks, or Kubecost/OpenCost, and adjust gradually.

  3. 3

    Bin-packing: Karpenter consolidation and Cluster Autoscaler scale-down remove underused nodes. Large pods on small nodes, or many node groups, leave gaps.

  4. 4

    Cheaper capacity: spot instances for stateless and batch workloads (with PDBs and diversified instance types), ARM instances (Graviton) with multi-arch images, savings plans for the steady baseline.

  5. 5

    Idle environments: scale dev and preview namespaces to zero at night (KEDA cron, or a scheduler), and delete preview environments when pull requests close.

  6. 6

    Visibility: label workloads with team and app, install OpenCost/Kubecost, and show each team its cost. Watch hidden costs too: cross-zone traffic, NAT gateways, load balancers, and storage.

Explain it without notes

01

Why do requests, not usage, drive Kubernetes cost?

02

Which workloads suit spot instances?

Practice

01

Find the 5 workloads with the biggest gap between requests and usage.

02

Plan a schedule that scales a dev namespace to zero at night and on weekends.

Trade-offs

  • ↔

    Aggressive right-sizing and consolidation save money but leave less headroom and cause more pod moves. Spot is much cheaper but interruptible. Visibility tools cost effort but make savings targeted rather than guessed.

Done when you can

  • I size requests from real usage.

  • I use spot and ARM capacity where it's safe.

  • Idle environments scale down and costs are visible per team.