Topic 2.5
ResourceQuotas & LimitRanges: Sharing a Cluster Fairly
In one line
When several teams share a cluster, one team's runaway deployment shouldn't use all the CPU, memory, or disks. A ResourceQuota caps the total a namespace can request (CPU, memory, storage, number of pods, Services, LoadBalancers). A LimitRange sets default and maximum requests and limits for each container in a namespace, so pods without resource settings still get sensible ones.
Think of it like this
A shared office kitchen. Each team gets a shelf of a set size in the fridge (quota), and there's a rule that no single container may be bigger than a 2-litre box, with a default box size if you don't bring your own (limit range).
Words you'll meet
New words in this topic, in plain English. Come back here whenever one feels fuzzy.
- ResourceQuota
- A limit on the total resources and object counts a namespace can use.
- LimitRange
- Per-container or per-pod defaults and bounds for resource requests and limits in a namespace.
- Admission
- The step where the API server checks a new object against rules before saving it.
- Requests vs limits
- Requests reserve resources for scheduling, limits cap what a container may use (Topic 4.1).
- Noisy neighbour
- A workload that uses so many shared resources it degrades others.
Step by step
01A quota and defaults for the dispatch team
The platform team gives the dispatch namespace a budget: 16 CPUs and 48 GiB of memory requested in total, 60 pods, one LoadBalancer, and 200 GiB of storage. A LimitRange supplies defaults for containers that don't set their own.
apiVersion: v1
kind: ResourceQuota
metadata: { name: dispatch-quota, namespace: dispatch }
spec:
hard:
requests.cpu: "16"
requests.memory: 48Gi
limits.memory: 72Gi
pods: "60"
services.loadbalancers: "1"
requests.storage: 200Gi
---
apiVersion: v1
kind: LimitRange
metadata: { name: dispatch-defaults, namespace: dispatch }
spec:
limits:
- type: Container
defaultRequest: { cpu: 100m, memory: 256Mi }
default: { memory: 512Mi } # default limit
max: { cpu: "4", memory: 8Gi }02Defaults filled in automatically
A developer deploys a debugging pod with no resources section. The LimitRange adds the defaults, so the pod is admitted and counted in the quota.
03Hitting the ceiling
During a load test, the dispatch team scales the matcher to 50 replicas. Pods stop being created once the quota is reached, and the reason is in the ReplicaSet's events.
Break it on purpose
Errors are the best teachers. Make each change, read the error, guess what went wrong, then reveal the answer.
Break #1
A quota without defaults blocks every deploy
The platform team adds a ResourceQuota with requests.cpu to the tiffin namespace but no LimitRange, and some manifests don't set requests.
Myth vs fact
Myth
A quota limits how much CPU pods actually use.
Fact
Quotas count what pods request (and their declared limits) at creation. Actual usage at runtime is controlled by limits and node capacity (Topic 4.1).
Pro corner
Extra depth for experienced readers. New to this? Skip it for now and come back later.
- ▸
Alert when a namespace uses more than 80% of a quota (
kube_resourcequotametrics from kube-state-metrics), so teams ask for more before deploys start failing during an incident.
Remember this
- 1
ResourceQuota (per namespace): limits like
requests.cpu: "20",requests.memory: 64Gi,limits.memory: 96Gi,pods: "100",services.loadbalancers: "2",requests.storage: 500Gi, and per StorageClass storage. - 2
When a quota covers CPU or memory, every new pod must set requests/limits for them, or it's rejected. A LimitRange's defaults solve that.
- 3
Quotas are checked at admission: creating a pod that would exceed the quota fails with
exceeded quota. Existing pods aren't evicted when you lower a quota. - 4
LimitRange:
default(limits) anddefaultRequestfor containers without them, plusmin/maxper container or pod, andmaxLimitRequestRatio. - 5
Deployments don't fail visibly when a quota blocks pods: the ReplicaSet can't create them, and the error appears in its events. Check
kubectl describe rs. - 6
Use quotas for fairness and cost control in shared clusters, and pair them with monitoring of quota usage (
kubectl describe quota).
Explain it without notes
What's the difference between a ResourceQuota and a LimitRange?
Why must pods specify requests when a CPU quota exists?
Practice
Create a namespace with a quota of 2 pods and try to create a third.
Add a LimitRange with defaults and check that a pod without resources gets them.
Trade-offs
- ↔
Quotas prevent one team from starving others and keep costs predictable, but too-tight quotas block legitimate scaling during spikes. Review them regularly against real usage and give teams a fast way to request more.
Done when you can
I can set ResourceQuotas and LimitRanges on a namespace.
I always pair CPU/memory quotas with defaults.
I check ReplicaSet events when a Deployment doesn't scale.