Topic 6.5
Logs, Metrics & Events: Observing a Cluster
In one line
Containers' logs vanish when pods are deleted, metrics-server keeps no history, and events expire after an hour, so production clusters need an observability stack. Prometheus scrapes metrics from apps, kubelets, and kube-state-metrics, Grafana shows dashboards, and Alertmanager pages people. A log agent (Fluent Bit, Promtail/Alloy) ships container logs to Loki, OpenSearch, or a cloud service. OpenTelemetry adds traces. Alert on symptoms users feel, backed by cluster signals.
Think of it like this
A hospital ward. Monitors track vital signs over time (metrics), nurses write notes (logs), and the ward log records admissions and transfers (events). Alarms (alerts) call the doctor only when a patient needs attention.
Words you'll meet
New words in this topic, in plain English. Come back here whenever one feels fuzzy.
- Observability
- Being able to understand a system's behaviour from its outputs: metrics, logs, events, and traces.
- Prometheus
- A time-series database that scrapes and stores metrics and evaluates alert rules.
- kube-state-metrics
- An exporter turning Kubernetes object state into metrics, such as restarts and replicas.
- Loki
- A log store that indexes logs by labels, often used with Grafana.
- ServiceMonitor
- A custom resource telling the Prometheus Operator which Services to scrape.
- Alertmanager
- The component that groups, routes, and sends Prometheus alerts to people.
Step by step
01The stack around the cluster
Tiffin installs kube-prometheus-stack for metrics and alerts, Fluent Bit shipping logs to Loki, and Grafana on top of both.
02Scraping the web app
Spring Boot exposes Micrometer metrics at /actuator/prometheus. A ServiceMonitor tells Prometheus to scrape every Service labelled app: web on its http port.
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: web
namespace: tiffin
labels: { release: monitoring } # picked up by the stack's Prometheus
spec:
selector: { matchLabels: { app: web } }
endpoints:
- port: http
path: /actuator/prometheus
interval: 30s03Questions the stack answers
With history in place, the team can ask useful questions: which pods restarted in the last day, what's the 5xx rate, and what did the crashed pod log before it died.
# pods with restarts in the last 24h (kube-state-metrics)
sort_desc(increase(kube_pod_container_status_restarts_total{namespace="tiffin"}[24h]) > 0)
# web 5xx ratio over 5 minutes (Micrometer)
sum(rate(http_server_requests_seconds_count{namespace="tiffin",status=~"5.."}[5m]))
/ sum(rate(http_server_requests_seconds_count{namespace="tiffin"}[5m]))
# CPU throttling ratio per pod (cAdvisor)
sum by (pod) (rate(container_cpu_cfs_throttled_periods_total{namespace="tiffin"}[5m]))
/ sum by (pod) (rate(container_cpu_cfs_periods_total{namespace="tiffin"}[5m]))
# Loki: errors from a pod that no longer exists
{namespace="tiffin", pod=~"orders-worker-6c8b7d-.*"} |= "ERROR"Break it on purpose
Errors are the best teachers. Make each change, read the error, guess what went wrong, then reveal the answer.
Break #1
High-cardinality labels melt Prometheus
A developer adds orderId as a label to an order-processing metric.
Myth vs fact
Myth
kubectl logs is our logging solution.
Fact
It reads logs from the node, which disappear when the pod is deleted or rotated. Without central logging, crashes and incidents leave no evidence.
Pro corner
Extra depth for experienced readers. New to this? Skip it for now and come back later.
- ▸
Start alerting from SLOs (error rate and latency burn rates) and a short list of cluster health alerts. kube-prometheus-stack ships many default alerts: review them and route only actionable ones to pagers.
Remember this
- 1
Logs: apps write to stdout/stderr, the container runtime saves them on the node under
/var/log/pods, and a DaemonSet agent (Topic 1.5) ships them with labels (namespace, pod, container) to a central store. - 2
Metrics: Prometheus scrapes
/metricsendpoints. Key sources: your apps (Micrometer in Spring Boot), kubelet/cAdvisor (container CPU, memory, throttling), kube-state-metrics (object state: replicas, restarts, pod phases), node-exporter (node hardware). - 3
kube-prometheus-stack (Helm) installs Prometheus, Alertmanager, Grafana, exporters, and ready-made dashboards and alerts. ServiceMonitor/PodMonitor objects tell Prometheus what to scrape.
- 4
Events: export them (kubernetes-event-exporter, or your logging agent) so 'why was this pod killed yesterday' has an answer.
- 5
Alerts that matter: SLO-based alerts on error rate and latency, plus a few cluster ones: pods crash-looping, pods Pending for over 10 minutes, nodes NotReady, PVCs nearly full, certificates expiring, HPAs at max.
- 6
Traces: OpenTelemetry instrumentation and a collector send traces to Tempo or Jaeger, linking requests across services (Observability course).
Explain it without notes
Why do you need kube-state-metrics in addition to kubelet metrics?
Why should logs go to a central store?
Practice
Install kube-prometheus-stack on kind and find a dashboard showing pod CPU and memory.
Write an alert for pods stuck Pending for more than 10 minutes.
Trade-offs
- ↔
Self-hosted Prometheus and Loki give control and low cost per GB, but need storage and operations. Managed services (Amazon Managed Prometheus, Grafana Cloud, Datadog) cost more and remove the operational work.
Done when you can
My cluster ships logs and events centrally.
Prometheus scrapes apps, kubelets, and kube-state-metrics.
I alert on user-facing symptoms and a few key cluster signals.