Command Palette

Search for a command to run...

Hectal
PHASE 6Intermediate ~15 min· topic 5 of 5

Topic 6.5

Logs, Metrics & Events: Observing a Cluster

In one line

Containers' logs vanish when pods are deleted, metrics-server keeps no history, and events expire after an hour, so production clusters need an observability stack. Prometheus scrapes metrics from apps, kubelets, and kube-state-metrics, Grafana shows dashboards, and Alertmanager pages people. A log agent (Fluent Bit, Promtail/Alloy) ships container logs to Loki, OpenSearch, or a cloud service. OpenTelemetry adds traces. Alert on symptoms users feel, backed by cluster signals.

0/5 · 0%

Think of it like this

A hospital ward. Monitors track vital signs over time (metrics), nurses write notes (logs), and the ward log records admissions and transfers (events). Alarms (alerts) call the doctor only when a patient needs attention.

Words you'll meet

New words in this topic, in plain English. Come back here whenever one feels fuzzy.

Observability
Being able to understand a system's behaviour from its outputs: metrics, logs, events, and traces.
Prometheus
A time-series database that scrapes and stores metrics and evaluates alert rules.
kube-state-metrics
An exporter turning Kubernetes object state into metrics, such as restarts and replicas.
Loki
A log store that indexes logs by labels, often used with Grafana.
ServiceMonitor
A custom resource telling the Prometheus Operator which Services to scrape.
Alertmanager
The component that groups, routes, and sends Prometheus alerts to people.

Step by step

01The stack around the cluster

Tiffin installs kube-prometheus-stack for metrics and alerts, Fluent Bit shipping logs to Loki, and Grafana on top of both.

terminal
$ helm install monitoring prometheus-community/kube-prometheus-stack -n monitoring --create-namespace
kubectl get pods -n monitoring
── expected output ──
NAME: monitoring
STATUS: deployed
NAME READY STATUS RESTARTS AGE
alertmanager-monitoring-kube-prometheus-alertmanager-0 2/2 Running 0 2m
monitoring-grafana-6d7c9b8f5-2xk8p 3/3 Running 0 2m
monitoring-kube-state-metrics-7b9d6c5f4-9pl4m 1/1 Running 0 2m
monitoring-prometheus-node-exporter-4k2xq 1/1 Running 0 2m
prometheus-monitoring-kube-prometheus-prometheus-0 2/2 Running 0 2m
The stack around the clusterdiagram
Rendering diagram…

02Scraping the web app

Spring Boot exposes Micrometer metrics at /actuator/prometheus. A ServiceMonitor tells Prometheus to scrape every Service labelled app: web on its http port.

k8s/web-servicemonitor.yamlwhole fileyaml
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
  name: web
  namespace: tiffin
  labels: { release: monitoring }        # picked up by the stack's Prometheus
spec:
  selector: { matchLabels: { app: web } }
  endpoints:
    - port: http
      path: /actuator/prometheus
      interval: 30s

03Questions the stack answers

With history in place, the team can ask useful questions: which pods restarted in the last day, what's the 5xx rate, and what did the crashed pod log before it died.

queries (PromQL and LogQL)whole filetext
# pods with restarts in the last 24h (kube-state-metrics)
sort_desc(increase(kube_pod_container_status_restarts_total{namespace="tiffin"}[24h]) > 0)

# web 5xx ratio over 5 minutes (Micrometer)
sum(rate(http_server_requests_seconds_count{namespace="tiffin",status=~"5.."}[5m]))
  / sum(rate(http_server_requests_seconds_count{namespace="tiffin"}[5m]))

# CPU throttling ratio per pod (cAdvisor)
sum by (pod) (rate(container_cpu_cfs_throttled_periods_total{namespace="tiffin"}[5m]))
  / sum by (pod) (rate(container_cpu_cfs_periods_total{namespace="tiffin"}[5m]))

# Loki: errors from a pod that no longer exists
{namespace="tiffin", pod=~"orders-worker-6c8b7d-.*"} |= "ERROR"

Break it on purpose

Errors are the best teachers. Make each change, read the error, guess what went wrong, then reveal the answer.

Break #1

High-cardinality labels melt Prometheus

A developer adds orderId as a label to an order-processing metric.

terminal
$ kubectl top pod -n monitoring prometheus-monitoring-kube-prometheus-prometheus-0
kubectl logs -n monitoring prometheus-monitoring-kube-prometheus-prometheus-0 -c prometheus --tail=1
── what you'll see ──
NAME CPU(cores) MEMORY(bytes)
prometheus-monitoring-kube-prometheus-prometheus-0 3820m 15302Mi
ts=2026-10-02T13:02:11Z level=warn msg="Head series count" series=8412004
Prometheus is OOMKilled every few hours, so dashboards and alerts go blank.

Myth vs fact

Myth

kubectl logs is our logging solution.

Fact

It reads logs from the node, which disappear when the pod is deleted or rotated. Without central logging, crashes and incidents leave no evidence.

Pro corner

Extra depth for experienced readers. New to this? Skip it for now and come back later.

  • ▸

    Start alerting from SLOs (error rate and latency burn rates) and a short list of cluster health alerts. kube-prometheus-stack ships many default alerts: review them and route only actionable ones to pagers.

Remember this

  1. 1

    Logs: apps write to stdout/stderr, the container runtime saves them on the node under /var/log/pods, and a DaemonSet agent (Topic 1.5) ships them with labels (namespace, pod, container) to a central store.

  2. 2

    Metrics: Prometheus scrapes /metrics endpoints. Key sources: your apps (Micrometer in Spring Boot), kubelet/cAdvisor (container CPU, memory, throttling), kube-state-metrics (object state: replicas, restarts, pod phases), node-exporter (node hardware).

  3. 3

    kube-prometheus-stack (Helm) installs Prometheus, Alertmanager, Grafana, exporters, and ready-made dashboards and alerts. ServiceMonitor/PodMonitor objects tell Prometheus what to scrape.

  4. 4

    Events: export them (kubernetes-event-exporter, or your logging agent) so 'why was this pod killed yesterday' has an answer.

  5. 5

    Alerts that matter: SLO-based alerts on error rate and latency, plus a few cluster ones: pods crash-looping, pods Pending for over 10 minutes, nodes NotReady, PVCs nearly full, certificates expiring, HPAs at max.

  6. 6

    Traces: OpenTelemetry instrumentation and a collector send traces to Tempo or Jaeger, linking requests across services (Observability course).

Explain it without notes

01

Why do you need kube-state-metrics in addition to kubelet metrics?

02

Why should logs go to a central store?

Practice

01

Install kube-prometheus-stack on kind and find a dashboard showing pod CPU and memory.

02

Write an alert for pods stuck Pending for more than 10 minutes.

Trade-offs

  • ↔

    Self-hosted Prometheus and Loki give control and low cost per GB, but need storage and operations. Managed services (Amazon Managed Prometheus, Grafana Cloud, Datadog) cost more and remove the operational work.

Done when you can

  • My cluster ships logs and events centrally.

  • Prometheus scrapes apps, kubelets, and kube-state-metrics.

  • I alert on user-facing symptoms and a few key cluster signals.