Alerts
- Alerting
The controller's Kubernetes Events on StageSet transitions, the opt-in PrometheusRule alert catalog with tunable thresholds, and how every alert and Ready reason links to a runbook.
- Controller pod down
A stageset-controller pod has been NotReady for the alert window.
- Observability
The four observability pillars for the controller — structured logging, OTLP tracing, Prometheus metrics, and the opt-in alert catalog — plus service level objectives and a ready-made dashboard.
- Operations
Metrics, alerts, events, and runbooks for running the controller day to day.
- Reconcile latency high
Reconcile p99 latency for the StageSet controller is above threshold.
- Service level objectives
The StageSet controller's SLOs — reconcile availability and reconcile latency — with their SLIs, targets, error budget, the dashboard that shows them, and how to tune them.
- Workqueue saturation
The controller cannot drain its reconcile queue fast enough.