Master Metrics & Dashboards
From metrics vs. logs vs. traces to PromQL, Grafana dashboards, alerting rules that don't page you for nothing, and the retention/cardinality traps that fill production disks - paired with real corporate-incident labs against a genuine Prometheus and Grafana stack.
Modules
- What is Monitoring, Really? - Metrics vs. logs vs. traces, pull vs. push monitoring, and why Prometheus's data model of metric name + labels became the industry default.
- Prometheus Fundamentals - Prometheus's architecture (scraper, TSDB, HTTP API), writing a scrape_config, and the first PromQL queries: instant vectors, range vectors, and rate().
- Grafana Fundamentals - How Grafana relates to Prometheus, datasources, dashboards vs. panels, and a plain-language intro to dashboard variables.
- Alerting: From Metrics to Notifications - Writing Prometheus alerting rules with expr, for, labels and annotations, why the for duration prevents alert fatigue, and Alertmanager's role in routing, grouping and silencing.
- Production Concerns - Retention and disk growth, cardinality explained plainly, a forward look at service discovery, and a wrap-up tying real incidents to the five corporate-incident labs.