SRE & Observability
The planned outline of 19 modules. Lessons are written, run and reviewed before they are published.
Before this track
Planned outline
- Module 1
Foundations: what SRE is and how reliability is measured
Coming soon - Module 2
Service level objectives: indicators, objectives and error budgets
Coming soon - Module 3
Signals: monitoring, observability and metric design
Coming soon - Module 4
Prometheus foundations: scraping, instrumenting and PromQL
Coming soon - Module 5
Prometheus in depth: advanced PromQL, rules, cardinality and scale
Coming soon - Module 6
SLOs in practice: SLI rules, burn-rate alerts and tooling
Coming soon - Module 7
Alerting, Alertmanager and sustainable on-call
Coming soon - Module 8
Distributed tracing and OpenTelemetry instrumentation
Coming soon - Module 9
The OpenTelemetry Collector and telemetry pipelines
Coming soon - Module 10
Logs, log pipelines and telemetry privacy
Coming soon - Module 11
Dashboards, debugging, black-box checks and profiling
Coming soon - Module 12
Incident response: roles, mitigation, communication and practice
Coming soon - Module 13
Postmortems and learning from incidents
Coming soon - Module 14
Capacity, load testing and overload
Coming soon - Module 15
Change safety: releases, canaries and delivery metrics
Coming soon - Module 16
Toil, automation and production readiness
Coming soon - Module 17
Resilience testing: chaos engineering and disaster recovery
Coming soon - Module 18
Observability at scale: Kubernetes, cost and AI features
Coming soon - Module 19
SRE interviews and next steps
Coming soon
Official documentation
- opentelemetry.io/docs (opentelemetry.io)
- prometheus.io/docs/introduction/overview (prometheus.io)
- sre.google/sre-book/table-of-contents (sre.google)
More in DevOps and cloud
- Git & GitHub (coming soon)
- Linux Command Line & Administration (coming soon)
- Docker & Containers (coming soon)
- Cloud Computing Fundamentals (coming soon)
- All DevOps & cloud tracks
- How we make lessons