Your country

Tools that support it use your country for local currency, number formats, units and paper size. Your choice is saved only in this browser.

Type a name or a two-letter code. Use the up and down arrow keys to move through the countries, Enter to choose one and Escape to close.

DevOps and cloud

SRE & Observability

The planned outline of 19 modules. Lessons are written, run and reviewed before they are published.

Coming soon 19 modules

Coming soon. This track is being written. Below is its planned outline: a lesson appears here once its examples have been run and its technical review is signed off.

Planned outline

  1. Module 1

    Foundations: what SRE is and how reliability is measured

    Coming soon
  2. Module 2

    Service level objectives: indicators, objectives and error budgets

    Coming soon
  3. Module 3

    Signals: monitoring, observability and metric design

    Coming soon
  4. Module 4

    Prometheus foundations: scraping, instrumenting and PromQL

    Coming soon
  5. Module 5

    Prometheus in depth: advanced PromQL, rules, cardinality and scale

    Coming soon
  6. Module 6

    SLOs in practice: SLI rules, burn-rate alerts and tooling

    Coming soon
  7. Module 7

    Alerting, Alertmanager and sustainable on-call

    Coming soon
  8. Module 8

    Distributed tracing and OpenTelemetry instrumentation

    Coming soon
  9. Module 9

    The OpenTelemetry Collector and telemetry pipelines

    Coming soon
  10. Module 10

    Logs, log pipelines and telemetry privacy

    Coming soon
  11. Module 11

    Dashboards, debugging, black-box checks and profiling

    Coming soon
  12. Module 12

    Incident response: roles, mitigation, communication and practice

    Coming soon
  13. Module 13

    Postmortems and learning from incidents

    Coming soon
  14. Module 14

    Capacity, load testing and overload

    Coming soon
  15. Module 15

    Change safety: releases, canaries and delivery metrics

    Coming soon
  16. Module 16

    Toil, automation and production readiness

    Coming soon
  17. Module 17

    Resilience testing: chaos engineering and disaster recovery

    Coming soon
  18. Module 18

    Observability at scale: Kubernetes, cost and AI features

    Coming soon
  19. Module 19

    SRE interviews and next steps

    Coming soon

Official documentation

More in DevOps and cloud

Quick answers and tool search

Type to search tools or to get a quick answer, for example 18% of 2500. Use the up and down arrow keys to move through the results, Enter to choose, and Escape to close.