Skip to main content

Site Reliability Engineering (SRE)


About this course

Site Reliability Engineering (SRE) is a discipline created by Google to build and maintain systems at scale. This course teaches you SRE principles and practices to improve the reliability of your services.


Prerequisites

  • Monitoring & Logs (previous course)
  • DevOps fundamentals
  • Kubernetes basics
  • Scripting (Python/Bash)

Course content

  1. Introduction to SRE

    • Origins at Google
    • SRE vs DevOps
    • Fundamental principles
  2. SLIs, SLOs, and SLAs

    • Service Level Indicators
    • Service Level Objectives
    • Service Level Agreements
  3. Error Budget

    • Concept and calculation
    • Policies
    • Decision making
  4. Toil Elimination

    • Definition of toil
    • Identification
    • Automation strategies
  5. Incident Management

    • Incident response
    • Roles and communication
    • Escalation
  6. Post-mortems

    • Blameless culture
    • Template and process
    • Action items
  7. Capacity Planning

    • Forecasting
    • Load testing
    • Resource management
  8. Release Engineering

    • Deployment strategies
    • Feature flags
    • Rollbacks
  9. Best practices

    • On-call best practices
    • Team structure
    • SRE maturity
  10. Exercises and Projects

    • Hands-on labs
    • SRE implementation
    • Certifications

Estimated duration

⏱️ 16-20 hours of training including hands-on exercises


Learning objectives

By the end of this course, you will be able to:

  • Define relevant SLIs and SLOs
  • Manage an Error Budget
  • Identify and eliminate Toil
  • Conduct blameless Post-mortems
  • Plan system capacity
  • Implement Release Engineering practices