Site Reliability Engineering (SRE)
About this course
Site Reliability Engineering (SRE) is a discipline created by Google to build and maintain systems at scale. This course teaches you SRE principles and practices to improve the reliability of your services.
Prerequisites
- Monitoring & Logs (previous course)
- DevOps fundamentals
- Kubernetes basics
- Scripting (Python/Bash)
Course content
-
- Origins at Google
- SRE vs DevOps
- Fundamental principles
-
- Service Level Indicators
- Service Level Objectives
- Service Level Agreements
-
- Concept and calculation
- Policies
- Decision making
-
- Definition of toil
- Identification
- Automation strategies
-
- Incident response
- Roles and communication
- Escalation
-
- Blameless culture
- Template and process
- Action items
-
- Forecasting
- Load testing
- Resource management
-
- Deployment strategies
- Feature flags
- Rollbacks
-
- On-call best practices
- Team structure
- SRE maturity
-
- Hands-on labs
- SRE implementation
- Certifications
Estimated duration
⏱️ 16-20 hours of training including hands-on exercises
Learning objectives
By the end of this course, you will be able to:
- Define relevant SLIs and SLOs
- Manage an Error Budget
- Identify and eliminate Toil
- Conduct blameless Post-mortems
- Plan system capacity
- Implement Release Engineering practices