SRE / DevOps / Platform · 300-level
SRE & Platform Engineering
A staff-level deep dive into running reliable systems at scale: SLOs and error budgets, incident response, observability, CI/CD and progressive delivery, infrastructure as code, Kubernetes operations, resilience engineering, and the platform practices that make the reliable path the default path.
SRE-3013 creditselectiveno prerequisites
What's inside
Sections & lessons
01
Reliability foundations
- SLIs: measuring the thing users actually feelconcept35 min
- SLOs and the math of ninesconcept40 min
- Spending the error budget: policy over heroicsconcept35 min
- The four golden signals and the war on toilconcept35 min
02
Incident response & on-call
- Severity levels that trigger the right responseconcept35 min
- Incident command: roles that hold under pressureconcept40 min
- Communicating during an incidentconcept35 min
- Blameless postmortems and sustainable on-callconcept40 min
03
Observability
- Metrics that answer questions: RED and USEconcept40 min
- Structured logs, correlation, and the cost of writing everything downconcept35 min
- Distributed tracing with OpenTelemetrydemo45 min
- Cardinality, alert fatigue, and paging on symptomsconcept35 min
04
CI/CD pipelines
- Anatomy of a deployment pipelineconcept40 min
- Progressive delivery: canary and blue-greendemo45 min
- Rollbacks: the fastest path back to known-goodconcept35 min
- Decoupling deploy from release: flags and safe migrationsconcept40 min
05
Infrastructure as Code
- Declarative infrastructure and the desired-state modelconcept35 min
- Terraform: modules, state, and blast radiusconcept40 min
- Drift: when reality diverges from codeconcept35 min
- Immutable infrastructure and golden imagesconcept40 min
06
Containers & Kubernetes
- Pods, Deployments, and Services: the core objectsconcept40 min
- Resource requests, limits, and QoSconcept40 min
- Health probes and autoscalingconcept40 min
- Debugging the common Kubernetes failure modesconcept35 min
07
Reliability engineering
- Capacity planning and headroomconcept40 min
- Load testing and chaos engineeringdemo40 min
- Backpressure, graceful degradation, and disaster recoveryconcept40 min
08
Platform engineering
- Golden paths and paved roadsconcept40 min
- Internal developer platforms and self-serviceconcept40 min
- Measuring developer experience and platform successreflection35 min
This module ends in a gate you can fail.
That's what makes passing it mean something. Take the DSAT, get placed, and start earning.