SRE / DevOps / Platform · 300-level

SRE & Platform Engineering

A staff-level deep dive into running reliable systems at scale: SLOs and error budgets, incident response, observability, CI/CD and progressive delivery, infrastructure as code, Kubernetes operations, resilience engineering, and the platform practices that make the reliable path the default path.

SRE-3013 creditselectiveno prerequisites
What's inside

Sections & lessons

01

Reliability foundations

  • SLIs: measuring the thing users actually feelconcept35 min
  • SLOs and the math of ninesconcept40 min
  • Spending the error budget: policy over heroicsconcept35 min
  • The four golden signals and the war on toilconcept35 min
02

Incident response & on-call

  • Severity levels that trigger the right responseconcept35 min
  • Incident command: roles that hold under pressureconcept40 min
  • Communicating during an incidentconcept35 min
  • Blameless postmortems and sustainable on-callconcept40 min
03

Observability

  • Metrics that answer questions: RED and USEconcept40 min
  • Structured logs, correlation, and the cost of writing everything downconcept35 min
  • Distributed tracing with OpenTelemetrydemo45 min
  • Cardinality, alert fatigue, and paging on symptomsconcept35 min
04

CI/CD pipelines

  • Anatomy of a deployment pipelineconcept40 min
  • Progressive delivery: canary and blue-greendemo45 min
  • Rollbacks: the fastest path back to known-goodconcept35 min
  • Decoupling deploy from release: flags and safe migrationsconcept40 min
05

Infrastructure as Code

  • Declarative infrastructure and the desired-state modelconcept35 min
  • Terraform: modules, state, and blast radiusconcept40 min
  • Drift: when reality diverges from codeconcept35 min
  • Immutable infrastructure and golden imagesconcept40 min
06

Containers & Kubernetes

  • Pods, Deployments, and Services: the core objectsconcept40 min
  • Resource requests, limits, and QoSconcept40 min
  • Health probes and autoscalingconcept40 min
  • Debugging the common Kubernetes failure modesconcept35 min
07

Reliability engineering

  • Capacity planning and headroomconcept40 min
  • Load testing and chaos engineeringdemo40 min
  • Backpressure, graceful degradation, and disaster recoveryconcept40 min
08

Platform engineering

  • Golden paths and paved roadsconcept40 min
  • Internal developer platforms and self-serviceconcept40 min
  • Measuring developer experience and platform successreflection35 min

This module ends in a gate you can fail.

That's what makes passing it mean something. Take the DSAT, get placed, and start earning.