Site Reliability Engineering
Betsy Beyer, Chris Jones, Jennifer Petoff, Niall Richard Murphy (eds.)
Google's account of how it runs large-scale production systems reliably: monitoring, capacity planning, incident response, postmortems, release engineering, and the SRE approach to managing risk at scale.
DevOpsAdvancedOfficial Free DistributionHTML