Site Reliability Engineering
470 items from dastergon/awesome-sre ★13,713
-
Open Guide to AWS :fire::fire::fire::fire::fire: github.com
Links to the Billing and Cost Management section which details the broad characteristics of billing for a cloud provider.
-
A collection of post-mortems github.com
-
Collection of Kubernetes Failure Stories github.com
-
A collection of postmortem templates github.com
-
Run Book / Operations Manual template github.com
-
The On-Call Handbook github.com
-
-
-
-
-
Weathering the Unexpected queue.acm.org
-
Incident Response at Heroku blog.heroku.com
-
Service Level Disagreements Part I blog.b3k.us
-
Load Testing & Black Friday capacity planning medium.com
How Back Market prepared for Black Friday with k6 based load testing.
-
Chaos Engineering: Crash test your applications manning.com
A book on how to design and execute controlled software failure experiments.
-
Blameless PostMortems and a Just Culture codeascraft.com
📰 - Etsy's Code as Craft blog discusses how they look at mistakes with a perspective of learning through blameless post-mortems.
-
SRE fundamentals: SLIs, SLAs and SLOs cloudplatform.googleblog.com
If you are in the business of cloud services, these metrics are certainly great KPIs.
-
PagerDuty Incident Response Documentation response.pagerduty.com
Documents that describe parts of the PagerDuty Incident Response process. It provides information not only on preparing for an incident, but also what to do during and after. Source is available on GitHub.
-
-
Resilience Roundup resilienceroundup.com
Weekly analysis of Resilience Engineering and Human Factors research designed for software systems
-
Monitoring Weekly monitoring.love
What's new in monitoring? Curated monitoring articles to your inbox each week.
-
Things I Learned Managing Site Reliability for Some of the World’s Busiest Gambling Sites zwischenzugs.wordpress.com
-
-
-
-
-
Time To Detect - Netflix youtube.com
-
-
About SRE and how (not) to apply it youtube.com
-
- next page of items loading…