dogear

enter for all results · esc to close

Site Reliability Engineering

470 items from dastergon/awesome-sre ★13,713

  1. Open Guide to AWS :fire::fire::fire::fire::fire: github.com

    Links to the Billing and Cost Management section which details the broad characteristics of billing for a cloud provider.

  2. A collection of post-mortems github.com
  3. Collection of Kubernetes Failure Stories github.com
  4. A collection of postmortem templates github.com
  5. Run Book / Operations Manual template github.com
  6. The On-Call Handbook github.com
  7. SRE cheat sheet github.com

    A cheat sheet for Site Reliability Engineering principles and numbers

  8. Extended Dreyfus Model for Incident Lifecycles github.com
  9. https://rachelbythebay.com/w/ rachelbythebay.com

    Techincal Blog Posts.

  10. http://highscalability.com/ highscalability.com

    Technical Blog Posts About Systems Architecture.

  11. Weathering the Unexpected queue.acm.org
  12. Incident Response at Heroku blog.heroku.com
  13. Service Level Disagreements Part I blog.b3k.us
  14. Load Testing & Black Friday capacity planning medium.com

    How Back Market prepared for Black Friday with k6 based load testing.

  15. Chaos Engineering: Crash test your applications manning.com

    A book on how to design and execute controlled software failure experiments.

  16. Blameless PostMortems and a Just Culture codeascraft.com

    📰 - Etsy's Code as Craft blog discusses how they look at mistakes with a perspective of learning through blameless post-mortems.

  17. SRE fundamentals: SLIs, SLAs and SLOs cloudplatform.googleblog.com

    If you are in the business of cloud services, these metrics are certainly great KPIs.

  18. PagerDuty Incident Response Documentation response.pagerduty.com

    Documents that describe parts of the PagerDuty Incident Response process. It provides information not only on preparing for an incident, but also what to do during and after. Source is available on GitHub.

  19. SRE Weekly sreweekly.com

    Weekly Site Reliability Newsletter.

  20. Resilience Roundup resilienceroundup.com

    Weekly analysis of Resilience Engineering and Human Factors research designed for software systems

  21. Monitoring Weekly monitoring.love

    What's new in monitoring? Curated monitoring articles to your inbox each week.

  22. Things I Learned Managing Site Reliability for Some of the World’s Busiest Gambling Sites zwischenzugs.wordpress.com
  23. SBSRE Meetup: Different SRE roles and challenges(Netflix) youtube.com
  24. Site Reliability Engineers — Keeping Google up and running 24/7 youtube.com
  25. "Practical Applications of the Dickerson Pyramid" by Nat Welch youtube.com
  26. How Your Systems Keep Running Day After Day - John Allspaw youtube.com
  27. Time To Detect - Netflix youtube.com
  28. Embracing Failure: Fault-Injection and Service Reliability youtube.com
  29. About SRE and how (not) to apply it youtube.com
  30. South Bay SRE Meetup - Netflix Cloud Performance Team youtube.com
  31. next page of items loading…