Site Reliability Engineering: How Google Runs Production Systems
暫譯: Site Reliability Engineering:Google 如何運行生產系統
Beyer, Betsy, Jones, Chris, Leng, Christof
商品描述
Google pioneered the discipline of Site Reliability Engineering, applying reliability to the entire user journey for consumer, enterprise, and infrastructure systems. In the years since, many organizations have followed suit, guided by the tenets laid out in this practical book. This fully revised edition brings Site Reliability Engineering up-to-date with fresh insights on engineering techniques, organizational processes, and case studies that will help you promote and implement greater reliability throughout the engineering lifecycle.
In this collection of essays and articles, key members of Google's Site Reliability Engineering team explore the company's current SRE practices and explain how they've evolved in the decade since the initial publication. New updates cover the value of reliability, cloud reliability, and the impact of AI. You'll learn the principles and practices that enable Google engineers to make some of the world's largest systems scalable, reliable, and efficient--lessons directly applicable to your organization.
- Train new Site Reliability Engineers based on the latest practices in the field
- Develop engineering organizations that support reliability as a feature
- Build online services that incorporate reliability principles
- Use AI to improve SRE across the organization and optimize critical areas such as automation and incident detection
商品描述(中文翻譯)
Google 首創網站可靠性工程(Site Reliability Engineering,SRE)這門學科,將可靠性應用於消費者、企業及基礎架構系統的完整使用者旅程。此後多年,許多組織紛紛效法 Google,並以本實用著作所闡述的原則為指引。全新修訂版的《Site Reliability Engineering》融入了工程技術、組織流程及案例研究方面的最新洞見,協助您在整個工程生命週期中推動並實踐更高程度的可靠性。
本書收錄一系列由 Google Site Reliability Engineering 團隊核心成員撰寫的論文與文章,探討 Google 目前的 SRE 實務,並說明這些實務在初版出版後十年間如何持續演進。新增內容涵蓋可靠性的價值、雲端可靠性,以及 AI 所帶來的影響。您將學習讓 Google 工程師得以打造全球最大型系統之一,使其具備可擴展性、可靠性與高效率的原則與實務;這些經驗同樣能直接應用於您的組織。
• 根據領域中的最新實務,培訓新進 Site Reliability Engineer
• 建立將可靠性視為功能的工程組織
• 打造融入可靠性原則的線上服務
• 運用 AI 在整個組織中改善 SRE,並最佳化自動化與事件偵測等關鍵領域