A large scale study of data center network reliability

Justin Meza, Kaushik Veeraraghavan, Tianyin Xu, Onur Mutlu

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

The ability to tolerate, remediate, and recover from network incidents (caused by device failures and fiber cuts, for example) is critical for building and operating highly-available web services. Achieving fault tolerance and failure preparedness requires system architects, software developers, and site operators to have a deep understanding of network reliability at scale, along with its implications on the software systems that run in data centers. Unfortunately, little has been reported on the reliability characteristics of large scale data center network infrastructure, let alone its impact on the availability of services powered by software running on that network infrastructure. This paper fills the gap by presenting a large scale, longitudinal study of data center network reliability based on operational data collected from the production network infrastructure at Facebook, one of the largest web service providers in the world. Our study covers reliability characteristics of both intra and inter data center networks. For intra data center networks, we study seven years of operation data comprising thousands of network incidents across two different data center network designs, a cluster network design and a state-of-the-art fabric network design. For inter data center networks, we study eighteen months of recent repair tickets from the field to understand reliability of Wide Area Network (WAN) backbones. In contrast to prior work, we study the effects of network reliability on software systems, and how these reliability characteristics evolve over time. We discuss the implications of network reliability on the design, implementation, and operation of large scale data center systems and how it affects highly-available web services. We hope our study forms a foundation for understanding the reliability of large scale network infrastructure, and inspires new reliability solutions to network incidents.

Original languageEnglish (US)
Title of host publicationIMC 2018 - Proceedings of the Internet Measurement Conference
PublisherAssociation for Computing Machinery
Pages393-407
Number of pages15
ISBN (Electronic)9781450356190
DOIs
StatePublished - Oct 31 2018
Event2018 Internet Measurement Conference, IMC 2018 - Boston, United States
Duration: Oct 31 2018Nov 2 2018

Publication series

NameProceedings of the ACM SIGCOMM Internet Measurement Conference, IMC

Other

Other2018 Internet Measurement Conference, IMC 2018
Country/TerritoryUnited States
CityBoston
Period10/31/1811/2/18

Keywords

  • Data centers
  • Fault tolerance
  • Networks
  • Reliability

ASJC Scopus subject areas

  • Software
  • Computer Networks and Communications

Fingerprint

Dive into the research topics of 'A large scale study of data center network reliability'. Together they form a unique fingerprint.

Cite this