California fault lines: Understanding the causes and impact of network failures

Daniel Turner, Kirill Levchenko, Alex C. Snoeren, Stefan Savage

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

Of the major factors affecting end-to-end service availability, network component failure is perhaps the least well understood. How often do failures occur, how long do they last, what are their causes, and how do they impact customers? Traditionally, answering questions such as these has required dedicated (and often expensive) instrumentation broadly deployed across a network. We propose an alternative approach: opportunistically mining "low-quality" data sources that are already available in modern network environments. We describe a methodology for recreating a succinct history of failure events in an IP network using a combination of structured data (router configurations and syslogs) and semi-structured data (email logs). Using this technique we analyze over five years of failure events in a large regional network consisting of over 200 routers; to our knowledge, this is the largest study of its kind.

Original languageEnglish (US)
Title of host publicationSIGCOMM'10 - Proceedings of the SIGCOMM 2010 Conference
Pages315-326
Number of pages12
DOIs
StatePublished - 2010
Externally publishedYes
Event7th International Conference on Autonomic Computing, SIGCOMM 2010 - New Delhi, India
Duration: Aug 30 2010Sep 3 2010

Publication series

NameSIGCOMM'10 - Proceedings of the SIGCOMM 2010 Conference

Other

Other7th International Conference on Autonomic Computing, SIGCOMM 2010
Country/TerritoryIndia
CityNew Delhi
Period8/30/109/3/10

Keywords

  • failure

ASJC Scopus subject areas

  • Computational Theory and Mathematics
  • Theoretical Computer Science

Fingerprint

Dive into the research topics of 'California fault lines: Understanding the causes and impact of network failures'. Together they form a unique fingerprint.

Cite this