California fault lines: Understanding the causes and impact of network failures

Daniel Turner, Kirill Levchenko, Alex C. Snoeren, Stefan Savage

Research output: Contribution to journalArticle

Abstract

Of the major factors affecting end-to-end service availability, network component failure is perhaps the least well understood. How often do failures occur, how long do they last, what are their causes, and how do they impact customers? Traditionally, answering questions such as these has required dedicated (and often expensive) instrumentation broadly deployed across a network. We propose an alternative approach: opportunistically mining "low-quality" data sources that are already available in modern network environments. We describe a methodology for recreating a succinct history of failure events in an IP network using a combination of structured data (router configurations and syslogs) and semi-structured data (email logs). Using this technique we analyze over five years of failure events in a large regional network consisting of over 200 routers; to our knowledge, this is the largest study of its kind.

Original languageEnglish (US)
Pages (from-to)315-326
Number of pages12
JournalComputer Communication Review
Volume40
Issue number4
DOIs
StatePublished - Dec 1 2010

Fingerprint

Routers
Network components
Electronic mail
Availability

Keywords

  • C.2.3 [computer-communication network]: Network operations
  • Measurement
  • Reliability

ASJC Scopus subject areas

  • Software
  • Computer Networks and Communications

Cite this

California fault lines : Understanding the causes and impact of network failures. / Turner, Daniel; Levchenko, Kirill; Snoeren, Alex C.; Savage, Stefan.

In: Computer Communication Review, Vol. 40, No. 4, 01.12.2010, p. 315-326.

Research output: Contribution to journalArticle

Turner, Daniel ; Levchenko, Kirill ; Snoeren, Alex C. ; Savage, Stefan. / California fault lines : Understanding the causes and impact of network failures. In: Computer Communication Review. 2010 ; Vol. 40, No. 4. pp. 315-326.
@article{8ae9ed338f094e60a2d431629368d531,
title = "California fault lines: Understanding the causes and impact of network failures",
abstract = "Of the major factors affecting end-to-end service availability, network component failure is perhaps the least well understood. How often do failures occur, how long do they last, what are their causes, and how do they impact customers? Traditionally, answering questions such as these has required dedicated (and often expensive) instrumentation broadly deployed across a network. We propose an alternative approach: opportunistically mining {"}low-quality{"} data sources that are already available in modern network environments. We describe a methodology for recreating a succinct history of failure events in an IP network using a combination of structured data (router configurations and syslogs) and semi-structured data (email logs). Using this technique we analyze over five years of failure events in a large regional network consisting of over 200 routers; to our knowledge, this is the largest study of its kind.",
keywords = "C.2.3 [computer-communication network]: Network operations, Measurement, Reliability",
author = "Daniel Turner and Kirill Levchenko and Snoeren, {Alex C.} and Stefan Savage",
year = "2010",
month = "12",
day = "1",
doi = "10.1145/1851275.1851220",
language = "English (US)",
volume = "40",
pages = "315--326",
journal = "Computer Communication Review",
issn = "0146-4833",
publisher = "Association for Computing Machinery (ACM)",
number = "4",

}

TY - JOUR

T1 - California fault lines

T2 - Understanding the causes and impact of network failures

AU - Turner, Daniel

AU - Levchenko, Kirill

AU - Snoeren, Alex C.

AU - Savage, Stefan

PY - 2010/12/1

Y1 - 2010/12/1

N2 - Of the major factors affecting end-to-end service availability, network component failure is perhaps the least well understood. How often do failures occur, how long do they last, what are their causes, and how do they impact customers? Traditionally, answering questions such as these has required dedicated (and often expensive) instrumentation broadly deployed across a network. We propose an alternative approach: opportunistically mining "low-quality" data sources that are already available in modern network environments. We describe a methodology for recreating a succinct history of failure events in an IP network using a combination of structured data (router configurations and syslogs) and semi-structured data (email logs). Using this technique we analyze over five years of failure events in a large regional network consisting of over 200 routers; to our knowledge, this is the largest study of its kind.

AB - Of the major factors affecting end-to-end service availability, network component failure is perhaps the least well understood. How often do failures occur, how long do they last, what are their causes, and how do they impact customers? Traditionally, answering questions such as these has required dedicated (and often expensive) instrumentation broadly deployed across a network. We propose an alternative approach: opportunistically mining "low-quality" data sources that are already available in modern network environments. We describe a methodology for recreating a succinct history of failure events in an IP network using a combination of structured data (router configurations and syslogs) and semi-structured data (email logs). Using this technique we analyze over five years of failure events in a large regional network consisting of over 200 routers; to our knowledge, this is the largest study of its kind.

KW - C.2.3 [computer-communication network]: Network operations

KW - Measurement

KW - Reliability

UR - http://www.scopus.com/inward/record.url?scp=84874697428&partnerID=8YFLogxK

UR - http://www.scopus.com/inward/citedby.url?scp=84874697428&partnerID=8YFLogxK

U2 - 10.1145/1851275.1851220

DO - 10.1145/1851275.1851220

M3 - Article

AN - SCOPUS:84874697428

VL - 40

SP - 315

EP - 326

JO - Computer Communication Review

JF - Computer Communication Review

SN - 0146-4833

IS - 4

ER -