Methods for mining frequent items in data streams: An overview

Hongyan Liu, Yuan Lin, Jiawei Han

Research output: Contribution to journalArticle

Abstract

In many real-world applications, information such as web click data, stock ticker data, sensor network data, phone call records, and traffic monitoring data appear in the form of data streams. Online monitoring of data streams has emerged as an important research undertaking. Estimating the frequency of the items on these streams is an important aggregation and summary technique for both stream mining and data management systems with a broad range of applications. This paper reviews the state-of-the-art progress on methods of identifying frequent items from data streams. It describes different kinds of models for frequent items mining task. For general models such as cash register and Turnstile, we classify existing algorithms into sampling-based, counting-based, and hashing-based categories. The processing techniques and data synopsis structure of each algorithm are described and compared by evaluation measures. Accordingly, as an extension of the general data stream model, four more specific models including time-sensitive model, distributed model, hierarchical and multi-dimensional model, and skewed data model are introduced. The characteristics and limitations of the algorithms of each model are presented, and open issues waiting for study and improvement are discussed.

Original languageEnglish (US)
Pages (from-to)1-30
Number of pages30
JournalKnowledge and Information Systems
Volume26
Issue number1
DOIs
StatePublished - 2011

Keywords

  • Data mining
  • Data stream
  • Frequent items
  • Mining methods and algorithms

ASJC Scopus subject areas

  • Software
  • Information Systems
  • Human-Computer Interaction
  • Hardware and Architecture
  • Artificial Intelligence

Fingerprint Dive into the research topics of 'Methods for mining frequent items in data streams: An overview'. Together they form a unique fingerprint.

  • Cite this