TY - GEN
T1 - A semi-supervised active-learning truth estimator for social networks
AU - Cui, Hang
AU - Abdelzaher, Tarek
AU - Kaplan, Lance
N1 - Research reported in this paper was sponsored in part by the U.S. Army Research Laboratory under Cooperative Agreements W911NF-17-2-0196 and W911NF-09-2-0053, DARPA contract W911NF-17-C-0099, and NSF grants CNS 13-29886 and CNS 16-18627. The views and conclusions contained in this document are those of the authors and should not be interpreted as representing the official policies, either expressed or implied, of the Army Research Laboratory or the U.S. Government. The U.S. Government is authorized to reproduce and distribute reprints for Government purposes notwithstanding any copyright notation here on.
PY - 2019/5/13
Y1 - 2019/5/13
N2 - This paper introduces an active-learning-based truth estimator for social networks, such as Twitter, that enhances estimation accuracy significantly by requesting a well-selected (small) fraction of data to be labeled. Data assessment and truth discovery from arbitrary open online sources are a hard problem due to uncertainty regarding source reliability. Multiple truth finding systems were developed to solve this problem. Their accuracy is limited by the noisy nature of the data, where distortions, fabrications, omissions, and duplication are introduced. This paper presents a semi-supervised truth estimator for social networks, in which a portion of inputs are carefully selected to be reliably verified. The challenge is to find the subset of observations to verify that would maximally enhance the overall fact-finding accuracy. This work extends previous passive approaches to recursive truth estimation, as well as semi-supervised approaches where the estimator has no control over the choice of data to be labeled. Results show that by optimally selecting claims to be verified, we improve estimated accuracy by 12% over unsupervised baseline, and by 5% over previous semi-supervised approaches.
AB - This paper introduces an active-learning-based truth estimator for social networks, such as Twitter, that enhances estimation accuracy significantly by requesting a well-selected (small) fraction of data to be labeled. Data assessment and truth discovery from arbitrary open online sources are a hard problem due to uncertainty regarding source reliability. Multiple truth finding systems were developed to solve this problem. Their accuracy is limited by the noisy nature of the data, where distortions, fabrications, omissions, and duplication are introduced. This paper presents a semi-supervised truth estimator for social networks, in which a portion of inputs are carefully selected to be reliably verified. The challenge is to find the subset of observations to verify that would maximally enhance the overall fact-finding accuracy. This work extends previous passive approaches to recursive truth estimation, as well as semi-supervised approaches where the estimator has no control over the choice of data to be labeled. Results show that by optimally selecting claims to be verified, we improve estimated accuracy by 12% over unsupervised baseline, and by 5% over previous semi-supervised approaches.
KW - Active Learning
KW - Maximum Likelihood Estimation
KW - Semi Supervision
KW - Social Sensing
KW - Truth Discovery
UR - https://www.scopus.com/pages/publications/85066905643
UR - https://www.scopus.com/pages/publications/85066905643#tab=citedBy
U2 - 10.1145/3308558.3313712
DO - 10.1145/3308558.3313712
M3 - Conference contribution
AN - SCOPUS:85066905643
T3 - The Web Conference 2019 - Proceedings of the World Wide Web Conference, WWW 2019
SP - 296
EP - 306
BT - The Web Conference 2019 - Proceedings of the World Wide Web Conference, WWW 2019
PB - Association for Computing Machinery
T2 - 2019 World Wide Web Conference, WWW 2019
Y2 - 13 May 2019 through 17 May 2019
ER -