TY - GEN
T1 - Integrating local context and global cohesiveness for open information extraction
AU - Zhu, Qi
AU - Ren, Xiang
AU - Shang, Jingbo
AU - Zhang, Yu
AU - El-Kishky, Ahmed
AU - Han, Jiawei
N1 - Research was sponsored in part by the U.S. Army Research Lab. under Cooperative Agreement No. W911NF-09-2-0053 (NSCTA), National Science Foundation IIS 16-18481, IIS 17-04532, and IIS-17-41317, and grant 1U54GM114838 awarded by NIGMS through funds provided by the trans-NIH Big Data to Knowledge (BD2K) initiative (www.bd2k.nih.gov). Xiang Ren's research has been supported in part by National Science Foundation SMA 18-29268. We thank Frank F. Xu and Ellen Wu for valuable feedback and discussions.
PY - 2019/1/30
Y1 - 2019/1/30
N2 - Extracting entities and their relations from text is an important task for understanding massive text corpora. Open information extraction (IE) systems mine relation tuples (i.e., entity arguments and a predicate string to describe their relation) from sentences. These relation tuples are not confined to a predefined schema for the relations of interests. However, current Open IE systems focus on modeling local context information in a sentence to extract relation tuples, while ignoring the fact that global statistics in a large corpus can be collectively leveraged to identify high-quality sentence-level extractions. In this paper, we propose a novel Open IE system, called ReMine, which integrates local context signals and global structural signals in a unified, distant-supervision framework. Leveraging facts from external knowledge bases as supervision, the new system can be applied to many different domains to facilitate sentence-level tuple extractions using corpus-level statistics. Our system operates by solving a joint optimization problem to unify (1) segmenting entity/relation phrases in individual sentences based on local context; and (2) measuring the quality of tuples extracted from individual sentences with a translating-based objective. Learning the two subtasks jointly helps correct errors produced in each subtask so that they can mutually enhance each other. Experiments on two real-world corpora from different domains demonstrate the effectiveness, generality, and robustness of ReMine when compared to state-of-the-art open IE systems.
AB - Extracting entities and their relations from text is an important task for understanding massive text corpora. Open information extraction (IE) systems mine relation tuples (i.e., entity arguments and a predicate string to describe their relation) from sentences. These relation tuples are not confined to a predefined schema for the relations of interests. However, current Open IE systems focus on modeling local context information in a sentence to extract relation tuples, while ignoring the fact that global statistics in a large corpus can be collectively leveraged to identify high-quality sentence-level extractions. In this paper, we propose a novel Open IE system, called ReMine, which integrates local context signals and global structural signals in a unified, distant-supervision framework. Leveraging facts from external knowledge bases as supervision, the new system can be applied to many different domains to facilitate sentence-level tuple extractions using corpus-level statistics. Our system operates by solving a joint optimization problem to unify (1) segmenting entity/relation phrases in individual sentences based on local context; and (2) measuring the quality of tuples extracted from individual sentences with a translating-based objective. Learning the two subtasks jointly helps correct errors produced in each subtask so that they can mutually enhance each other. Experiments on two real-world corpora from different domains demonstrate the effectiveness, generality, and robustness of ReMine when compared to state-of-the-art open IE systems.
UR - https://www.scopus.com/pages/publications/85061742382
UR - https://www.scopus.com/pages/publications/85061742382#tab=citedBy
U2 - 10.1145/3289600.3291030
DO - 10.1145/3289600.3291030
M3 - Conference contribution
AN - SCOPUS:85061742382
T3 - WSDM 2019 - Proceedings of the 12th ACM International Conference on Web Search and Data Mining
SP - 42
EP - 50
BT - WSDM 2019 - Proceedings of the 12th ACM International Conference on Web Search and Data Mining
PB - Association for Computing Machinery
T2 - 12th ACM International Conference on Web Search and Data Mining, WSDM 2019
Y2 - 11 February 2019 through 15 February 2019
ER -