TY - GEN
T1 - Zero-shot Faithful Factual Error Correction
AU - Huang, Kung Hsiang
AU - Chan, Hou Pong
AU - Ji, Heng
N1 - This research is based upon work supported by U.S. DARPA SemaFor Program No. HR001120C0123, DARPA AIDA Program No. FA8750-18-2-0014, and DARPA MIPs Program No. HR00112290105. The views and conclusions contained herein are those of the authors and should not be interpreted as necessarily representing the official policies, either expressed or implied, of DARPA, or the U.S. Government. The U.S. Government is authorized to reproduce and distribute reprints for governmental purposes notwithstanding any copyright annotation therein. Hou Pong Chan was supported in part by the Science and Technology Development Fund, Macau SAR (Grant Nos. FDCT/060/2022/AFJ, FDCT/0070/2022/AMJ) and the Multi-year Research Grant from the University of Macau (Grant No. MYRG2020-00054-FST).
PY - 2023
Y1 - 2023
N2 - Faithfully correcting factual errors is critical for maintaining the integrity of textual knowledge bases and preventing hallucinations in generative models. Drawing on humans' ability to identify and correct factual errors, we present a zero-shot framework that formulates questions about input claims, looks for correct answers in the given evidence, and assesses the faithfulness of each correction based on its consistency with the evidence. Our zero-shot framework outperforms fully-supervised approaches, as demonstrated by experiments on the FEVER and SCIFACT datasets, where our outputs are shown to be more faithful. More importantly, the decomposability nature of our framework inherently provides interpretability. Additionally, to reveal the most suitable metrics for evaluating factual error corrections, we analyze the correlation between commonly used metrics with human judgments in terms of three different dimensions regarding intelligibility and faithfulness.
AB - Faithfully correcting factual errors is critical for maintaining the integrity of textual knowledge bases and preventing hallucinations in generative models. Drawing on humans' ability to identify and correct factual errors, we present a zero-shot framework that formulates questions about input claims, looks for correct answers in the given evidence, and assesses the faithfulness of each correction based on its consistency with the evidence. Our zero-shot framework outperforms fully-supervised approaches, as demonstrated by experiments on the FEVER and SCIFACT datasets, where our outputs are shown to be more faithful. More importantly, the decomposability nature of our framework inherently provides interpretability. Additionally, to reveal the most suitable metrics for evaluating factual error corrections, we analyze the correlation between commonly used metrics with human judgments in terms of three different dimensions regarding intelligibility and faithfulness.
UR - https://www.scopus.com/pages/publications/85172430751
UR - https://www.scopus.com/pages/publications/85172430751#tab=citedBy
U2 - 10.18653/v1/2023.acl-long.311
DO - 10.18653/v1/2023.acl-long.311
M3 - Conference contribution
AN - SCOPUS:85172430751
T3 - Proceedings of the Annual Meeting of the Association for Computational Linguistics
SP - 5660
EP - 5676
BT - Long Papers
PB - Association for Computational Linguistics (ACL)
T2 - 61st Annual Meeting of the Association for Computational Linguistics, ACL 2023
Y2 - 9 July 2023 through 14 July 2023
ER -