构建真实数值事实核查数据集,模拟人类拆解与检索证据的流程。
A Benchmark for Open-Domain Numerical Fact-Checking Enhanced by Claim Decomposition
- 通过模拟人类拆解数值声明的方式收集证据,提升真实性。
- 数据集无时间泄漏,包含自然数值主张与对应相关证据。
- 适合研究自动事实核查、信息检索及可解释验证的学者使用。
数值类事实核查至关重要,因数字常给人以真实感,但虚假数字可能对社会造成灾难性影响。现有自动事实验证研究较少聚焦自然语言中的数值主张。人类核查者通常先检索与主张各数值维度相关的证据,再进行推理判断。因此,检索过程是验证的基础技能。现有基准采用启发式拆解和弱监督网络搜索获取证据,常导致证据不相关、来源噪声大且存在时间泄露,难以模拟真实场景。为此,我们提出 QuanTemp++:一个由自然数值主张、开放域语料库及对应相关证据构成的数据集。证据通过近似人类核查者方式的主张分解流程收集,并确保无时间泄露。基于该数据集,我们评估了关键主张分解范式的检索性能,并分析其对验证流水线结果的影响,得出实用洞察。数据管道代码及数据链接见 https://github.com/VenkteshV/QuanTemp_Plus。
原文摘要 · Abstract (English)
Fact-checking numerical claims is critical as the presence of numbers provide mirage of veracity despite being fake potentially causing catastrophic impacts on society. The prior works in automatic fact verification do not primarily focus on natural numerical claims. A typical human fact-checker first retrieves relevant evidence addressing the different numerical aspects of the claim and then reasons about them to predict the veracity of the claim. Hence, the search process of a human fact-checker is a crucial skill that forms the foundation of the verification process. Emulating a real-world setting is essential to aid in the development of automated methods that encompass such skills. However, existing benchmarks employ heuristic claim decomposition approaches augmented with weakly supervised web search to collect evidences for verifying claims. This sometimes results in less relevant evidences and noisy sources with temporal leakage rendering a less realistic retrieval setting for claim verification. Hence, we introduce QuanTemp++: a dataset consisting of natural numerical claims, an open domain corpus, with the corresponding relevant evidence for each claim. The evidences are collected through a claim decomposition process approximately emulating the approach of human fact-checker and veracity labels ensuring there is no temporal leakage. Given this dataset, we also characterize the retrieval performance of key claim decomposition paradigms. Finally, we observe their effect on the outcome of the verification pipeline and draw insights. The code for data pipeline along with link to data can be found at https://github.com/VenkteshV/QuanTemp_Plus
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。