提出新评估方法,让自动查证更准更稳。
Ev2R: Evaluating Evidence Retrieval in Automated Fact-Checking
- 融合参考答案与结论预测,双重评估证据质量。
- 相比旧方法,与人工判断相关性更高,抗干扰更强。
- 适合研究自动查证、证据评估的学者和开发者。
当前自动查证方法通常通过预测结论隐式评估证据,或依赖维基百科等封闭知识源进行精确匹配,但这些方法受限于非专门设计的评估指标和知识源限制。本文提出Ev²R,结合基于参考的标准评估与基于结论的代理评分,联合评估证据与真实参考的一致性及其对结论的支持可靠性,克服了以往方法的不足。在三类证据评估方法(基于参考、代理参考、无参考)上对比测试,结果表明Ev²R在准确性和鲁棒性方面均优于现有方法,与人工评分相关性更强,对抗性扰动更具韧性,可作为自动查证中证据评估的可靠基准。
原文摘要 · Abstract (English)
Current automated fact-checking (AFC) approaches typically evaluate evidence either implicitly via the predicted verdicts or through exact matches with predefined closed knowledge sources, such as Wikipedia. However, these methods are limited due to their reliance on evaluation metrics originally designed for other purposes and constraints from closed knowledge sources. In this work, we introduce \textbf{\textcolor{skyblue}{Ev\textsuperscript{2}}\textcolor{orangebrown}{R}} which combines the strengths of reference-based evaluation and verdict-level proxy scoring. Ev\textsuperscript{2}R jointly assesses how well the evidence aligns with the gold references and how reliably it supports the verdict, addressing the shortcomings of prior methods. We evaluate Ev\textsuperscript{2}R against three types of evidence evaluation approaches: reference-based, proxy-reference, and reference-less baselines. Assessments against human ratings and adversarial tests demonstrate that Ev\textsuperscript{2}R consistently outperforms existing scoring approaches in accuracy and robustness. It achieves stronger correlation with human judgments and greater robustness to adversarial perturbations, establishing it as a reliable metric for evidence evaluation in AFC.\footnote{Code is available at \href{https://github.com/mubasharaak/fc-evidence-evaluation}{https://github.com/mubasharaak/fc-evidence-evaluation}.}
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。