arXiv:2504.05229cs.AI2025-04EMNLP被引 2

提出细粒度评估框架,量化可行动性解释效果

FinGrAct: A Framework for FINe-GRrained Evaluation of ACTionability in Explainable Automatic Fact-Checking

  • 构建能访问网络的评估框架,通过明确标准衡量解释的可行动性
  • 在人类判断相关性上表现最优,皮尔逊与肯德尔相关系数最高
  • 适合研究可信自动事实核查的学者和开发可解释系统的技术人员

可解释的自动事实核查(AFC)旨在通过清晰易懂的解释提升自动化验证系统的透明度与可信度。然而,解释的有效性取决于其可行动性——即是否能帮助用户做出明智决策并遏制虚假信息传播。尽管可行动性是高质量解释的关键属性,但此前尚无专门方法对其进行评估。本文提出FinGrAct,一个可访问网络的细粒度评估框架,通过定义明确的标准和专用数据集,实现对AFC解释可行动性的量化评估。FinGrAct在与人类判断的相关性上超越现有最先进评估器,皮尔逊相关系数与肯德尔等级相关系数均达最高,同时表现出最低的自我中心偏差,展现出更强的评估鲁棒性。

原文摘要 · Abstract (English)

The field of explainable Automatic Fact-Checking (AFC) aims to enhance the transparency and trustworthiness of automated fact-verification systems by providing clear and comprehensible explanations. However, the effectiveness of these explanations depends on their actionability --their ability to empower users to make informed decisions and mitigate misinformation. Despite actionability being a critical property of high-quality explanations, no prior research has proposed a dedicated method to evaluate it. This paper introduces FinGrAct, a fine-grained evaluation framework that can access the web, and it is designed to assess actionability in AFC explanations through well-defined criteria and an evaluation dataset. FinGrAct surpasses state-of-the-art (SOTA) evaluators, achieving the highest Pearson and Kendall correlation with human judgments while demonstrating the lowest ego-centric bias, making it a more robust evaluation approach for actionability evaluation in AFC.

可解释性事实核查评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。