让AI自解释为何不确定,基于证据冲突与一致性的自然语言说明。
Explaining Sources of Uncertainty in Automated Fact-Checking
- 通过无监督识别文本片段间的证据冲突与一致,定位不确定性来源。
- 生成解释后,人类评估认为更可信、更少冗余且逻辑更连贯。
- 无需微调或改架构,可直接用于任何白盒语言模型。
理解模型预测时的不确定性来源对人机协作至关重要。现有方法仅依赖数值不确定度或模糊表达,无法解释由矛盾证据引发的不确定性,导致用户难以解决分歧或信任输出。我们提出CLUE(冲突与一致性感知的语言模型不确定性解释框架),首次通过无监督方式识别文本跨度间揭示主张-证据或证据间冲突与一致的关系,从而驱动模型的预测不确定性;并利用提示工程与注意力引导生成自然语言解释,明确描述这些关键互动。在三个语言模型和两个事实核查数据集上,结果表明,相比无跨度交互指导的提示,CLUE生成的解释更忠实于模型不确定性,且更符合事实核查结论。人类评估显示,我们的解释更具帮助性、信息量更高、冗余更少、逻辑更一致。该方法无需微调或结构修改,可即插即用适配任意白盒语言模型。通过显式关联不确定性与证据冲突,为事实核查提供实用支持,并可推广至需复杂信息推理的其他任务。
原文摘要 · Abstract (English)
Understanding sources of a model's uncertainty regarding its predictions is crucial for effective human-AI collaboration. Prior work proposes using numerical uncertainty or hedges ("I'm not sure, but ..."), which do not explain uncertainty that arises from conflicting evidence, leaving users unable to resolve disagreements or rely on the output. We introduce CLUE (Conflict-and-Agreement-aware Language-model Uncertainty Explanations), the first framework to generate natural language explanations of model uncertainty by (i) identifying relationships between spans of text that expose claim-evidence or inter-evidence conflicts and agreements that drive the model's predictive uncertainty in an unsupervised way, and (ii) generating explanations via prompting and attention steering that verbalize these critical interactions. Across three language models and two fact-checking datasets, we show that CLUE produces explanations that are more faithful to the model's uncertainty and more consistent with fact-checking decisions than prompting for uncertainty explanations without span-interaction guidance. Human evaluators judge our explanations to be more helpful, more informative, less redundant, and more logically consistent with the input than this baseline. CLUE requires no fine-tuning or architectural changes, making it plug-and-play for any white-box language model. By explicitly linking uncertainty to evidence conflicts, it offers practical support for fact-checking and generalises readily to other tasks that require reasoning over complex information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。