arXiv:2410.04068cs.CLcs.AI2024-10EMNLP被引 24

提出检测与解决证据冲突的新方法,助力AI系统应对虚假信息。

ECon: On the Detection and Resolution of Evidence Conflicts

  • 构建多样化的验证性证据冲突数据集模拟真实误导场景。
  • NLI和大模型在冲突检测中精度高,但小模型召回率低。
  • 大模型解冲突时常偏袒一方且缺乏解释,依赖已有认知。

大语言模型(LLMs)的兴起显著影响了决策系统中的信息质量,导致生成内容泛滥,难以识别错误信息及管理相互矛盾的信息(即“跨证据冲突”)。本研究提出一种生成多样化、经验证的证据冲突的方法,以模拟现实中的虚假信息场景。我们评估了自然语言推理(NLI)模型、事实一致性(FC)模型及大语言模型在这些冲突上的表现(研究问题1),并分析了大模型在冲突解决中的行为特征(研究问题2)。主要发现包括:(1)NLI和大语言模型在检测答案冲突方面具有高精度,但较弱模型存在召回率不足的问题;(2)FC模型在处理词汇相似的答案冲突时表现较差,而NLI和大语言模型则更具优势;(3)更强的模型如GPT-4展现出稳健性能,尤其在处理细微冲突时。对于冲突解决,大语言模型常单方面偏好某条证据且无合理解释,若其已有先验信念,则更依赖内部知识进行判断。

原文摘要 · Abstract (English)

The rise of large language models (LLMs) has significantly influenced the quality of information in decision-making systems, leading to the prevalence of AI-generated content and challenges in detecting misinformation and managing conflicting information, or "inter-evidence conflicts." This study introduces a method for generating diverse, validated evidence conflicts to simulate real-world misinformation scenarios. We evaluate conflict detection methods, including Natural Language Inference (NLI) models, factual consistency (FC) models, and LLMs, on these conflicts (RQ1) and analyze LLMs' conflict resolution behaviors (RQ2). Our key findings include: (1) NLI and LLM models exhibit high precision in detecting answer conflicts, though weaker models suffer from low recall; (2) FC models struggle with lexically similar answer conflicts, while NLI and LLM models handle these better; and (3) stronger models like GPT-4 show robust performance, especially with nuanced conflicts. For conflict resolution, LLMs often favor one piece of conflicting evidence without justification and rely on internal knowledge if they have prior beliefs.

证据冲突大模型评估信息真实性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。