arXiv:2412.18004cs.CL2024-12被引 22

揭示检索增强生成中引用不忠问题,强调可信溯源需兼顾正确性与真实依赖。

Correctness is not Faithfulness in RAG Attributions

  • 区分引用正确性与忠实性,指出前者不能保证真实引用。
  • 实验发现高达57%的引用存在后理性化现象,非真实依赖。
  • 适合关注大模型可解释性与可信生成的研究者。

检索相关上下文是减少幻觉、提升回答可靠性常用方法。明确引用源文档可让用户验证生成内容,增强信任。以往研究主要评估引用正确性——即引用的文档是否支持对应陈述。但仅靠正确性不足。为建立可信的引用,必须同时考察引用正确性与引用忠实性。本文首次厘清这两个概念,它们在先前研究中常被混淆。忠实性确保模型对引用文档的依赖是真实的,而非仅表面契合既有认知(称作后理性化)。我们设计实验揭示后理性化普遍存在,削弱了可靠引用,可能导致错误信任。研究发现当前引用答案中高达57%缺乏忠实性,表明在语言模型中实现可信引用,必须同时评估正确性与忠实性。

原文摘要 · Abstract (English)

Retrieving relevant context is a common approach to reduce hallucinations and enhance answer reliability. Explicitly citing source documents allows users to verify generated responses and increases trust. Prior work largely evaluates citation correctness - whether cited documents support the corresponding statements. But citation correctness alone is insufficient. To establish trust in attributed answers, we must examine both citation correctness and citation faithfulness. In this work, we first disentangle the notions of citation correctness and faithfulness, which have been applied inconsistently in previous studies. Faithfulness ensures that the model's reliance on cited documents is genuine, reflecting actual reference use rather than superficial alignment with prior beliefs, which we call post-rationalization. We design an experiment that reveals the prevalent issue of post-rationalization, which undermines reliable attribution and may result in misplaced trust. Our findings suggest that current attributed answers often lack citation faithfulness (up to 57 percent of the citations), highlighting the need to evaluate correctness and faithfulness for trustworthy attribution in language models.

RAG可解释性信任

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。