arXiv:2601.06189cs.AIcs.LG2026-01ACL被引 1

分析大模型在检索增强生成中的推理机制,发现其更依赖重复和顺序而非证据真实强度。

Rational Synthesizers or Heuristic Followers? Analyzing LLMs in RAG-based Question-Answering

  • 构建含1635个争议问题的GroupQA数据集,标注立场与证据强度。
  • 模型更倾向首次出现的证据,且大模型更难被新证据说服。
  • 模型解释不忠实,暴露其依赖启发式而非理性推理的弱点。

检索增强生成(RAG)是当前大语言模型(LLM)实现知识对齐的主要范式,但模型如何整合存在冲突的多份检索证据仍不清晰。大模型的回答是基于事实强度、先验信念,还是仅仅因为信息被频繁重复?为此,我们提出了GroupQA,一个包含1,635个争议性问题和15,058份来源多样证据文档的标注数据集,每条证据均标注立场与定性强度。通过控制实验,我们揭示了群体级证据聚合的动态特征:改写论点比提供独立支持更具说服力;模型偏好最先呈现的证据而非最后的;更大的模型对新证据的适应性反而降低。此外,我们发现大模型对群体回答的解释缺乏忠实性。综上,大模型表现出一致的启发式跟随者特征,对改进RAG系统设计具有直接意义。

原文摘要 · Abstract (English)

Retrieval-Augmented Generation (RAG) is the prevailing paradigm for grounding Large Language Models (LLMs), yet the mechanisms governing how models integrate groups of conflicting retrieved evidence remain opaque. Does an LLM answer a certain way because the evidence is factually strong, because of a prior belief, or merely because it is repeated frequently? To answer this, we introduce GroupQA, a curated dataset of 1,635 controversial questions paired with 15,058 diversely-sourced evidence documents, annotated for stance and qualitative strength. Through controlled experiments, we characterize group-level evidence aggregation dynamics: Paraphrasing an argument can be more persuasive than providing distinct independent support; Models favor evidence presented first rather than last, and Larger models are increasingly resistant to adapt to presented evidence. Additionally, we find that LLM explanations to group-based answers are unfaithful. Together, we show that LLMs behave consistently as vulnerable heuristic followers, with direct implications for improving RAG system design.

RAG大模型推理证据聚合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。