揭示检索错误对增强型语言模型的致命影响,提出新评估方法
Toward Robust RALMs: Revealing the Impact of Imperfect Retrieval on Retrieval-Augmented Language Models
- 构建生成式对抗攻击方法GenADV与新指标RAD,系统测试模型鲁棒性
- 发现模型在无法回答、矛盾、对抗等场景下常产生幻觉,性能显著下降
- 适用于研究模型安全性的学者,尤其关注真实场景可靠性问题
检索增强语言模型(RALMs)因能生成准确答案并提升效率而受到广泛关注。然而,由于依赖不完美的检索器或知识源,RALMs天然易受错误信息影响。我们识别出三种常见场景——无法回答、对抗性、冲突性——其检索文档集可能误导模型,且具有现实代表性。本文首次全面评估了RALMs在这些异常场景下的检测与处理能力。为系统检验对抗鲁棒性,我们提出一种基于生成模型的对抗攻击方法GenADV和新指标鲁棒性附加文档度量(RAD)。研究发现,RALMs往往无法识别文档集的不可回答性或矛盾性,导致频繁产生幻觉。此外,加入对抗样本显著降低模型表现,当对抗与不可回答场景重叠时,模型脆弱性进一步加剧。本研究指出了评估与提升RALMs鲁棒性的关键方向,为构建更可靠的模型奠定基础。
原文摘要 · Abstract (English)
Retrieval Augmented Language Models (RALMs) have gained significant attention for their ability to generate accurate answer and improve efficiency. However, RALMs are inherently vulnerable to imperfect information due to their reliance on the imperfect retriever or knowledge source. We identify three common scenarios-unanswerable, adversarial, conflicting-where retrieved document sets can confuse RALM with plausible real-world examples. We present the first comprehensive investigation to assess how well RALMs detect and handle such problematic scenarios. Among these scenarios, to systematically examine adversarial robustness we propose a new adversarial attack method, Generative model-based ADVersarial attack (GenADV) and a novel metric Robustness under Additional Document (RAD). Our findings reveal that RALMs often fail to identify the unanswerability or contradiction of a document set, which frequently leads to hallucinations. Moreover, we show the addition of an adversary significantly degrades RALM's performance, with the model becoming even more vulnerable when the two scenarios overlap (adversarial+unanswerable). Our research identifies critical areas for assessing and enhancing the robustness of RALMs, laying the foundation for the development of more robust models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。