arXiv:2504.18041cs.CLcs.AI2025-04NAACL被引 38

RAG框架可能让大模型更不安全,需专门的安全评估方法。

RAG LLMs are Not Safer: A Safety Analysis of Retrieval-Augmented Generation for Large Language Models

  • 对比11个大模型的RAG与非RAG框架,分析安全表现差异。
  • 即使使用安全模型和安全文档,RAG仍会生成不安全内容。
  • 现有红队测试方法在RAG场景下效果显著下降,需针对性改进。

为确保大语言模型(LLMs)的安全性,已有大量工作聚焦于安全微调、评估与红队测试。然而,尽管检索增强生成(RAG)框架被广泛使用,安全研究仍集中于标准LLMs,对RAG应用场景下的安全影响了解有限。本文对11个大模型在RAG与非RAG框架下进行详细对比分析,发现RAG可能降低模型安全性并改变其安全特征。我们探究了成因,发现即便结合安全模型与安全文档,仍可能导致不安全输出。此外,评估现有红队测试方法在RAG场景下的有效性,结果表明其效果显著低于非RAG设置。本研究强调必须发展针对RAG LLMs的专门安全研究与红队方法。

原文摘要 · Abstract (English)

Efforts to ensure the safety of large language models (LLMs) include safety fine-tuning, evaluation, and red teaming. However, despite the widespread use of the Retrieval-Augmented Generation (RAG) framework, AI safety work focuses on standard LLMs, which means we know little about how RAG use cases change a model's safety profile. We conduct a detailed comparative analysis of RAG and non-RAG frameworks with eleven LLMs. We find that RAG can make models less safe and change their safety profile. We explore the causes of this change and find that even combinations of safe models with safe documents can cause unsafe generations. In addition, we evaluate some existing red teaming methods for RAG settings and show that they are less effective than when used for non-RAG settings. Our work highlights the need for safety research and red-teaming methods specifically tailored for RAG LLMs.

RAG安全红队测试LLM安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。