攻击者可通过污染检索系统放大大模型偏见,误导生成结果。
Bias Amplification in RAG: Poisoning Knowledge Retrieval to Steer LLMs
- 设计多目标奖励生成对抗文档,操控检索结果
- 在多个大模型上使偏见增强达37%以上
- 揭示检索增强架构的公平性风险,适合安全与伦理研究者
大型语言模型中的检索增强生成(RAG)系统虽能提升性能,但也引入新安全风险。现有研究多关注中毒攻击对输出质量的影响,却忽视其放大模型偏见的潜在危害。例如,查询家庭暴力受害者时,被攻陷的RAG系统可能优先检索女性受害的文档,导致模型生成内容强化性别刻板印象,即使原查询中性。本文提出偏见检索与奖励攻击(BRRA)框架,系统分析通过操纵RAG系统放大语言模型偏见的路径。设计基于多目标奖励函数的对抗文档生成方法,采用子空间投影技术操控检索结果,并构建循环反馈机制实现持续偏见放大。在多个主流大模型上的实验表明,BRRA可显著增强模型在多个维度的偏见。此外,我们探索了双阶段防御机制,有效缓解攻击影响。本研究揭示了RAG系统中毒攻击会直接放大模型输出偏见,阐明了系统安全与模型公平之间的关系,提示需关注RAG系统的公平性问题。
原文摘要 · Abstract (English)
In Large Language Models, Retrieval-Augmented Generation (RAG) systems can significantly enhance the performance of large language models by integrating external knowledge. However, RAG also introduces new security risks. Existing research focuses mainly on how poisoning attacks in RAG systems affect model output quality, overlooking their potential to amplify model biases. For example, when querying about domestic violence victims, a compromised RAG system might preferentially retrieve documents depicting women as victims, causing the model to generate outputs that perpetuate gender stereotypes even when the original query is gender neutral. To show the impact of the bias, this paper proposes a Bias Retrieval and Reward Attack (BRRA) framework, which systematically investigates attack pathways that amplify language model biases through a RAG system manipulation. We design an adversarial document generation method based on multi-objective reward functions, employ subspace projection techniques to manipulate retrieval results, and construct a cyclic feedback mechanism for continuous bias amplification. Experiments on multiple mainstream large language models demonstrate that BRRA attacks can significantly enhance model biases in dimensions. In addition, we explore a dual stage defense mechanism to effectively mitigate the impacts of the attack. This study reveals that poisoning attacks in RAG systems directly amplify model output biases and clarifies the relationship between RAG system security and model fairness. This novel potential attack indicates that we need to keep an eye on the fairness issues of the RAG system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。