arXiv:2509.22486cs.IRcs.CR2025-09EMNLP被引 10

RAG模型存在隐蔽的公平性漏洞,攻击者可长期操控生成内容偏见。

Your RAG is Unfair: Exposing Fairness Vulnerabilities in Retrieval-Augmented Generation via Backdoor Attacks

  • 通过双阶段攻击污染检索与生成交互,植入隐蔽偏见
  • 攻击成功率达高,且不影响内容相关性与可用性
  • 揭示RAG在社会偏见传播上的新风险,适合安全与伦理研究者

检索增强生成(RAG)通过融合检索机制与生成模型提升事实准确性,但引入了新的攻击面,尤其在后门攻击方面。现有研究多关注虚假信息威胁,而公平性漏洞尚未被充分探索。不同于依赖直接触发-目标映射的传统后门,本研究提出基于公平性的攻击,利用检索与生成模型间的交互,操纵目标群体与社会偏见之间的语义关系,实现持久且隐蔽的内容影响。论文提出BiasRAG框架,包含两个阶段:预训练阶段污染查询编码器,使目标群体与特定社会偏见对齐,确保长期影响;部署后阶段向知识库注入对抗性文档,强化后门,细微影响检索结果,且在常规公平性评估下难以察觉。实验表明,BiasRAG在保持上下文相关性和实用性的同时,实现高攻击成功率,构成对RAG公平性的持续且演化的威胁。

原文摘要 · Abstract (English)

Retrieval-augmented generation (RAG) enhances factual grounding by integrating retrieval mechanisms with generative models but introduces new attack surfaces, particularly through backdoor attacks. While prior research has largely focused on disinformation threats, fairness vulnerabilities remain underexplored. Unlike conventional backdoors that rely on direct trigger-to-target mappings, fairness-driven attacks exploit the interaction between retrieval and generation models, manipulating semantic relationships between target groups and social biases to establish a persistent and covert influence on content generation. This paper introduces BiasRAG, a systematic framework that exposes fairness vulnerabilities in RAG through a two-phase backdoor attack. During the pre-training phase, the query encoder is compromised to align the target group with the intended social bias, ensuring long-term persistence. In the post-deployment phase, adversarial documents are injected into knowledge bases to reinforce the backdoor, subtly influencing retrieved content while remaining undetectable under standard fairness evaluations. Together, BiasRAG ensures precise target alignment over sensitive attributes, stealthy execution, and resilience. Empirical evaluations demonstrate that BiasRAG achieves high attack success rates while preserving contextual relevance and utility, establishing a persistent and evolving threat to fairness in RAG.

RAG后门攻击公平性偏见注入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。