提出针对RAG模型的议题型对抗操纵攻击,可系统性误导多角度观点生成。
Topic-FlipRAG: Topic-Orientated Adversarial Opinion Manipulation Attacks to Retrieval-Augmented Generation Models
- 分两阶段设计对抗扰动,利用大模型推理能力实现语义级操纵。
- 实验表明攻击能显著改变模型在特定议题上的输出倾向。
- 揭示现有防御手段无效,对LLM安全研究有重要警示意义。
基于大语言模型的检索增强生成(RAG)系统已成为问答和内容生成等任务的关键技术。然而,其在公共舆论与信息传播中日益重要的影响,使其成为安全研究的重点,因其固有的脆弱性。以往研究主要关注事实性或单查询的攻击。本文针对更实际的场景——面向议题的对抗性意见操纵攻击,即要求大模型综合多个视角进行推理,使其极易受到系统性知识污染的影响。我们提出Topic-FlipRAG,一种两阶段操纵攻击流程,通过精心设计的对抗扰动影响相关查询的输出观点。该方法结合传统对抗排序攻击技术,并利用大模型内部丰富的相关知识与推理能力,实施语义层级的扰动。实验表明,该攻击能有效改变模型在特定议题上的输出意见,显著影响用户信息认知。现有缓解方法无法有效防御此类攻击,凸显了加强RAG系统防护的必要性,并为大模型安全研究提供了关键洞见。
原文摘要 · Abstract (English)
Retrieval-Augmented Generation (RAG) systems based on Large Language Models (LLMs) have become essential for tasks such as question answering and content generation. However, their increasing impact on public opinion and information dissemination has made them a critical focus for security research due to inherent vulnerabilities. Previous studies have predominantly addressed attacks targeting factual or single-query manipulations. In this paper, we address a more practical scenario: topic-oriented adversarial opinion manipulation attacks on RAG models, where LLMs are required to reason and synthesize multiple perspectives, rendering them particularly susceptible to systematic knowledge poisoning. Specifically, we propose Topic-FlipRAG, a two-stage manipulation attack pipeline that strategically crafts adversarial perturbations to influence opinions across related queries. This approach combines traditional adversarial ranking attack techniques and leverages the extensive internal relevant knowledge and reasoning capabilities of LLMs to execute semantic-level perturbations. Experiments show that the proposed attacks effectively shift the opinion of the model's outputs on specific topics, significantly impacting user information perception. Current mitigation methods cannot effectively defend against such attacks, highlighting the necessity for enhanced safeguards for RAG systems, and offering crucial insights for LLM security research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。