通过语义网络攻击让大模型整体观点偏移,隐蔽性强且效果显著。
DiscourseFlip: An Oblique Discourse-Level Opinion Manipulation Attack against Black-box Retrieval-Augmented Generation

- 基于图结构动态分配污染预算,实现跨话题协同操纵。
- 在多主题查询网络中实现稳定的目标观点偏移,覆盖范围更广。
- 现有防御手段无效,适合研究安全与对抗攻击的学者参考。
检索增强生成(RAG)系统广泛部署,但依赖外部语料使其面临毒化检索内容的新安全风险。现有攻击主要针对单个查询或局部主题查询集,实用性有限且难以伪装。本文提出话语级观点操纵新威胁模型:通过语义查询网络中的协同影响,实现对多主题查询空间的整体观点偏移。我们在黑盒环境下形式化该威胁,并提出DiscourseFlip——一种代理式、图引导的攻击方法,能动态分配有限污染预算以最大化话语级观点偏离。大量实验表明,DiscourseFlip在上下文查询网络中持续引发目标观点偏移,显著优于现有基线,在覆盖率和有效性上均表现更优。用户研究表明其有效且不易被察觉。系统性分析显示,现有缓解策略对话语级操纵无效,凸显亟需更鲁棒、自适应的防御机制应对此类新型漏洞。
原文摘要 · Abstract (English)
Retrieval-Augmented Generation (RAG) systems are widely deployed and increasingly influential, but their reliance on external corpora exposes new security risks from poisoned retrieval content. Existing RAG attacks are largely focusing on individual queries or narrow topic-local query sets, which limits their practical reach and offers limited camouflage in real-world settings. In this paper, we introduce discourse-level opinion manipulation, a new threat model in which coordinated influence across a semantic query network induces opinion shifts over a holistic, multi-topic query space. We formalize this threat in a black-box setting and propose DiscourseFlip, an agentic, graph-guided attack that dynamically allocates a limited poisoning budget to maximize discourse-level opinion deviation. Extensive experiments demonstrate that DiscourseFlip consistently induces targeted opinion shifts across the contextualized query network and significantly outperforms existing baselines in terms of coverage and effectiveness. User studies further confirm that DiscourseFlip is effective while remaining well camouflaged from user detection. Moreover, systematic analyses show that existing mitigation strategies are ineffective against discourse-level manipulation, underscoring the urgent need for more robust and adaptive defenses to address discourse-level vulnerabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。