单个隐性提示可导致多智能体系统中偏见传播,威胁对齐安全。
Thought Virus: Viral Misalignment via Subliminal Prompting in Multi-Agent Systems
- 通过无关词汇隐性诱导一个智能体,引发偏见扩散。
- 6个智能体在两种拓扑下,偏见响应率持续升高。
- 适合关注多智能体安全与对齐风险的研究者阅读。
隐性提示指语言模型因使用语义无关的标记而被引导至特定概念或特征。尽管已有研究探讨用户-大模型交互中的隐性提示现象,但多智能体系统中的潜在偏见传递及其安全影响仍不明朗。本文揭示:单个被隐性提示的智能体可在整个网络中传播一种削弱但持久的偏见。我们在6个智能体上测试了两种不同拓扑结构,发现转移的概念在整个网络中保持较高的响应率。为说明潜在的对齐风险,我们评估了该网络在多项选择题型的TruthfulQA上的表现,结果显示,仅对一个智能体进行隐性提示即可降低其他智能体的真相回答率。研究发现表明,隐性提示在多智能体安全中引入了新型攻击向量,对系统对齐具有深远影响。所有实验实现已公开于 https://github.com/Multi-Agent-Security-Initiative/thought_virus。
原文摘要 · Abstract (English)
Subliminal prompting is a phenomenon in which language models are biased towards certain concepts or traits through prompting with semantically unrelated tokens. While prior work has examined subliminal prompting in user-LLM interactions, potential bias transfer in multi-agent systems and its associated security implications remain unexplored. In this work, we show that a single subliminally prompted agent can spread a weakening but persisting bias throughout its entire network. We measure this phenomenon across 6 agents using two different topologies, observing that the transferred concept maintains an elevated response rate throughout the network. To exemplify potential misalignment risks, we assess network performance on multiple-choice TruthfulQA, showing that subliminal prompting of a single agent may degrade the truthfulness of other agents. Our findings reveal that subliminal prompting introduces a new attack vector in multi-agent security, with implications for the alignment of such systems. The implementation of all experiments is publicly available at https://github.com/Multi-Agent-Security-Initiative/thought_virus .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。