arXiv:2504.05689cs.CLcs.CR2025-04被引 3

发现角色分隔符会引发对话模型偏见,可被攻击者利用操控模型行为。

Separator Injection Attack: Uncovering Dialogue Biases in Large Language Models Caused by Role Separators

  • 利用角色分隔符构造新型注入攻击,触发模型位置偏差
  • 自动攻击使成功率达100%,手动攻击提升18.2%成功率
  • 揭示对话系统深层安全缺陷,适合关注模型安全的研究者

对话型大语言模型(LLMs)因具备指令遵循能力而受到广泛关注。为确保模型正确响应,通常使用角色分隔符区分对话参与者。然而,角色分隔符的引入可能带来安全隐患。不当使用角色分隔符会导致提示注入攻击,使模型行为偏离用户意图,引发严重安全问题。尽管已有多种提示注入攻击方法,但近期研究大多忽视了角色分隔符对模型安全的影响。本文揭示了由角色分隔符引发的建模缺陷,发现其存在固有的位置偏差,该偏差源于对话建模格式,可通过插入角色分隔符被触发。我们进一步提出分离符注入攻击(SIA),一种基于角色分隔符的新型正交攻击方法。实验结果表明,SIA在操纵模型行为方面高效且广泛:手动方法平均提升18.2%攻击效果,自动方法将攻击成功率提升至100%。

原文摘要 · Abstract (English)

Conversational large language models (LLMs) have gained widespread attention due to their instruction-following capabilities. To ensure conversational LLMs follow instructions, role separators are employed to distinguish between different participants in a conversation. However, incorporating role separators introduces potential vulnerabilities. Misusing roles can lead to prompt injection attacks, which can easily misalign the model's behavior with the user's intentions, raising significant security concerns. Although various prompt injection attacks have been proposed, recent research has largely overlooked the impact of role separators on safety. This highlights the critical need to thoroughly understand the systemic weaknesses in dialogue systems caused by role separators. This paper identifies modeling weaknesses caused by role separators. Specifically, we observe a strong positional bias associated with role separators, which is inherent in the format of dialogue modeling and can be triggered by the insertion of role separators. We further develop the Separators Injection Attack (SIA), a new orthometric attack based on role separators. The experiment results show that SIA is efficient and extensive in manipulating model behavior with an average gain of 18.2% for manual methods and enhances the attack success rate to 100% with automatic methods.

模型安全提示攻击对话系统角色分隔符

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。