用提示词攻击让大模型代理行为失控,暴露内部指令。
Doppelganger Method: Breaking Role Consistency in LLM Agent via Prompt-based Transferable Adversarial Attack
- 设计可迁移的对抗性提示,诱导代理违背原设定。
- 攻击成功率超90%,能成功泄露系统指令和内部信息。
- 提出防御方案CAT,有效抵御此类攻击,适合安全研究者参考。
自大型语言模型问世以来,提示工程使得快速、低成本创建多样化自主代理成为可能,已广泛部署。然而这种便利性也带来了安全、鲁棒性和行为一致性方面的严峻挑战,尤其是防范用户通过提示攻击暴露底层提示内容。本文提出「Doppelganger方法」,展示代理被劫持的风险,从而暴露系统指令与内部信息。我们定义了「对抗迁移下的提示对齐崩溃(PACAT)」评估等级以衡量此类攻击的脆弱性。同时提出「警惕对抗迁移(CAT)」提示作为防御手段。实验表明,Doppelganger方法能有效破坏代理的一致性并泄露内部信息;而CAT提示可实现有效防御。
原文摘要 · Abstract (English)
Since the advent of large language models, prompt engineering now enables the rapid, low-effort creation of diverse autonomous agents that are already in widespread use. Yet this convenience raises urgent concerns about the safety, robustness, and behavioral consistency of the underlying prompts, along with the pressing challenge of preventing those prompts from being exposed to user's attempts. In this paper, we propose the ''Doppelganger method'' to demonstrate the risk of an agent being hijacked, thereby exposing system instructions and internal information. Next, we define the ''Prompt Alignment Collapse under Adversarial Transfer (PACAT)'' level to evaluate the vulnerability to this adversarial transfer attack. We also propose a ''Caution for Adversarial Transfer (CAT)'' prompt to counter the Doppelganger method. The experimental results demonstrate that the Doppelganger method can compromise the agent's consistency and expose its internal information. In contrast, CAT prompts enable effective defense against this adversarial attack.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。