arXiv:2502.01925cs.CLcs.CR2025-02ICML被引 7

通过优化对话模板提升大模型越狱成功率,更隐蔽地绕过安全限制。

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling

  • 用正向肯定和负向示范改造伪造对话,增强诱导性。
  • 在长上下文场景下,越狱成功率显著高于基线方法。
  • 适用于研究模型安全漏洞的人员,尤其关注长文本攻击者。

多轮越狱攻击通过在目标提示前添加数百个伪造的用户-模型对话,利用大模型处理长序列的能力绕过安全对齐。这些对话通常从危险问答对池中随机采样,使模型看似已响应恶意指令。本文提出PANDAS:一种混合技术,通过正向肯定、负向示范和针对目标话题优化的自适应采样,改进伪造对话质量。我们还构建了ManyHarm数据集(含有害问答对),并通过大量实验验证,PANDAS在长上下文场景中显著优于基线方法。注意力分析揭示了长上下文漏洞的利用机制,并说明PANDAS如何进一步增强越狱效果。

原文摘要 · Abstract (English)

Many-shot jailbreaking circumvents the safety alignment of LLMs by exploiting their ability to process long input sequences. To achieve this, the malicious target prompt is prefixed with hundreds of fabricated conversational exchanges between the user and the model. These exchanges are randomly sampled from a pool of unsafe question-answer pairs, making it appear as though the model has already complied with harmful instructions. In this paper, we present PANDAS: a hybrid technique that improves many-shot jailbreaking by modifying these fabricated dialogues with Positive Affirmations, Negative Demonstrations, and an optimized Adaptive Sampling method tailored to the target prompt's topic. We also introduce ManyHarm, a dataset of harmful question-answer pairs, and demonstrate through extensive experiments that PANDAS significantly outperforms baseline methods in long-context scenarios. Through attention analysis, we provide insights into how long-context vulnerabilities are exploited and show how PANDAS further improves upon many-shot jailbreaking.

越狱攻击大模型安全长文本生成对抗样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。