arXiv:2502.01154cs.CLcs.AI2025-02NAACL被引 2

用通用多提示攻击大模型,一次训练可攻破多种任务。

Jailbreaking with Universal Multi-Prompts

  • 设计通用多提示策略,一次优化适配多个任务。
  • 在未见过的任务上仍能成功越狱,攻击成功率更高。
  • 既可用于攻击,也可用于防御,适用范围广。

近年来大型语言模型(LLMs)快速发展,广泛应用于各类场景并显著提升效率。但随之而来的是伦理风险和新型攻击,如越狱攻击。现有提示技术多针对单个案例优化,处理大规模数据时计算成本高;而针对通用攻击者训练的研究较少。本文提出JUMP方法,通过提示优化实现基于通用多提示的越狱攻击,并将该思路扩展至防御,称为DUMP。实验表明,该方法在优化通用多提示方面优于现有技术,在未见任务上仍具强迁移能力。

原文摘要 · Abstract (English)

Large language models (LLMs) have seen rapid development in recent years, revolutionizing various applications and significantly enhancing convenience and productivity. However, alongside their impressive capabilities, ethical concerns and new types of attacks, such as jailbreaking, have emerged. While most prompting techniques focus on optimizing adversarial inputs for individual cases, resulting in higher computational costs when dealing with large datasets. Less research has addressed the more general setting of training a universal attacker that can transfer to unseen tasks. In this paper, we introduce JUMP, a prompt-based method designed to jailbreak LLMs using universal multi-prompts. We also adapt our approach for defense, which we term DUMP. Experimental results demonstrate that our method for optimizing universal multi-prompts outperforms existing techniques.

越狱攻击通用提示大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。