arXiv:2503.20320cs.CLcs.AI2025-03被引 1

用迭代说服技巧逐步突破大模型安全限制,成功率最高达90%。

Iterative Prompting with Persuasion Skills in Jailbreaking Large Language Models

  • 通过多轮优化提示词,动态调整攻击策略
  • 对GPT-4和ChatGLM的攻击成功率最高达90%
  • 适合研究模型安全与对抗性提示的学者

大型语言模型(LLMs)旨在其响应中与人类价值观对齐。本研究采用迭代提示技术,通过在多轮中系统性地修改和优化提示词,逐步提升其在越狱攻击中的有效性。该方法分析GPT-3.5、GPT-4、LLaMa2、Vicuna和ChatGLM等模型的响应模式,以调整并优化提示词,从而绕过模型的伦理与安全约束。结合说服策略可增强提示效果,同时保持恶意意图的一致性。实验结果表明,随着提示词不断优化,攻击成功率(ASR)持续上升,对GPT-4和ChatGLM的最高ASR达到90%,对LLaMa2的最低为68%。该方法在ASR上优于基线技术(PAIR和PAP),且与GCG和ArtPrompt表现相当。

原文摘要 · Abstract (English)

Large language models (LLMs) are designed to align with human values in their responses. This study exploits LLMs with an iterative prompting technique where each prompt is systematically modified and refined across multiple iterations to enhance its effectiveness in jailbreaking attacks progressively. This technique involves analyzing the response patterns of LLMs, including GPT-3.5, GPT-4, LLaMa2, Vicuna, and ChatGLM, allowing us to adjust and optimize prompts to evade the LLMs' ethical and security constraints. Persuasion strategies enhance prompt effectiveness while maintaining consistency with malicious intent. Our results show that the attack success rates (ASR) increase as the attacking prompts become more refined with the highest ASR of 90% for GPT4 and ChatGLM and the lowest ASR of 68% for LLaMa2. Our technique outperforms baseline techniques (PAIR and PAP) in ASR and shows comparable performance with GCG and ArtPrompt.

越狱攻击提示工程安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。