用人格提示词增强大模型越狱攻击,成功率提升10%-20%。
Enhancing Jailbreak Attacks on LLMs via Persona Prompts
- 用遗传算法自动生成人格提示词绕过安全机制。
- 在多款大模型上使拒绝率降低50%-70%。
- 可与现有攻击方法协同,显著提升越狱成功率。
越狱攻击旨在通过诱导大语言模型(LLMs)生成有害内容,揭示其安全漏洞。理解并应对这些攻击对推进大模型安全至关重要。以往方法主要关注直接操纵有害意图,较少关注人格提示的影响。本研究系统探索了人格提示在突破大模型防御中的有效性。我们提出一种基于遗传算法的方法,自动构建人格提示以绕过安全机制。实验表明:(1)所生成的人格提示在多个大模型上将拒绝率降低50%-70%;(2)该提示与现有攻击方法结合时,成功率提升10%-20%。代码与数据已公开于https://github.com/CjangCjengh/Generic_Persona。
原文摘要 · Abstract (English)
Jailbreak attacks aim to exploit large language models (LLMs) by inducing them to generate harmful content, thereby revealing their vulnerabilities. Understanding and addressing these attacks is crucial for advancing the field of LLM safety. Previous jailbreak approaches have mainly focused on direct manipulations of harmful intent, with limited attention to the impact of persona prompts. In this study, we systematically explore the efficacy of persona prompts in compromising LLM defenses. We propose a genetic algorithm-based method that automatically crafts persona prompts to bypass LLM's safety mechanisms. Our experiments reveal that: (1) our evolved persona prompts reduce refusal rates by 50-70% across multiple LLMs, and (2) these prompts demonstrate synergistic effects when combined with existing attack methods, increasing success rates by 10-20%. Our code and data are available at https://github.com/CjangCjengh/Generic_Persona.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。