设计多步伦理伪装提示,测试大模型防御恶意攻击的能力。
"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks
- 用职场晋升场景伪装多步提示,绕过模型安全机制。
- GPT-4o等五款模型均被成功攻破,生成言语攻击内容。
- Claude 3.5 Sonnet防御更强,适合安全评估研究者参考。
随着大语言模型在各领域的广泛应用,识别有害内容生成及防护机制的有效性面临更高挑战。本研究通过黑盒测试,评估GPT-4o、Grok-2 Beta、Llama 3.1(405B)、Gemini 1.5和Claude 3.5 Sonnet的防护能力,采用看似符合伦理的多步提示,模拟“企业中层管理者争夺晋升”的情境进行道德攻击。实验结果表明,上述大模型的防护机制均被绕过,成功生成了言语攻击内容。其中,Claude 3.5 Sonnet对多步越狱提示的抵抗能力更为明显。为保证实验可复现与客观性,本文已将实验流程、黑盒测试代码及增强防护代码上传至GitHub:https://github.com/brucewang123456789/GeniusTrail.git。
原文摘要 · Abstract (English)
As the application of large language models continues to expand in various fields, it poses higher challenges to the effectiveness of identifying harmful content generation and guardrail mechanisms. This research aims to evaluate the guardrail effectiveness of GPT-4o, Grok-2 Beta, Llama 3.1 (405B), Gemini 1.5, and Claude 3.5 Sonnet through black-box testing of seemingly ethical multi-step jailbreak prompts. It conducts ethical attacks by designing an identical multi-step prompts that simulates the scenario of "corporate middle managers competing for promotions." The data results show that the guardrails of the above-mentioned LLMs were bypassed and the content of verbal attacks was generated. Claude 3.5 Sonnet's resistance to multi-step jailbreak prompts is more obvious. To ensure objectivity, the experimental process, black box test code, and enhanced guardrail code are uploaded to the GitHub repository: https://github.com/brucewang123456789/GeniusTrail.git.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。