用心理学说服策略破解大模型安全防线,发现模型有独特响应模式
Uncovering the Persuasive Fingerprint of LLMs in Jailbreaking Attacks
- 基于社会心理学说服理论设计攻击提示
- 多种对齐模型中说服型提示绕过率显著提升
- 揭示模型在越狱响应中的独特语言指纹
尽管取得进展,大语言模型仍易受越狱攻击,此类攻击可绕过对齐安全机制并诱导有害输出。现有研究多关注攻击策略在可读性与迁移性上的差异,但较少探讨影响模型脆弱性的语言与心理机制。本文结合社会科学中的说服理论,探索如何通过具有说服结构的提示,触发大模型的安全漏洞。我们假设,因训练数据包含大量人类生成文本,大模型可能更易响应具有说服力的提示。实验在多个对齐大模型上验证了该假设,结果表明:具备说服意识的提示能显著绕过安全约束,诱发越狱行为。同时,研究发现模型在越狱响应中表现出独特的语言特征,即‘说服指纹’。本工作强调跨学科视角对应对大模型安全挑战的重要性。代码与数据已公开。
原文摘要 · Abstract (English)
Despite recent advances, Large Language Models remain vulnerable to jailbreak attacks that bypass alignment safeguards and elicit harmful outputs. While prior research has proposed various attack strategies differing in human readability and transferability, little attention has been paid to the linguistic and psychological mechanisms that may influence a model's susceptibility to such attacks. In this paper, we examine an interdisciplinary line of research that leverages foundational theories of persuasion from the social sciences to craft adversarial prompts capable of circumventing alignment constraints in LLMs. Drawing on well-established persuasive strategies, we hypothesize that LLMs, having been trained on large-scale human-generated text, may respond more compliantly to prompts with persuasive structures. Furthermore, we investigate whether LLMs themselves exhibit distinct persuasive fingerprints that emerge in their jailbreak responses. Empirical evaluations across multiple aligned LLMs reveal that persuasion-aware prompts significantly bypass safeguards, demonstrating their potential to induce jailbreak behaviors. This work underscores the importance of cross-disciplinary insight in addressing the evolving challenges of LLM safety. The code and data are available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。