用简单对话就能让大模型泄露有害内容,暴露安全漏洞。
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
- 通过多轮跨语言互动诱导模型输出可执行的有害指令
- 新攻击框架使成功率平均提升0.319,危害评分提高0.426
- 揭示普通用户也能利用常见交互方式发起攻击
尽管进行了大量安全对齐工作,大型语言模型(LLMs)仍易受越狱攻击,被诱导产生有害行为。现有研究多关注需技术背景的攻击方法,但两个关键问题仍未充分探讨:(1)越狱响应是否真能帮助普通用户实施有害行为?(2)在更常见的简单人机交互中是否存在安全漏洞?本文表明,当模型响应兼具可操作性与信息性时,最能促成有害行为——而这可通过多步、多语言交互轻松激发。基于此,我们提出HarmScore越狱评估指标,以及Speak Easy攻击框架。将Speak Easy应用于直接请求和越狱基线后,在四个安全基准上,开源与专有模型的攻击成功率平均提升0.319,HarmScore提升0.426。本工作揭示了一个常被忽视的关键漏洞:恶意用户可轻易利用普遍交互模式实现有害目的。
原文摘要 · Abstract (English)
Despite extensive safety alignment efforts, large language models (LLMs) remain vulnerable to jailbreak attacks that elicit harmful behavior. While existing studies predominantly focus on attack methods that require technical expertise, two critical questions remain underexplored: (1) Are jailbroken responses truly useful in enabling average users to carry out harmful actions? (2) Do safety vulnerabilities exist in more common, simple human-LLM interactions? In this paper, we demonstrate that LLM responses most effectively facilitate harmful actions when they are both actionable and informative--two attributes easily elicited in multi-step, multilingual interactions. Using this insight, we propose HarmScore, a jailbreak metric that measures how effectively an LLM response enables harmful actions, and Speak Easy, a simple multi-step, multilingual attack framework. Notably, by incorporating Speak Easy into direct request and jailbreak baselines, we see an average absolute increase of 0.319 in Attack Success Rate and 0.426 in HarmScore in both open-source and proprietary LLMs across four safety benchmarks. Our work reveals a critical yet often overlooked vulnerability: Malicious users can easily exploit common interaction patterns for harmful intentions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。