提出新型音频越狱攻击,可绕过语音模型安全防护。
AudioJailbreak: Jailbreak Attacks against End-to-End Large Audio-Language Models
- 设计异步音频扰动,不需与提示对齐即可生效
- 单个扰动可适配多类提示,且在空中播放仍有效
- 隐匿恶意意图,适用于真实场景中的弱攻击者
近期针对端到端大音频语言模型(LALMs)的越狱攻击主要集中于攻击者可完全操控用户提示的强对抗场景,但效果有限。本文首次系统评估发现,先进文本越狱方法难以通过文本转语音技术移植至端到端LALMs。为此,提出AUDIOJAILBREAK,具备四大特性:(1)异步性——通过尾部扰动音频实现时间轴非对齐;(2)通用性——将多个提示融合生成单一扰动;(3)隐蔽性——采用多种策略隐藏恶意意图;(4)空中鲁棒性——在生成中加入混响模拟真实播放环境。相比现有攻击,本方法同时满足上述四点。更重要的是,其适用于更现实的弱对抗场景(攻击者无法完全操控提示)。在目前最全面的LALMs测试中,该方法成功越狱OpenAI的GPT-4o-Audio,并绕过Meta的Llama-Guard-3防护,在弱对抗下仍具高有效性。
原文摘要 · Abstract (English)
Jailbreak attacks to Large audio-language models (LALMs) are studied recently, but they exclusively focused on the attack scenario where the adversary can fully manipulate user prompts (named strong adversary) and limited in effectiveness, applicability, and practicability. In this work, we first conduct an extensive evaluation showing that advanced text jailbreak attacks cannot be easily ported to end-to-end LALMs via text-to-speech (TTS) techniques. We then propose AUDIOJAILBREAK, a novel audio jailbreak attack, featuring (1) asynchrony: the jailbreak audios do not need to align with user prompts in the time axis by crafting suffixal jailbreak audios; (2) universality: a single jailbreak perturbation is effective for different prompts by incorporating multiple prompts into the perturbation generation; (3) stealthiness: the malicious intent of jailbreak audios is concealed by proposing various intent concealment strategies; and (4) over-the-air robustness: the jailbreak audios remain effective when being played over the air by incorporating reverberation into the perturbation generation. In contrast, all prior audio jailbreak attacks cannot offer asynchrony, universality, stealthiness, and/or over-the-air robustness. Moreover, AUDIOJAILBREAK is also applicable to a more practical and broader attack scenario where the adversary cannot fully manipulate user prompts (named weak adversary). Extensive experiments with thus far the most LALMs demonstrate the high effectiveness of AUDIOJAILBREAK, in particular, it can jailbreak openAI's GPT-4o-Audio and bypass Meta's Llama-Guard-3 safeguard, in the weak adversary scenario. We highlight that our work peeks into the security implications of audio jailbreak attacks against LALMs, and realistically fosters improving their robustness, especially for the newly proposed weak adversary.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。