通过模仿人类语音风格,破解音频大模型的对齐安全漏洞。
StyleBreak: Revealing Alignment Vulnerabilities in Large Audio-Language Models via Style-Aware Audio Jailbreak
- 设计双阶段风格可控音频转换,同时改变语义和语音特征。
- 在多个模型上实现更高成功率与更少查询次数的攻击效果。
- 揭示语音表达多样性对模型安全的影响,适合安全研究者参考。
大型音频语言模型(LAM)通过结合音频编码器与大语言模型(LLM),实现了强大的语音交互能力。然而,针对LAM的对抗攻击,特别是通过恶意音频提示绕过对齐机制的安全性,仍缺乏深入研究。现有方法主要依赖将文本攻击转为语音或施加浅层信号扰动,忽视了人类语音中丰富的表现力变化对模型对齐鲁棒性的影响。为此,本文提出StyleBreak——一种新型风格感知音频越狱框架,系统探究多种人类语音属性对LAM对齐鲁棒性的影响。StyleBreak采用两阶段风格感知转换流程,同时扰动文本内容与音频信号,以控制语言、副语言及超语言属性。此外,设计了一种查询自适应策略网络,自动搜索对抗性语音风格,提升越狱探索效率。大量实验证明,当暴露于多样化语音属性时,LAM表现出严重安全漏洞。同时,StyleBreak在多种攻击范式下显著提升攻击有效性和效率,凸显了强化LAM对齐鲁棒性的紧迫性。
原文摘要 · Abstract (English)
Large Audio-language Models (LAMs) have recently enabled powerful speech-based interactions by coupling audio encoders with Large Language Models (LLMs). However, the security of LAMs under adversarial attacks remains underexplored, especially through audio jailbreaks that craft malicious audio prompts to bypass alignment. Existing efforts primarily rely on converting text-based attacks into speech or applying shallow signal-level perturbations, overlooking the impact of human speech's expressive variations on LAM alignment robustness. To address this gap, we propose StyleBreak, a novel style-aware audio jailbreak framework that systematically investigates how diverse human speech attributes affect LAM alignment robustness. Specifically, StyleBreak employs a two-stage style-aware transformation pipeline that perturbs both textual content and audio to control linguistic, paralinguistic, and extralinguistic attributes. Furthermore, we develop a query-adaptive policy network that automatically searches for adversarial styles to enhance the efficiency of LAM jailbreak exploration. Extensive evaluations demonstrate that LAMs exhibit critical vulnerabilities when exposed to diverse human speech attributes. Moreover, StyleBreak achieves substantial improvements in attack effectiveness and efficiency across multiple attack paradigms, highlighting the urgent need for more robust alignment in LAMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。