arXiv:2605.18168cs.CRcs.SD2026-05被引 1

用特定声学特征干扰大模型安全机制,实现无需恶意内容的通用越狱。

Acoustic Interference: A New Paradigm Weaponizing Acoustic Latent Semantic for Universal Jailbreak against Large Audio Language Models

论文配图:Acoustic Interference: A New Paradigm Weaponizing Acoustic Latent Semantic for Universal Jailbreak against Large Audio Language Models
图 1 · 摘自论文原文
  • 利用音频内在的声学语义特征作为干扰信号,不依赖恶意内容注入
  • 在10个大音频模型上达成顶尖越狱成功率,且无需针对每条指令优化
  • 揭示了跨模态对齐的深层漏洞,适合安全研究者和防御开发者参考

将音频模态融入大型音频语言模型(LALMs)显著扩大了其攻击面。现有越狱方法多将音频视为恶意载荷的载体,通过语义优化、声学参数控制或添加扰动来嵌入有害内容。本文挑战这一必要性,提出新范式:音频角色从内容注入转为安全对齐干扰。我们发现,仅凭特定的声学潜在语义(ALS),即音频生成模型先验中的内在副语言特征,即可破坏LALM的安全对齐。不同于以往通过显式声学参数仅风格化恶意音频的方法,我们证明:内容无害但注入特定ALS的干扰音频,可作为通用越狱触发器。基于此,我们提出声学干扰攻击(AIA),将攻击载荷与音频解耦。AIA使用一组通用、指令无关的干扰音频,使标准恶意文本查询无需实例化优化即可绕过安全对齐。在五个数据集上的10个LALMs上广泛实验表明,AIA达到当前最优攻击成功率。此外,可解释性分析揭示了AIA引发的推理路径偏移,并识别出ALS中固有的有效模式,暴露出LALMs跨模态对齐的根本脆弱性。

原文摘要 · Abstract (English)

The integration of audio modality into Large Audio Language Models (LALMs) significantly expands their attack surface. Existing jailbreak paradigms predominantly treat audio as a carrier for malicious payloads, relying on semantic optimization, acoustic parameter control, or additive perturbation to embed harmful content into the audio signal. In this work, we challenge this necessity and propose a new paradigm in which the role of audio shifts from content injection to safety alignment interference. We reveal that LALM safety alignment can be compromised solely by specific Acoustic Latent Semantics (ALS), the underlying paralinguistic features intrinsic to the priors of audio generative models. Distinct from previous works that leverage explicit acoustic parameters to merely style malicious audio, we demonstrate that interference audio, benign in content but infused with specific ALS, can serve as a universal jailbreak trigger. Leveraging this insight, we propose the Acoustic Interference Attack (AIA), which decouples the attack payload from the audio. Specifically, AIA employs a set of universal, instruction-neutral interference audio, enabling standard malicious text queries to bypass safety alignment without instance-specific optimization. Extensive experiments on 10 LALMs across five datasets demonstrate that AIA achieves the state-of-the-art attack success rate. Furthermore, our interpretability analysis uncovers the inference path drift induced by AIA and identifies the inherent effective patterns within ALS, revealing the fundamental vulnerability of cross-modal alignment in LALMs.

安全攻击音频模型越狱对抗样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。