arXiv:2508.02175cs.SDcs.CL2025-08AAAI被引 11

提出隐蔽声学触发器攻击音频大模型,暴露其安全漏洞。

Hidden in the Noise: Unveiling Backdoors in Audio LLMs Alignment through Latent Acoustic Pattern Triggers

  • 通过修改音频时序与频谱噪声,植入隐蔽声学特征
  • 90%以上攻击成功率,且模型对音量变化不敏感
  • 适合关注音频模型安全的开发者与研究者

随着音频大语言模型(ALLMs)在语音处理中的兴起,其安全性问题亟需关注。尽管文本与视觉安全已有广泛研究,但音频的独特特性带来了显著挑战。本文首次探究:ALLM是否易受基于声学触发器的后门攻击?为此,我们提出隐藏在噪声中(HIN)的新型后门攻击框架,通过修改原始音频波形中的时序动态和注入特定频谱噪声,引入被音频特征编码器捕获的稳定模式,从而在音频流中嵌入鲁棒触发信号。为评估ALLM对声学特征触发器的鲁棒性,我们构建了AudioSafe基准,涵盖九类风险类型。在AudioSafe及三个现有安全数据集上的大量实验表明,现有ALLMs存在严重漏洞:(I)环境噪声、语速变化等音频特征可实现超过90%的平均攻击成功率;(II)ALLM对不同声学特征敏感度差异显著,尤其对音量变化几乎无响应;(III)中毒样本仅引起微小损失曲线波动,凸显攻击的高度隐蔽性。

原文摘要 · Abstract (English)

As Audio Large Language Models (ALLMs) emerge as powerful tools for speech processing, their safety implications demand urgent attention. While considerable research has explored textual and vision safety, audio's distinct characteristics present significant challenges. This paper first investigates: Is ALLM vulnerable to backdoor attacks exploiting acoustic triggers? In response to this issue, we introduce Hidden in the Noise (HIN), a novel backdoor attack framework designed to exploit subtle, audio-specific features. HIN applies acoustic modifications to raw audio waveforms, such as alterations to temporal dynamics and strategic injection of spectrally tailored noise. These changes introduce consistent patterns that an ALLM's acoustic feature encoder captures, embedding robust triggers within the audio stream. To evaluate ALLM robustness against audio-feature-based triggers, we develop the AudioSafe benchmark, assessing nine distinct risk types. Extensive experiments on AudioSafe and three established safety datasets reveal critical vulnerabilities in existing ALLMs: (I) audio features like environment noise and speech rate variations achieve over 90% average attack success rate. (II) ALLMs exhibit significant sensitivity differences across acoustic features, particularly showing minimal response to volume as a trigger, and (III) poisoned sample inclusion causes only marginal loss curve fluctuations, highlighting the attack's stealth.

音频安全后门攻击声学特征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。