arXiv:2502.00718cs.LGcs.SD2025-02被引 9

发现可绕过对齐机制的通用音频越狱攻击,且隐含第一人称有毒语义。

"I am bad": Interpreting Stealthy, Universal and Robust Audio Jailbreaks in Audio-Language Models

  • 构建跨提示、任务和音频样本的通用对抗扰动,实现隐蔽越狱。
  • 攻击在模拟真实场景下仍有效,且扰动中编码不可察觉的有毒语音特征。
  • 揭示多模态模型对音频输入的脆弱性,适合安全与防御研究者阅读。

多模态大语言模型的兴起带来了创新的人机交互范式,但也带来了机器学习安全的重大挑战。音频-语言模型(ALMs)因其自然的口语交流特性尤为关键,但其失效模式尚不明确。本文研究针对ALMs的音频越狱攻击,重点分析其绕过对齐机制的能力。我们构建了可跨提示、任务甚至基础音频样本泛化的对抗扰动,首次展示音频模态下的通用越狱攻击,并验证其在模拟真实世界条件下的有效性。除证明攻击可行性外,我们分析了ALMs如何解析这些音频对抗样本,发现其编码了不可察觉的第一人称有毒话语——表明最有效的扰动通过在音频信号中嵌入特定语言特征来诱发有害输出。该结果对理解多模态模型中不同模态的交互具有重要意义,并为增强对抗音频攻击的防御提供了可行洞见。

原文摘要 · Abstract (English)

The rise of multimodal large language models has introduced innovative human-machine interaction paradigms but also significant challenges in machine learning safety. Audio-Language Models (ALMs) are especially relevant due to the intuitive nature of spoken communication, yet little is known about their failure modes. This paper explores audio jailbreaks targeting ALMs, focusing on their ability to bypass alignment mechanisms. We construct adversarial perturbations that generalize across prompts, tasks, and even base audio samples, demonstrating the first universal jailbreaks in the audio modality, and show that these remain effective in simulated real-world conditions. Beyond demonstrating attack feasibility, we analyze how ALMs interpret these audio adversarial examples and reveal them to encode imperceptible first-person toxic speech - suggesting that the most effective perturbations for eliciting toxic outputs specifically embed linguistic features within the audio signal. These results have important implications for understanding the interactions between different modalities in multimodal models, and offer actionable insights for enhancing defenses against adversarial audio attacks.

音频攻击越狱多模态安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。