arXiv:2412.08608cs.SDcs.AI2024-12ICLR被引 40

首个针对大音频语言模型的隐蔽越狱攻击框架,突破梯度破碎难题。

AdvWave: Stealthy Adversarial Jailbreak Attack against Large Audio-Language Models

  • 分两阶段优化,解决音频编码器梯度破碎问题
  • 自适应目标搜索算法,提升越狱成功率40%
  • 用城市环境音生成噪声,保持听觉自然性

大型音频语言模型(LALMs)的兴起使语音交互成为可能,显著提升了用户体验并加速了实际应用部署。然而,确保其安全性至关重要,以避免引发社会关注或违反AI监管的风险。尽管如此,由于这类模型较新且相比传统DNN音频模型面临更多技术挑战,针对其越狱攻击的研究仍十分有限。具体而言,LALMs中的音频编码器涉及离散化操作,常导致梯度破碎,阻碍基于梯度的攻击有效性;模型行为的变异性也增加了有效攻击目标的识别难度;同时,对对抗音频波形的隐蔽性要求进一步压缩了可行解空间,使优化过程更加困难。为此,我们提出AdvWave——首个针对LALMs的越狱攻击框架。通过双阶段优化方法缓解梯度破碎,实现有效的端到端梯度优化;设计自适应对抗目标搜索算法,根据模型对特定查询的响应模式动态调整攻击目标;为保障对抗音频对人类听觉的自然性,采用分类器引导的优化策略,生成类似常见城市声音的对抗噪声。在多个先进LALMs上的广泛评估表明,AdvWave优于基线方法,平均越狱成功率提升40%。

原文摘要 · Abstract (English)

Recent advancements in large audio-language models (LALMs) have enabled speech-based user interactions, significantly enhancing user experience and accelerating the deployment of LALMs in real-world applications. However, ensuring the safety of LALMs is crucial to prevent risky outputs that may raise societal concerns or violate AI regulations. Despite the importance of this issue, research on jailbreaking LALMs remains limited due to their recent emergence and the additional technical challenges they present compared to attacks on DNN-based audio models. Specifically, the audio encoders in LALMs, which involve discretization operations, often lead to gradient shattering, hindering the effectiveness of attacks relying on gradient-based optimizations. The behavioral variability of LALMs further complicates the identification of effective (adversarial) optimization targets. Moreover, enforcing stealthiness constraints on adversarial audio waveforms introduces a reduced, non-convex feasible solution space, further intensifying the challenges of the optimization process. To overcome these challenges, we develop AdvWave, the first jailbreak framework against LALMs. We propose a dual-phase optimization method that addresses gradient shattering, enabling effective end-to-end gradient-based optimization. Additionally, we develop an adaptive adversarial target search algorithm that dynamically adjusts the adversarial optimization target based on the response patterns of LALMs for specific queries. To ensure that adversarial audio remains perceptually natural to human listeners, we design a classifier-guided optimization approach that generates adversarial noise resembling common urban sounds. Extensive evaluations on multiple advanced LALMs demonstrate that AdvWave outperforms baseline methods, achieving a 40% higher average jailbreak attack success rate.

音频安全越狱攻击对抗样本大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。