arXiv:2505.15406cs.SDcs.AI2025-05被引 22

首个专用于评估大音频模型越狱漏洞的基准测试,揭示当前模型安全性严重不足。

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models

  • 构建1495个对抗性音频提示,覆盖10类违规内容,模拟真实越狱攻击
  • 引入动态扰动工具,微小改动即让顶尖模型安全性能大幅下降
  • 适合关注AI安全、音频生成与防御机制的研究者和开发者

大型音频语言模型(LAMs)的兴起带来了潜在价值与风险,其音频输出可能包含有害或不道德内容。然而,现有研究缺乏对LAM安全性的系统性、量化评估,尤其是针对越狱攻击——这因语音的时间性和语义性而极具挑战。为填补这一空白,我们提出AJailBench,首个专门用于评估LAM越狱漏洞的基准。首先构建AJailBench-Base数据集,包含1,495个对抗性音频提示,涵盖10类政策违规内容,由文本越狱攻击通过逼真语音合成转换而来。基于该数据集,我们评估多个先进LAMs,发现它们在各类攻击下均无一致鲁棒性。为进一步强化测试并模拟更真实攻击场景,我们提出音频扰动工具(APT),在时域、频域、幅域施加定向扰动,同时保持原始越狱意图的语义一致性,并采用贝叶斯优化高效搜索隐蔽且高效的扰动,形成扩展数据集AJailBench-APT。结果表明,即使微小且语义保持的扰动,也能显著削弱领先LAMs的安全性能,凸显亟需更强大、语义感知的防御机制。

原文摘要 · Abstract (English)

The rise of Large Audio Language Models (LAMs) brings both potential and risks, as their audio outputs may contain harmful or unethical content. However, current research lacks a systematic, quantitative evaluation of LAM safety especially against jailbreak attacks, which are challenging due to the temporal and semantic nature of speech. To bridge this gap, we introduce AJailBench, the first benchmark specifically designed to evaluate jailbreak vulnerabilities in LAMs. We begin by constructing AJailBench-Base, a dataset of 1,495 adversarial audio prompts spanning 10 policy-violating categories, converted from textual jailbreak attacks using realistic text to speech synthesis. Using this dataset, we evaluate several state-of-the-art LAMs and reveal that none exhibit consistent robustness across attacks. To further strengthen jailbreak testing and simulate more realistic attack conditions, we propose a method to generate dynamic adversarial variants. Our Audio Perturbation Toolkit (APT) applies targeted distortions across time, frequency, and amplitude domains. To preserve the original jailbreak intent, we enforce a semantic consistency constraint and employ Bayesian optimization to efficiently search for perturbations that are both subtle and highly effective. This results in AJailBench-APT, an extended dataset of optimized adversarial audio samples. Our findings demonstrate that even small, semantically preserved perturbations can significantly reduce the safety performance of leading LAMs, underscoring the need for more robust and semantically aware defense mechanisms.

音频安全越狱攻击对抗样本模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。