arXiv:2505.17568cs.CRcs.AI2025-05被引 16

首个专用于评估音频大模型越狱漏洞的基准,覆盖超1000小时音频。

JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models

  • 构建包含1.1万文本和24.5万音频样本的越狱攻击评测集
  • 发现文本安全对齐可部分迁移至音频输入,跨模态策略更稳健
  • 适合研究音频安全、模型鲁棒性及对抗防御的学者与工程师

大型音频语言模型(LALMs)已取得显著进展,但在实际应用中面临越狱攻击带来的安全风险,此类攻击可绕过安全对齐机制。然而,目前缺乏专门针对LALMs的对抗性音频数据集和统一评估框架。为此,我们提出JALMBench,一个全面的基准测试体系,用于评估LALM在越狱攻击下的安全性,包含11,316个文本样本和245,355个音频样本(总时长超过1,000小时)。该基准支持12种主流LALMs、8种攻击方法(4种文本转移型,4种音频原生型)以及5种防御策略。我们深入分析了攻击效率、话题敏感性、语音多样性及模型架构的影响。此外,探索了在提示和响应层面的缓解策略。系统性评估表明,LALMs的安全性强烈受模态和架构选择影响:基于文本的安全对齐可部分迁移到音频输入,而交错式音文策略能实现更强的跨模态泛化能力。现有通用内容审核方法仅略微提升安全性,凸显为LALMs量身定制防御机制的必要性。我们希望本工作能为构建更鲁棒的LALMs提供设计指导。

原文摘要 · Abstract (English)

Large Audio Language Models (LALMs) have made significant progress. While increasingly deployed in real-world applications, LALMs face growing safety risks from jailbreak attacks that bypass safety alignment. However, there remains a lack of an adversarial audio dataset and a unified framework specifically designed to evaluate and compare jailbreak attacks against them. To address this gap, we introduce JALMBench, a comprehensive benchmark that assesses LALM safety against jailbreak attacks, comprising 11,316 text samples and 245,355 audio samples (>1,000 hours). JALMBench supports 12 mainstream LALMs, 8 attack methods (4 text-transferred and 4 audio-originated), and 5 defenses. We conduct in-depth analysis on attack efficiency, topic sensitivity, voice diversity, and model architecture. Additionally, we explore mitigation strategies for the attacks at both the prompt and response levels. Our systematic evaluation reveals that LALMs' safety is strongly influenced by modality and architectural choices: text-based safety alignment can partially transfer to audio inputs, and interleaved audio-text strategies enable more robust cross-modal generalization. Existing general-purpose moderation methods only slightly improve security, highlighting the need for defense methods specifically designed for LALMs. We hope our work can shed light on the design principles for building more robust LALMs.

音频安全越狱攻击多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。