arXiv:2605.30031cs.SDcs.AI2026-05

揭示大音频模型的语音越狱风险,提出统一评估框架。

Audio Jailbreaks in Large Audio-Language Models: Taxonomy, Attack-Defense Analysis, and Cost-Aware Evaluation

论文配图:Audio Jailbreaks in Large Audio-Language Models: Taxonomy, Attack-Defense Analysis, and Cost-Aware Evaluation
图 1 · 摘自论文原文
  • 按语义、声学、信号等四类划分越狱攻击,分三类防御策略。
  • 声学类攻击成功率高,叙事诱导低延迟但隐蔽性强。
  • 现有防御降低鲁棒性且影响正常响应,需综合成本考量。

大型音频语言模型(LALMs)将越狱风险从文本层面扩展至完整的语音感知-推理流程,可通过语义、声学风格、信号伪影或内部表征诱发不安全行为。现有研究在异构威胁模型和评估协议下展开,难以比较攻击实用性或防御有效性。本文提出统一分类体系,并对十款开源LALM进行受控实证评估。将已有工作归纳为语义、声学、信号与嵌入层攻击;基于防护机制、无训练与训练型防御;以及跨模态、纯音频与交互式评测基准。评估涵盖攻击成功率、良性拒绝率及延迟。结果表明:声学最优-多选(Acoustic Best-of-N)暴露严重音频空间漏洞,叙事框架是高效低延迟语义威胁,当前防御在鲁棒性与可用性间存在权衡。研究支持以成本与效用为导向的评估,作为仅依赖成功率的安全基准必要补充。

原文摘要 · Abstract (English)

Large Audio Language Models (LALMs) expand jailbreak risks from token-level prompting to the full speech perception-to-reasoning pipeline, where unsafe behavior can be induced through semantics, acoustic style, signal artifacts, or internal representations. Existing work studies these risks under heterogeneous threat models and evaluation protocols, making it difficult to compare attack practicality or defense utility. This paper provides a unified taxonomy and a controlled empirical evaluation of LALM jailbreak attacks and defenses. We organize prior work into semantic, acoustic, signal, and embedding-layer attacks; guard-based, training-free, and training-based defenses; and cross-modal, audio-native, and interactive benchmarks. We then evaluate representative attacks and defenses across ten open-source LALMs, measuring not only attack success rate but also benign refusal and latency. Our results show that Acoustic Best-of-N reveals strong worst-case audio-space vulnerabilities, Narrative Framing is an effective low-latency semantic threat, and current defenses trade robustness against benign usability. These findings support cost- and utility-aware evaluation as a necessary complement to success-rate-only LALM safety benchmarks.

音频安全越狱攻击大模型评测防御机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。