arXiv:2504.01094cs.SDcs.AI2025-04被引 31

多语言多口音音频攻击让语音大模型漏洞扩大,成功率最高提升57.25%。

Multilingual and Multi-Accent Jailbreaking of Audio LLMs

  • 构建首个多语言多口音对抗音频数据集,系统测试跨语言声学干扰攻击
  • 混响等声学扰动与跨语言发音结合,使攻击成功率最高提升57.25个百分点
  • 语音攻击比文本攻击成功率高3.1倍,揭示多模态模型的薄弱环节

大型音频语言模型(LALMs)虽显著提升音频理解能力,但引入严重安全风险,尤其是音频越狱攻击。现有研究集中于英语攻击,本文揭示更严重威胁:利用语言和声学差异的多语言多口音对抗攻击,能大幅提高攻击成功率。我们提出 Multi-AudioJail 框架,包含(1)首个对抗性多语言/多口音音频越狱提示数据集,(2)分层评估流程,显示声学扰动(如混响、回声、耳语效果)与跨语言语音特征相互作用,使越狱成功率(JSR)最高提升57.25个百分点(如对 MERaLiON 的肯尼亚口音混响攻击)。关键发现:多模态模型比单模态系统更易受攻——攻击者只需突破非英语音频输入这一弱环节,即可整体入侵模型,实证表明多语言音频攻击成功率是文本攻击的3.1倍。我们计划开源数据集,推动跨模态防御研究,呼吁社区关注随LALMs发展而扩大的攻击面。

原文摘要 · Abstract (English)

Large Audio Language Models (LALMs) have significantly advanced audio understanding but introduce critical security risks, particularly through audio jailbreaks. While prior work has focused on English-centric attacks, we expose a far more severe vulnerability: adversarial multilingual and multi-accent audio jailbreaks, where linguistic and acoustic variations dramatically amplify attack success. In this paper, we introduce Multi-AudioJail, the first systematic framework to exploit these vulnerabilities through (1) a novel dataset of adversarially perturbed multilingual/multi-accent audio jailbreaking prompts, and (2) a hierarchical evaluation pipeline revealing that how acoustic perturbations (e.g., reverberation, echo, and whisper effects) interacts with cross-lingual phonetics to cause jailbreak success rates (JSRs) to surge by up to +57.25 percentage points (e.g., reverberated Kenyan-accented attack on MERaLiON). Crucially, our work further reveals that multimodal LLMs are inherently more vulnerable than unimodal systems: attackers need only exploit the weakest link (e.g., non-English audio inputs) to compromise the entire model, which we empirically show by multilingual audio-only attacks achieving 3.1x higher success rates than text-only attacks. We plan to release our dataset to spur research into cross-modal defenses, urging the community to address this expanding attack surface in multimodality as LALMs evolve.

音频安全多语言越狱攻击多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。