测试多语言语音中大模型的安全漏洞,发现混语场景下攻击成功率极高。
SpeechJBB: Probing Safety Alignment and Comprehension in Large Audio Language Models under Code-Switched Speech

- 构建多语言语音对抗数据集SpeechJBB,含混语变体和伪装词
- 非英语混语语音攻击成功率最高,伪词插入使拒绝率持续下降
- 模型理解力强但仍易被攻破,适合安全与语音研究者参考
大型音频语言模型(LALMs)在真实场景中应用日益广泛,但其安全性评估仍主要基于单语文本有害指令。这使得它们在多语言及口语场景,尤其是混语语音下的泛化能力未被充分探索。为此,我们提出SpeechJBB,一个用于评估前沿LALMs的音频越狱数据集,涵盖英语、法语、德语、意大利语和西班牙语五种欧洲语言及其两两组合的混语变体。通过在关键安全词汇周围插入音素上合理的伪词,模拟局部混淆,进一步探测安全弱点。实验显示,混语有害音频导致显著高越狱成功率(JSR),非英语单语及非英语混语对的攻击成功率最高。伪词插入密度越高,模型拒绝率越低,即使模型极少将这些插入词视为有害含义。理解能力测试表明,这些失败并非源于多语言误解,因为部分具备强语音识别、口语理解和推理能力的模型反而最易被攻破。
原文摘要 · Abstract (English)
Large audio language models (LALMs) are increasingly deployed in real-world applications, yet their safety alignment is still primarily evaluated on monolingual, text-based harmful prompts. This leaves their generalizability under multilingual and spoken settings, particularly code-switched speech, largely underexplored. To address this gap, we introduce SpeechJBB, an audio jailbreak dataset for benchmarking state-of-the-art LALMs across five European languages: English, French, German, Italian, and Spanish, as well as code-switched variants combining pairs of these languages. The extent of safety weaknesses is further probed by introducing an augmented setting where phonologically plausible pseudo-words are inserted around safety-critical terms to simulate localized obfuscation. Across models, code-switched harmful audio yields substantially high jailbreak success rates (JSR), with non-English monolingual and non-English code-switched pairs exhibiting the highest attack success. Pseudo-word insertion monotonically reduces refusal as insertion density increases, even though models rarely attribute harmful meaning to the inserted tokens. Comprehension benchmarks show that these failures are not reducible to multilingual misunderstanding, as several models with strong ASR, spoken language understanding, and spoken reasoning performance are among the most vulnerable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。