测试发现音频大模型安全防护远不如文本,易被语音攻击突破。
Audio Is the Achilles' Heel: Red Teaming Audio Large Multimodal Models
- 用三类攻击测试五款音频多模态模型,包括语音与文本混合攻击。
- 开源模型在有害语音问题上平均攻击成功率高达69.14%。
- 非语音噪音干扰可降低模型安全判断,适合安全研究者参考。
大型多模态模型(LMMs)通过结合大语言模型(LLMs)与模态编码器,实现视觉与听觉信息与文本的对齐,在真实场景中与人类交互。然而,这类模型是否在文本安全对齐的同时也具备一致的多模态输入防护能力,成为新的安全挑战。尽管视觉类LMMs的安全对齐研究已有进展,音频类模型的安全性仍缺乏系统探索。本文针对五款先进音频LMMs,在三种情境下进行全面红队测试:(i) 音频与文本双格式的有害提问;(ii) 文本格式的有害提问搭配非语音干扰音频;(iii) 针对语音的特定越狱攻击。结果显示,开源音频LMMs在有害音频提问上的平均攻击成功率达69.14%,且在非语音音频干扰下表现出安全漏洞。对Gemini-1.5-Pro的语音越狱攻击在有害查询基准上达到70.67%的成功率。本文揭示了这些安全错位可能的原因。警告:本文包含不当示例。
原文摘要 · Abstract (English)
Large Multimodal Models (LMMs) have demonstrated the ability to interact with humans under real-world conditions by combining Large Language Models (LLMs) and modality encoders to align multimodal information (visual and auditory) with text. However, such models raise new safety challenges of whether models that are safety-aligned on text also exhibit consistent safeguards for multimodal inputs. Despite recent safety-alignment research on vision LMMs, the safety of audio LMMs remains under-explored. In this work, we comprehensively red team the safety of five advanced audio LMMs under three settings: (i) harmful questions in both audio and text formats, (ii) harmful questions in text format accompanied by distracting non-speech audio, and (iii) speech-specific jailbreaks. Our results under these settings demonstrate that open-source audio LMMs suffer an average attack success rate of 69.14% on harmful audio questions, and exhibit safety vulnerabilities when distracted with non-speech audio noise. Our speech-specific jailbreaks on Gemini-1.5-Pro achieve an attack success rate of 70.67% on the harmful query benchmark. We provide insights on what could cause these reported safety-misalignments. Warning: this paper contains offensive examples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。