arXiv:2505.18864cs.CL2025-05被引 6

针对语音模型设计新攻击,能绕过安全防护生成违规内容。

Audio Jailbreak Attacks: Exposing Vulnerabilities in SpeechGPT in a White-Box Framework

  • 通过语音分词层生成对抗性令牌序列,实现精准攻击。
  • 在SpeechGPT上达成89%攻击成功率,显著优于现有方法。
  • 适合研究语音安全与多模态模型防御的学者参考。

多模态大语言模型(MLLMs)通过融合文本、视觉和音频模态,显著提升了人机交互的自然性与灵活性。其中,语音驱动模型如SpeechGPT在实用性上取得显著进展,可实现富有表现力且情绪响应灵敏的交互,增强真实场景中的沟通深度。然而,语音特性(如语速、发音差异、语音转文字转换)也带来了新的安全风险:攻击者可利用这些特征,构造绕过防御机制的输入,其攻击方式远超传统文本类越狱。尽管已有大量文本越狱研究,语音模态的攻击策略与防御仍严重不足。本文在白盒环境下提出一种针对对齐式MLLM语音输入的新对抗攻击方法。我们引入一种基于令牌级别的攻击,利用对模型语音分词过程的访问权限,生成对抗性令牌序列,并将其合成音频提示。该方法可有效绕过对齐防护机制,诱导模型输出禁止内容。在SpeechGPT上评估显示,该方法在多个受限任务中达到最高89%的攻击成功率,显著超越现有语音越狱手段。研究揭示了语音增强型多模态系统的脆弱性,为下一代更鲁棒的MLLM发展提供重要指导。

原文摘要 · Abstract (English)

Recent advances in Multimodal Large Language Models (MLLMs) have significantly enhanced the naturalness and flexibility of human computer interaction by enabling seamless understanding across text, vision, and audio modalities. Among these, voice enabled models such as SpeechGPT have demonstrated considerable improvements in usability, offering expressive, and emotionally responsive interactions that foster deeper connections in real world communication scenarios. However, the use of voice introduces new security risks, as attackers can exploit the unique characteristics of spoken language, such as timing, pronunciation variability, and speech to text translation, to craft inputs that bypass defenses in ways not seen in text-based systems. Despite substantial research on text based jailbreaks, the voice modality remains largely underexplored in terms of both attack strategies and defense mechanisms. In this work, we present an adversarial attack targeting the speech input of aligned MLLMs in a white box scenario. Specifically, we introduce a novel token level attack that leverages access to the model's speech tokenization to generate adversarial token sequences. These sequences are then synthesized into audio prompts, which effectively bypass alignment safeguards and to induce prohibited outputs. Evaluated on SpeechGPT, our approach achieves up to 89 percent attack success rate across multiple restricted tasks, significantly outperforming existing voice based jailbreak methods. Our findings shed light on the vulnerabilities of voice-enabled multimodal systems and to help guide the development of more robust next-generation MLLMs.

语音安全对抗攻击多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。