arXiv:2511.10222cs.SDcs.AI2025-11被引 13

用语音音频组合攻击多模态大模型,暴露安全漏洞并提出新防御方案。

Speech-Audio Compositional Attacks on Multimodal LLMs and Their Mitigation with SALMONN-Guard

  • 通过语音与音频混合构造黑盒攻击,模拟真实场景中的有害内容共存
  • 即使顶级模型Gemini 2.5 Pro在攻击下仍有66%成功率,验证防护短板
  • 提出SALMONN-Guard联合分析语音、音频和文本,将攻击成功率降至20%

近年来,大语言模型对音频信号的理解能力提升,但也暴露出复杂音频输入带来的新型安全风险,现有防护机制难以应对。本文提出SACRED-Bench(Speech-Audio Composition for RED-teaming),用于评估多模态大模型在复杂音频攻击下的鲁棒性。不同于依赖噪声优化或白盒访问的已有方法,SACRED-Bench利用语音-音频组合实现有效黑盒攻击,包含三种机制:(a) 有害与无害语音重叠,(b) 无害语音与有害非语音音频混合,(c) 多说话人对话。这些设置聚焦于良性与有害意图共存的听觉场景,且问题设计隐含音频内容引用,使文本提示本身不显含有害信息。实验表明,即便启用完整安全防护的Gemini 2.5 Pro,攻击成功率仍达66%。为此,本文提出SALMONN-Guard,首个联合审查语音、音频与文本的安全判断模型,将攻击成功率降至20%。结果凸显音频感知防御对保障多模态大模型安全的重要性。数据集与模型检查点详见https://huggingface.co/datasets/tsinghua-ee/SACRED-Bench。

原文摘要 · Abstract (English)

Recent progress in LLMs has enabled understanding of audio signals, but has also exposed new safety risks arising from complex audio inputs that are inadequately handled by current safeguards. We introduce SACRED-Bench (Speech-Audio Composition for RED-teaming) to evaluate the robustness of LLMs under complex audio-based attacks. Unlike existing perturbation-based methods that rely on noise optimization or white-box access, SACRED-Bench exploits speech-audio composition to enable effective black-box attacks. SACRED-Bench adopts three composition mechanisms: (a) overlap of harmful and benign speech, (b) mixture of benign speech with harmful non-speech audio, and (c) multi-speaker dialogue. These mechanisms focus on evaluating safety in settings where benign and harmful intents co-occur within a single auditory scene. Moreover, questions in SACRED-Bench are designed to implicitly refer to content in the audio, such that no explicit harmful information appears in the text prompt alone. Experiments demonstrate that even Gemini 2.5 Pro, a state-of-the-art proprietary LLM with safety guardrails fully enabled, still exhibits a 66% attack success rate. To bridge this gap, we propose SALMONN-Guard, the first guard model that jointly inspects speech, audio, and text for safety judgments, reducing the attack success rate to 20%. Our results highlight the need for audio-aware defenses to ensure the safety of multimodal LLMs. The dataset and SALMONN-Guard checkpoints can be found at https://huggingface.co/datasets/tsinghua-ee/SACRED-Bench.

多模态安全语音攻击大模型防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。