arXiv:2505.14286cs.CLcs.SD2025-05EMNLP被引 5

用一段固定音频可操控语音大模型输出,且能精准触发特定条件。

Universal Acoustic Adversarial Attacks for Flexible Control of Speech-LLMs

  • 用一段通用噪声音频插入原语音,即可干扰模型输出。
  • 攻击可让模型完全无响应或执行被篡改的任务。
  • 仅在特定说话人性别或语言时生效,实现精准控制。

将预训练语音编码器与大语言模型结合,催生了可处理多种语音任务的语音大模型。然而其灵活性也带来了新风险:本文研究针对语音大模型的通用声学对抗攻击。通过在原始音频前添加一段固定的、通用的对抗性音频片段,可实现对模型行为的干扰。初始攻击使模型无输出或执行被篡改的任务;进一步扩展为选择性攻击,仅在输入包含特定属性(如说话人性别或语言)时触发,其他输入不受影响,从而实现细粒度控制。实验发现,Qwen2-Audio 和 Granite-Speech 存在严重漏洞,表明类似语音大模型可能普遍易受此类攻击,亟需更强的鲁棒训练策略与抗攻击能力。

原文摘要 · Abstract (English)

The combination of pre-trained speech encoders with large language models has enabled the development of speech LLMs that can handle a wide range of spoken language processing tasks. While these models are powerful and flexible, this very flexibility may make them more vulnerable to adversarial attacks. To examine the extent of this problem, in this work we investigate universal acoustic adversarial attacks on speech LLMs. Here a fixed, universal, adversarial audio segment is prepended to the original input audio. We initially investigate attacks that cause the model to either produce no output or to perform a modified task overriding the original prompt. We then extend the nature of the attack to be selective so that it activates only when specific input attributes, such as a speaker gender or spoken language, are present. Inputs without the targeted attribute should be unaffected, allowing fine-grained control over the model outputs. Our findings reveal critical vulnerabilities in Qwen2-Audio and Granite-Speech and suggest that similar speech LLMs may be susceptible to universal adversarial attacks. This highlights the need for more robust training strategies and improved resistance to adversarial attacks.

语音大模型对抗攻击安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。