用普通语音悄悄骗过语音模型,让它说出有害内容。
When Good Sounds Go Adversarial: Jailbreaking Audio-Language Models with Benign Inputs
- 分两阶段攻击:先诱使模型输出有害回复,再将恶意代码藏进正常语音中
- 在5个主流模型上成功率60%-78%,对天气查询等常见语音有效
- 揭示了语音交互中的隐蔽威胁,适合关注AI安全的研究者
随着大语言模型日益融入日常生活,语音成为人机交互的关键接口。然而,这种便利也带来了新漏洞,使语音可能成为攻击面。本研究提出WhisperInject,一种两阶段对抗性音频攻击框架,可操纵先进语音语言模型生成有害内容。方法通过在人类可理解的音频输入中嵌入微小扰动,将有害载荷隐藏其中。第一阶段采用基于奖励的白盒优化方法——强化学习与投影梯度下降(RL-PGD),诱导目标模型输出有害原生响应;第二阶段将该有害响应作为目标,利用梯度优化技术将细微扰动嵌入良性音频载体(如天气查询、问候语)。在两个基准测试和五个多模态大模型上,平均攻击成功率达60%-78%,并通过多种评估框架验证。本工作揭示了一类新型实用且隐蔽的音频原生威胁,从理论突破迈向现实可行的攻击手段。
原文摘要 · Abstract (English)
As large language models (LLMs) become increasingly integrated into daily life, audio has emerged as a key interface for human-AI interaction. However, this convenience also introduces new vulnerabilities, making audio a potential attack surface for adversaries. Our research introduces WhisperInject, a two-stage adversarial audio attack framework that manipulates state-of-the-art audio language models to generate harmful content. Our method embeds harmful payloads as subtle perturbations into audio inputs that remain intelligible to human listeners. The first stage uses a novel reward-based white-box optimization method, Reinforcement Learning with Projected Gradient Descent (RL-PGD), to jailbreak the target model and elicit harmful native responses. This native harmful response then serves as the target for Stage 2, Payload Injection, where we use gradient-based optimization to embed subtle perturbations into benign audio carriers, such as weather queries or greeting messages. Our method achieves average attack success rates of 60-78% across two benchmarks and five multimodal LLMs, validated by multiple evaluation frameworks. Our work demonstrates a new class of practical, audio-native threats, moving beyond theoretical exploits to reveal a feasible and covert method for manipulating multimodal AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。