用几乎听不到的音频指令劫持大模型,让语音助手干坏事
Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection

- 设计可无视上下文、无声无息的恶意音频生成框架
- 在13个主流模型上实现79%~96%成功率,且对未知场景有效
- 适合安全研究者与语音系统开发者关注
现代大型音频-语言模型(LALMs)通过深度融合音频与文本实现智能语音交互,但这一融合也扩大了攻击面,尤其在连续高维音频通道中引入新漏洞。现有研究多聚焦文本越狱,而对仅通过音频注入实施恶意行为操纵的风险仍被忽视。本文揭示了一种新型威胁——感知隐蔽的音频提示劫持,并提出通用框架AudioHijack,可在仅访问音频数据且需高度隐蔽的条件下生成对抗性音频。该框架采用基于采样的梯度估计实现端到端优化,绕过非可微音频分词;通过注意力监督与多上下文训练,引导模型关注恶意音频并泛化至未见用户场景;还设计卷积混合方法,将扰动融入自然混响,极大降低可察觉性。在13个先进LALMs上的实验显示,该方法在6类错误行为中均表现一致,对未见过的用户上下文平均成功率达79%-96%,同时保持高声学保真度。真实世界测试表明,Mistral AI和Microsoft Azure的商用语音代理均可被诱导执行未经授权操作。研究暴露了LALMs的关键安全缺陷,亟需针对性防御。
原文摘要 · Abstract (English)
Modern Large audio-language models (LALMs) power intelligent voice interactions by tightly integrating audio and text. This integration, however, expands the attack surface beyond text and introduces vulnerabilities in the continuous, high-dimensional audio channel. While prior work studied audio jailbreaks, the security risks of malicious audio injection and downstream behavior manipulation remain underexamined. In this work, we reveal a previously overlooked threat, auditory prompt injection, under realistic constraints of audio data-only access and strong perceptual stealth. To systematically analyze this threat, we propose \textit{AudioHijack}, a general framework that generates context-agnostic and imperceptible adversarial audio to hijack LALMs. \textit{AudioHijack} employs sampling-based gradient estimation for end-to-end optimization across diverse models, bypassing non-differentiable audio tokenization. Through attention supervision and multi-context training, it steers model attention toward adversarial audio and generalizes to unseen user contexts. We also design a convolutional blending method that modulates perturbations into natural reverberation, making them highly imperceptible to users. Extensive experiments on 13 state-of-the-art LALMs show consistent hijacking across 6 misbehavior categories, achieving average success rates of 79\%-96\% on unseen user contexts with high acoustic fidelity. Real-world studies demonstrate that commercial voice agents from Mistral AI and Microsoft Azure can be induced to execute unauthorized actions on behalf of users. These findings expose critical vulnerabilities in LALMs and highlight the urgent need for dedicated defense.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。