arXiv:2603.13847cs.CRcs.AI2026-03

用近超声波无声攻击语音大模型,让其执行恶意指令。

Sirens' Whisper: Inaudible Near-Ultrasonic Jailbreaks of Speech-Driven LLMs

  • 将指令编码为近超声波,经麦克风非线性还原成可读语音。
  • 在商用模型上实现最高0.94的拒绝规避率和0.925的说服力。
  • 攻击对人耳完全不可见,适合研究语音接口安全漏洞。

语音驱动的大语言模型正通过语音接口广泛接入,开放的声学通道引入新型安全风险。我们提出Sirens' Whisper(SWhisper),首个在真实黑盒条件下使用消费级硬件实现隐蔽提示攻击的框架。SWhisper通过将目标基带音频(包括长而结构化的提示)编码为近超声波形,在声学传输与麦克风非线性作用后仍能精准解调,实现鲁棒的无声传递。该方法基于对设备与环境非线性特性的轻量建模及预补偿,构建高保真隐蔽信道。在此基础上,设计语音感知的越狱生成方法,确保指令在语音接口下的可理解性、简洁性与跨模型迁移性。实验覆盖商业与开源语音驱动LLM,均表现强大黑盒有效性:在商用模型上,最高达成0.94的非拒绝率(NR)与0.925的特定说服力(SC)。受控用户研究表明,注入的越狱音频对人类听者而言与纯背景音无异。尽管越狱为案例研究,但底层隐蔽声学信道可支持更广泛的高保真提示注入与命令执行攻击。

原文摘要 · Abstract (English)

Speech-driven large language models (LLMs) are increasingly accessed through speech interfaces, introducing new security risks via open acoustic channels. We present Sirens' Whisper (SWhisper), the first practical framework for covert prompt-based attacks against speech-driven LLMs under realistic black-box conditions using commodity hardware. SWhisper enables robust, inaudible delivery of arbitrary target baseband audio-including long and structured prompts-on commodity devices by encoding it into near-ultrasound waveforms that demodulate faithfully after acoustic transmission and microphone nonlinearity. This is achieved through a simple yet effective approach to modeling nonlinear channel characteristics across devices and environments, combined with lightweight channel-inversion pre-compensation. Building on this high-fidelity covert channel, we design a voice-aware jailbreak generation method that ensures intelligibility, brevity, and transferability under speech-driven interfaces. Experiments across both commercial and open-source speech-driven LLMs demonstrate strong black-box effectiveness. On commercial models, SWhisper achieves up to 0.94 non-refusal (NR) and 0.925 specific-convincing (SC). A controlled user study further shows that the injected jailbreak audio is perceptually indistinguishable from background-only playback for human listeners. Although jailbreaks serve as a case study, the underlying covert acoustic channel enables a broader class of high-fidelity prompt-injection and commandexecution attacks.

语音安全越狱攻击隐蔽信道大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。