arXiv:2608.10405cs.SDcs.AI2026-08

通过微小声波扰动让语音大模型无限输出,耗尽计算资源。

Never Stop Speaking: a Denial-of-Service Attack on End-to-End Speech Language Models

论文配图:Never Stop Speaking: a Denial-of-Service Attack on End-to-End Speech Language Models
图 1 · 摘自论文原文
  • 用不可察觉的声波扰动直接干扰模型生成过程。
  • 攻击后生成长度提升超10倍,显存占用增长近6倍。
  • 适合研究语音模型安全或对抗攻击的学者使用。

大量研究表明,精心构造的输入可导致大型语言模型生成过长输出,造成显著计算开销与资源消耗。尽管现有拒绝服务(DoS)攻击多针对纯文本大模型,端到端(E2E)语音大模型正快速兴起。现有文本类DoS攻击依赖提示工程(如对抗后缀或语义诱导),利用文本离散特性,无法直接迁移至连续语音输入。此外,先前语音模型安全研究主要聚焦于语音识别(ASR)或文本转语音(TTS)系统,而对E2E语音大模型的DoS漏洞研究尚属空白。为填补此空白,我们提出基于扰动的DoS攻击方法,针对E2E语音模型。该方法不通过提示操纵诱导长输出,而是优化不可察觉的声学扰动,直接干预模型自回归生成过程,同时保持原始输入长度不变。具体地,将攻击建模为复合优化目标,联合抑制结束符(EOS)生成、延长解码过程,并通过加权EOS损失、top-k logits损失、长度损失及语义一致性损失保持语义连贯性。为进一步增强隐蔽性,采用语音活动检测(VAD)仅在发声区域注入扰动。在三个开源E2E语音大模型上的广泛实验表明,本方法实现了稳定的攻击成功率,生成长度显著增加,显存消耗提升近6倍,揭示了现代语音大模型的安全风险。

原文摘要 · Abstract (English)

Many studies have shown that specially crafted inputs can induce large language models (LLMs) to generate excessively long outputs, resulting in significant computational overhead and resource consumption. While most existing denial-of-service (DoS) attacks target text-only LLMs, end-to-end (E2E) speech LLMs are rapidly emerging. Existing text-based DoS attacks primarily rely on prompt engineering, such as adversarial suffixes or semantic inducement, which exploit the discrete nature of text inputs and therefore cannot be directly transferred to continuous speech inputs. Moreover, prior studies on speech model security mainly focus on ASR or TTS systems, leaving the DoS vulnerability of E2E speech LLMs largely unexplored. To address this gap, we propose the perturbation-based DoS attack targeting E2E speech models. Instead of inducing long outputs through prompt manipulation, our method optimizes imperceptible acoustic perturbations to directly influence the model's autoregressive generation process while preserving the original input length. Specifically, we formulate the attack as a composite optimization objective that jointly suppresses EOS generation, encourages prolonged decoding, and largely preserves semantic consistency by integrating weighted EOS loss, top-k logit loss, length loss, and semantic alignment loss. To further improve stealthiness, we employ voice activity detection (VAD) to inject perturbations only into voiced regions. Extensive experiments on three open-source E2E speech LLMs demonstrate that our method achieves stable attack success rate while significantly increasing generation length and GPU resource consumption, revealing security risks in modern ALLMs.

语音安全对抗攻击大模型漏洞

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。