arXiv:2606.20893cs.SDcs.AI2026-06中稿 · Interspeech 2026

用音频编码器潜空间生成对抗音频,单次推理即可高成功率攻击语音识别系统。

Exploiting Neural Audio Codec Latents for Adversarial Audio Attacks

论文配图:Exploiting Neural Audio Codec Latents for Adversarial Audio Attacks
图 1 · 摘自论文原文
  • 在神经音频编码器的连续潜空间生成对抗扰动,单次前向传播完成攻击
  • 目标攻击成功率最高达99%,推理延迟低于7毫秒
  • 适合需要低延迟、高隐蔽性的实时对抗攻击场景

基于深度学习的语音分类系统(如说话人验证)易受对抗攻击。现有优化类方法(如PGD、Carlini-Wagner)需在高维波形域进行多轮迭代,成本高昂;生成式攻击虽可单次合成,但常引入可感知伪影或依赖计算量大的模型,扩散与自回归方法则存在高推理延迟。为弥补这一差距,本文提出一种在神经音频编码器潜空间中运行的生成式攻击框架。通过条件生成器在单次前向传播中合成特定类别扰动,并解码为对抗波形。该方法在目标攻击下成功率最高达99%,推理延迟低于7毫秒,相比生成基线提升24倍效率。

原文摘要 · Abstract (English)

Deep learning-based audio classification systems, including automatic speaker verification, are vulnerable to adversarial attacks. Realistic real-time threat assessment remains difficult because optimization-based methods, such as projected gradient descent (PGD) and Carlini-Wagner, require costly iterative updates in the high-dimensional waveform domain. Generative attacks allow single-shot synthesis but often introduce perceptible artifacts or depend on computationally intensive architectures, while diffusion and autoregressive approaches incur high inference latency. To address this gap, we propose a generative attack framework operating in the continuous latent space of a neural audio codec. A conditional generator synthesizes class-specific perturbations in a single forward pass and decodes them into adversarial waveforms. Our method achieves targeted attack success rates up to 99% with sub-7 ms inference, outperforming generative baselines while reducing latency by 24x.

对抗攻击音频安全潜空间低延迟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。