深度影响音频编码器抗攻击能力,中等深度最稳。
Trade-offs Between Capacity and Robustness in Neural Audio Codecs for Adversarially Robust Speech Recognition
- 用残差向量量化深度调控编码粒度,调节抗扰性
- 中等量化深度下误识别率最低,平衡内容与抗扰
- 代码本变化量与识别错误强相关,适合对抗训练场景
对抗扰动会利用自动语音识别(ASR)系统的漏洞,同时保持人类感知的语言内容。神经音频编码器通过离散瓶颈抑制与对抗噪声相关的细微信号变化。我们研究了该瓶颈的粒度(由残差向量量化,RVQ,深度控制)如何影响对抗鲁棒性。在梯度攻击下观察到非单调权衡:浅层量化抑制对抗扰动但破坏语音内容,深层量化则同时保留内容和扰动。中等深度平衡二者,使转录误差最小。我们进一步发现,对抗诱导的离散码本词元变化与转录错误高度相关。这些优势在自适应攻击下依然存在,神经编码器配置优于传统压缩防御方法。
原文摘要 · Abstract (English)
Adversarial perturbations exploit vulnerabilities in automatic speech recognition (ASR) systems while preserving human perceived linguistic content. Neural audio codecs impose a discrete bottleneck that can suppress fine-grained signal variations associated with adversarial noise. We examine how the granularity of this bottleneck, controlled by residual vector quantization (RVQ) depth, shapes adversarial robustness. We observe a non-monotonic trade-off under gradient-based attacks: shallow quantization suppresses adversarial perturbations but degrades speech content, while deeper quantization preserves both content and perturbations. Intermediate depths balance these effects and minimize transcription error. We further show that adversarially induced changes in discrete codebook tokens strongly correlate with transcription error. These gains persist under adaptive attacks, where neural codec configurations outperform traditional compression defenses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。