arXiv:2607.00899eess.AS2026-07

提出新型噪声增强方法,提升语音验证系统抗攻击能力

Positive-Incentive Noise Predictor for Adversarial Purification in Speaker Verification

论文配图:Positive-Incentive Noise Predictor for Adversarial Purification in Speaker Verification
图 1 · 摘自论文原文
  • 将净化任务重构为可学习的加噪过程,引入正向激励噪声
  • 在四种主流模型上实现强防御,实时因子低至0.014
  • 适合需高效防护的真实语音验证场景

现代自动说话人验证(ASV)系统易受对抗扰动影响。基于扩散模型的净化方法虽有效,但逆向去噪需迭代采样,导致推理延迟高。我们发现前向加噪过程贡献了主要的鲁棒性提升。受此启发,我们将对抗净化重构为可学习的加噪问题,提出首个显式引入正向激励噪声(π-noise)的框架——正向激励噪声预测器(PnP)。PnP学习输入自适应的π-noise,并将其与输入混合以增强下游ASV系统的鲁棒性。在四个先进ASV主干网络上的实验表明,PnP能有效防御对抗攻击,同时保持自然语音性能。相比代表性净化基线,该框架在白盒、黑盒及防御者感知自适应攻击下,实现了防御效果、对真实语音影响和推理效率的优异平衡,实时因子低至0.014。此外,PnP可级联扩散去噪器,进一步提升净化语音的听觉质量。代码与净化音频示例见https://eurecom-asp.github.io/pnp/

原文摘要 · Abstract (English)

Modern automatic speaker verification (ASV) systems are vulnerable to adversarial perturbations. Diffusion-based purification has recently shown strong effectiveness against such perturbations, but its reverse denoising process requires iterative sampling and leads to high inference latency. We find that the forward noising process provides most of the robustness gain. Motivated by this observation, we reformulate adversarial purification as a learnable noising problem, and propose the Positive-Incentive Noise Predictor (PnP), the first framework that explicitly introduces positive-incentive noise (π-noise) into the purification task. PnP learns input-adaptive π-noise and mixes it with the input to improve the robustness of downstream ASV systems. Experiments on four advanced ASV backbones show that PnP effectively defends against adversarial attacks while preserving performance on natural speech. Compared with representative purification baselines, the proposed framework provides a competitive balance among defense effectiveness, impact on genuine utterances, and inference efficiency under white-box, black-box, and defender-aware adaptive attacks, with a real-time factor as low as 0.014. Moreover, PnP can be cascaded with a diffusion denoiser to further improve the perceptual quality of purified utterances. Code and purified audio examples are available at https://eurecom-asp.github.io/pnp/

语音安全对抗防御噪声注入实时优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。