arXiv:2508.19180eess.AScs.SD2025-08中稿 · APSIPA ASC 2025被引 1

用掩码扩散模型检测并净化语音验证中的对抗攻击

MDD: a Mask Diffusion Detector to Protect Speaker Verification Systems from Adversarial Perturbations

  • 基于文本条件掩码扩散模型,无需对抗样本即可训练
  • 在多个数据集上检测准确率超现有方法,净化后识别性能接近干净状态
  • 适合需要高安全性的语音验证系统部署,尤其对无对抗样本场景有效

语音验证系统在安防应用中日益普及,但极易受到对抗扰动的影响。本文提出一种新型对抗检测与净化框架——掩码扩散检测器(MDD),基于文本条件掩码扩散模型。训练时,MDD对梅尔频谱图进行部分掩码,并通过前向扩散过程逐步添加噪声,模拟语音特征的退化;逆过程则在输入文本条件下重建原始清晰表示。与以往方法不同,MDD无需对抗样本或大规模预训练。实验表明,MDD在对抗检测方面表现优异,优于现有扩散模型及神经编码器方法;同时能有效净化被恶意篡改的语音信号,使语音验证性能恢复至接近干净条件水平。这些结果证明了基于扩散的掩码策略在构建安全可靠的语音验证系统中的潜力。

原文摘要 · Abstract (English)

Speaker verification systems are increasingly deployed in security-sensitive applications but remain highly vulnerable to adversarial perturbations. In this work, we propose the Mask Diffusion Detector (MDD), a novel adversarial detection and purification framework based on a \textit{text-conditioned masked diffusion model}. During training, MDD applies partial masking to Mel-spectrograms and progressively adds noise through a forward diffusion process, simulating the degradation of clean speech features. A reverse process then reconstructs the clean representation conditioned on the input transcription. Unlike prior approaches, MDD does not require adversarial examples or large-scale pretraining. Experimental results show that MDD achieves strong adversarial detection performance and outperforms prior state-of-the-art methods, including both diffusion-based and neural codec-based approaches. Furthermore, MDD effectively purifies adversarially-manipulated speech, restoring speaker verification performance to levels close to those observed under clean conditions. These findings demonstrate the potential of diffusion-based masking strategies for secure and reliable speaker verification systems.

语音验证对抗攻击扩散模型安全检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。