通过优化潜在空间提升语音伪造检测的泛化能力
Generalizable Audio Deepfake Detection via Latent Space Refinement and Augmentation
- 在潜在空间中引入可学习原型,捕捉伪造语音的复杂差异
- 直接在潜在空间进行数据增强,扩展模型对新型攻击的识别能力
- 在多个数据集上表现优异,适合应对未知伪造攻击
语音合成技术(如文本转语音TTS和语音转换VC)的进步使语音伪造检测愈发困难。现有反伪造方法在面对未见攻击时泛化能力不足。为此,我们提出一种新策略,结合潜在空间精炼(LSR)与潜在空间增强(LSA),提升检测系统的泛化性能。LSR为伪造类引入多个可学习原型,优化潜在空间以更准确捕获伪造数据的复杂变化;LSA则在潜在空间中直接应用增强技术,进一步丰富伪造样本表示,使模型学习到更广泛的伪造模式。我们在ASVspoof 2019 LA、ASVspoof 2021 LA与DF、In-The-Wild四个代表性数据集上评估,结果表明LSR与LSA单独使用均有效,二者结合达到当前最优水平,甚至超越部分现有方法。
原文摘要 · Abstract (English)
Advances in speech synthesis technologies, like text-to-speech (TTS) and voice conversion (VC), have made detecting deepfake speech increasingly challenging. Spoofing countermeasures often struggle to generalize effectively, particularly when faced with unseen attacks. To address this, we propose a novel strategy that integrates Latent Space Refinement (LSR) and Latent Space Augmentation (LSA) to improve the generalization of deepfake detection systems. LSR introduces multiple learnable prototypes for the spoof class, refining the latent space to better capture the intricate variations within spoofed data. LSA further diversifies spoofed data representations by applying augmentation techniques directly in the latent space, enabling the model to learn a broader range of spoofing patterns. We evaluated our approach on four representative datasets, i.e. ASVspoof 2019 LA, ASVspoof 2021 LA and DF, and In-The-Wild. The results show that LSR and LSA perform well individually, and their integration achieves competitive results, matching or surpassing current state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。