通过双重掩码增强语音伪造检测,提升模型泛化能力。
Unmasking Deepfakes: Leveraging Augmentations and Features Variability for Deepfake Speech Detection
- 采用频谱与潜在特征双阶段掩码,增强对局部失真的鲁棒性。
- 在低资源下提升特征多样性,保持预训练表示完整性。
- 适用于语音安全、反伪造系统开发人员,尤其关注实际部署场景。
深度伪造语音检测面临生成音频技术不断进步的挑战。本文提出一种混合训练框架,通过创新的增强策略提升检测性能。首先,引入双阶段掩码方法,在频谱层(MaskedSpec)和潜在特征空间(MaskedFeature)分别操作,提供互补正则化,增强对局部失真的容忍度并促进泛化学习。其次,在自监督过程中采用压缩感知策略,提升低资源场景下的特征可变性,同时保持学习表征的完整性,使预训练特征更适配深度伪造检测任务。该框架将可学习的自监督特征提取器与ResNet分类头整合于统一训练流程中,实现声学表示与判别模式的联合优化。在ASVspoof5挑战赛(第1赛道)闭集条件下,系统取得4.08%的等错误率(EER),通过融合不同预训练特征提取器的模型进一步降至2.71%。在ASVspoof2019数据集上训练后,该系统在ASVspoof2019评测集上达到0.18% EER,ASVspoof2021 DF任务中为2.92% EER,均处于领先水平。
原文摘要 · Abstract (English)
Deepfake speech detection presents a growing challenge as generative audio technologies continue to advance. We propose a hybrid training framework that advances detection performance through novel augmentation strategies. First, we introduce a dual-stage masking approach that operates both at the spectrogram level (MaskedSpec) and within the latent feature space (MaskedFeature), providing complementary regularization that improves tolerance to localized distortions and enhances generalization learning. Second, we introduce compression-aware strategy during self-supervised to increase variability in low-resource scenarios while preserving the integrity of learned representations, thereby improving the suitability of pretrained features for deepfake detection. The framework integrates a learnable self-supervised feature extractor with a ResNet classification head in a unified training pipeline, enabling joint adaptation of acoustic representations and discriminative patterns. On the ASVspoof5 Challenge (Track~1), the system achieves state-of-the-art results with an Equal Error Rate (EER) of 4.08% under closed conditions, further reduced to 2.71% through fusion of models with diverse pretrained feature extractors. when trained on ASVspoof2019, our system obtaining leading performance on the ASVspoof2019 evaluation set (0.18% EER) and the ASVspoof2021 DF task (2.92% EER).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。