通过两阶段学习提升语音伪造检测鲁棒性,尤其在跨域场景下表现优异。
Wav2DF-TSL: Two-stage Learning with Efficient Pre-training and Hierarchical Experts Fusion for Robust Audio Deepfake Detection
- 先用3000小时伪语音训练模型,高效学习伪造特征
- 跨域测试中等错误率降低27.5%,优于现有方法
- 适合需要高鲁棒性的语音安全检测应用场景
近年来,自监督学习(SSL)模型在语音伪造检测(ADD)任务中取得显著进展。然而,现有SSL模型主要依赖大规模真实语音进行预训练,缺乏对伪造样本的学习,导致在ADD任务微调过程中易受领域偏移影响。为此,我们提出一种基于预训练和分层专家融合的两阶段学习策略(Wav2DF-TSL),以提升音频伪造检测的鲁棒性。在预训练阶段,利用适配器高效学习3000小时未标注伪造语音中的伪影特征,增强前端特征适应性并缓解灾难性遗忘。在微调阶段,提出分层自适应专家混合(HA-MoE)方法,通过门控路由实现多专家协作,动态融合多层级伪造线索。实验结果表明,该方法在所有四个基准数据集上均显著优于基线系统,尤其在跨域In-the-wild数据集上,等错误率(EER)相对降低27.5%,超越现有最先进方法。
原文摘要 · Abstract (English)
In recent years, self-supervised learning (SSL) models have made significant progress in audio deepfake detection (ADD) tasks. However, existing SSL models mainly rely on large-scale real speech for pre-training and lack the learning of spoofed samples, which leads to susceptibility to domain bias during the fine-tuning process of the ADD task. To this end, we propose a two-stage learning strategy (Wav2DF-TSL) based on pre-training and hierarchical expert fusion for robust audio deepfake detection. In the pre-training stage, we use adapters to efficiently learn artifacts from 3000 hours of unlabelled spoofed speech, improving the adaptability of front-end features while mitigating catastrophic forgetting. In the fine-tuning stage, we propose the hierarchical adaptive mixture of experts (HA-MoE) method to dynamically fuse multi-level spoofing cues through multi-expert collaboration with gated routing. Experimental results show that the proposed method significantly outperforms the baseline system on all four benchmark datasets, especially on the cross-domain In-the-wild dataset, achieving a 27.5% relative improvement in equal error rate (EER), outperforming the existing state-of-the-art systems. Index Terms: audio deepfake detection, self-supervised learning, parameter-efficient fine-tuning, mixture of experts
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。