对比多种模型与增强策略,提升音频深度伪造检测的泛化能力。
Generalizable Detection of Audio Deepfakes
- 测试Wav2Vec2、WavLM、Whisper等预训练模型在多数据集上的表现
- 采用不同数据增强与损失函数,使检测性能超越ASVspoof 5挑战冠军系统
- 为音频伪造检测提供可复用的优化方法,适合安全与可信通信研究者
本文开展全面研究,旨在提升音频深度伪造检测模型的泛化能力。我们评估了多种预训练骨干网络(包括Wav2Vec2、WavLM和Whisper)在多样化数据集上的表现,涵盖ASVspoof挑战及其他来源数据。实验重点分析不同数据增强策略和损失函数对模型性能的影响。结果表明,所提方法显著增强了检测模型的泛化能力,性能超过ASVspoof 5挑战中排名最高的单系统。本研究为构建更鲁棒的音频深度伪造检测模型提供了关键洞见,推动该领域未来发展。
原文摘要 · Abstract (English)
In this paper, we present our comprehensive study aimed at enhancing the generalization capabilities of audio deepfake detection models. We investigate the performance of various pre-trained backbones, including Wav2Vec2, WavLM, and Whisper, across a diverse set of datasets, including those from the ASVspoof challenges and additional sources. Our experiments focus on the effects of different data augmentation strategies and loss functions on model performance. The results of our research demonstrate substantial enhancements in the generalization capabilities of audio deepfake detection models, surpassing the performance of the top-ranked single system in the ASVspoof 5 Challenge. This study contributes valuable insights into the optimization of audio models for more robust deepfake detection and facilitates future research in this critical area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。