arXiv:2410.07379eess.AScs.AI2024-10中稿 · to ASVspoof5 works…被引 8

通过自监督学习增强语音真伪判别能力,提升模型泛化性。

Learn from Real: Reality Defender's Submission to ASVspoof5 Challenge

  • 采用自监督对比学习提取语音风格与语义的依赖关系特征
  • 在ASVspoof5上实现0.1499最小DCF和5.5%误报率
  • 轻量级设计适合实际部署,对真实环境噪声鲁棒

语音深度伪造检测对于防范人工智能合成语音的恶意使用至关重要。在社区诸多努力中,ASVspoof挑战赛已成为评估检测模型泛化性和鲁棒性的基准之一。本文介绍Reality Defender对ASVspoof5挑战赛的提交方案,提出一种新颖的预训练策略,在保持低训练计算成本的同时显著提升模型泛化能力。我们的系统SLIM利用自监督对比学习,从多种真实语音中学习风格-语言依赖嵌入表示,通过关注风格与语义间的关联来区分伪造与真实语音。我们在ASVspoof5、ASV2019和In-the-wild数据集上进行了评估,提交结果在ASVspoof5 Track 1上达到最小代价函数(minDCF)0.1499、等错误率(EER)5.5%;在ASV2019和In-the-wild上的EER分别为7.4%和10.8%。

原文摘要 · Abstract (English)

Audio deepfake detection is crucial to combat the malicious use of AI-synthesized speech. Among many efforts undertaken by the community, the ASVspoof challenge has become one of the benchmarks to evaluate the generalizability and robustness of detection models. In this paper, we present Reality Defender's submission to the ASVspoof5 challenge, highlighting a novel pretraining strategy which significantly improves generalizability while maintaining low computational cost during training. Our system SLIM learns the style-linguistics dependency embeddings from various types of bonafide speech using self-supervised contrastive learning. The learned embeddings help to discriminate spoof from bonafide speech by focusing on the relationship between the style and linguistics aspects. We evaluated our system on ASVspoof5, ASV2019, and In-the-wild. Our submission achieved minDCF of 0.1499 and EER of 5.5% on ASVspoof5 Track 1, and EER of 7.4% and 10.8% on ASV2019 and In-the-wild respectively.

语音检测深度伪造自监督学习ASVspoof

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。