用语音大模型集成检测假唱,准确率达1.79%等误率。
Speech Foundation Model Ensembles for the Controlled Singing Voice Deepfake Detection (CtrSVDD) Challenge 2024
- 融合多个语音大模型的预测结果,提升检测鲁棒性。
- 在挑战赛测试集上达到1.79%的等误率,性能领先。
- 提出新聚合方法SEA,有效整合多模型特征。
本工作介绍了我们在2024年可控假唱检测挑战赛(CtrSVDD)中取得领先成绩的方法,系统在评测集上的合并等误率(EER)为1.79%。生成式AI模型的快速发展给假唱语音检测带来严峻挑战,引发研究关注。2024年歌唱语音深度伪造检测(SVDD)挑战赛旨在应对这一复杂任务。本文探索了基于语音基础模型的集成方法,构建稳健的歌唱语音反欺骗系统,并提出一种新型挤压-激励聚合(Squeeze-and-Excitation Aggregation, SEA)方法,高效整合语音基础模型的表征特征,优于其他单一系统。实验结果验证了该方法在检测假唱语音方面的有效性。代码已开源:https://github.com/Anmol2059/SVDD2024。
原文摘要 · Abstract (English)
This work details our approach to achieving a leading system with a 1.79% pooled equal error rate (EER) on the evaluation set of the Controlled Singing Voice Deepfake Detection (CtrSVDD). The rapid advancement of generative AI models presents significant challenges for detecting AI-generated deepfake singing voices, attracting increased research attention. The Singing Voice Deepfake Detection (SVDD) Challenge 2024 aims to address this complex task. In this work, we explore the ensemble methods, utilizing speech foundation models to develop robust singing voice anti-spoofing systems. We also introduce a novel Squeeze-and-Excitation Aggregation (SEA) method, which efficiently and effectively integrates representation features from the speech foundation models, surpassing the performance of our other individual systems. Evaluation results confirm the efficacy of our approach in detecting deepfake singing voices. The codes can be accessed at https://github.com/Anmol2059/SVDD2024.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。