XMUspeech系统通过多尺度融合与自监督特征,有效提升语音伪造检测性能。
XMUspeech Systems for the ASVspoof 5 Challenge
- 采用自监督模型提取伪造音频中的多层级特征
- 融合Transformer层特征与手工特征,显著降低误报率
- 在闭集和开集条件下分别实现0.4783和0.2245的minDCF
本文介绍XMUspeech系统在ASVspoof 5挑战赛语音深度伪造检测任务中的表现。相比以往挑战,ASVspoof 5数据库中音频时长显著增加。我们发现仅调整输入音频长度即可显著提升系统性能。为捕捉多层次伪造痕迹,我们测试了AASIST、HM-Conformer、Hubert和Wav2vec2在不同输入特征与损失函数下的表现。为获取伪造相关特征,我们在包含伪造语句的数据集上训练自监督模型作为特征提取器,并提出自适应多尺度特征融合(AMFF)方法,将多层Transformer特征与手工特征融合以增强检测能力。此外,我们对单类损失函数进行了广泛实验,提供了更契合反伪造任务的优化配置。融合系统在闭集条件下达到minDCF 0.4783、EER 20.45%;在开集条件下达到minDCF 0.2245、EER 9.36%。
原文摘要 · Abstract (English)
In this paper, we present our submitted XMUspeech systems to the speech deepfake detection track of the ASVspoof 5 Challenge. Compared to previous challenges, the audio duration in ASVspoof 5 database has significantly increased. And we observed that merely adjusting the input audio length can substantially improve system performance. To capture artifacts at multiple levels, we explored the performance of AASIST, HM-Conformer, Hubert, and Wav2vec2 with various input features and loss functions. Specifically, in order to obtain artifact-related information, we trained self-supervised models on the dataset containing spoofing utterances as the feature extractors. And we applied an adaptive multi-scale feature fusion (AMFF) method to integrate features from multiple Transformer layers with the hand-crafted feature to enhance the detection capability. In addition, we conducted extensive experiments on one-class loss functions and provided optimized configurations to better align with the anti-spoofing task. Our fusion system achieved a minDCF of 0.4783 and an EER of 20.45% in the closed condition, and a minDCF of 0.2245 and an EER of 9.36% in the open condition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。