构建鲁棒语音欺骗检测与识别框架,提升真实场景下说话人验证性能。
DFKI-Speech System for WildSpoof Challenge: A robust framework for SASV In-the-Wild
- 双模块协同:欺骗检测器结合自监督特征提取与图神经网络,实现高精度伪造语音识别。
- 多尺度融合:采用分层混合专家模型融合高低层特征,有效提升对伪造语音的判别能力。
- 轻量级优化:通过对比圆损失和固定假身份共现归一化,增强模型区分难样本的能力。
本文介绍了为野环境语音欺骗挑战赛(WildSpoof Challenge)中欺骗感知自动说话人验证(SASV)赛道设计的DFKI-Speech系统。提出一种鲁棒的SASV框架,其中欺骗检测器与说话人验证(SV)网络协同工作。欺骗检测器采用自监督语音嵌入提取器作为前端,结合先进的图神经网络后端;同时使用基于前3层的混合专家(MoE)模型融合高层与低层特征,以实现有效的伪造语音检测。对于说话人验证,采用低复杂度卷积神经网络,在多尺度上融合2D与1D特征,并使用SphereFace损失进行训练。此外,引入对比圆损失,自适应调整每批次中的正负样本权重,使网络更好地区分难易样本对。最后,采用固定假身份共现归一化(AS Norm)与模型集成,进一步提升说话人验证系统的判别能力。
原文摘要 · Abstract (English)
This paper presents the DFKI-Speech system developed for the WildSpoof Challenge under the Spoofing aware Automatic Speaker Verification (SASV) track. We propose a robust SASV framework in which a spoofing detector and a speaker verification (SV) network operate in tandem. The spoofing detector employs a self-supervised speech embedding extractor as the frontend, combined with a state-of-the-art graph neural network backend. In addition, a top-3 layer based mixture-of-experts (MoE) is used to fuse high-level and low-level features for effective spoofed utterance detection. For speaker verification, we adapt a low-complexity convolutional neural network that fuses 2D and 1D features at multiple scales, trained with the SphereFace loss. Additionally, contrastive circle loss is applied to adaptively weight positive and negative pairs within each training batch, enabling the network to better distinguish between hard and easy sample pairs. Finally, fixed imposter cohort based AS Norm score normalization and model ensembling are used to further enhance the discriminative capability of the speaker verification system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。