提升语音识别在真实场景下的防欺骗能力,有效应对未知合成语音干扰。
BUT Systems for WildSpoof Challenge: SASV in the Wild
- 融合通用音频与语音专用模型,用轻量注意力机制聚合特征
- 引入分布不确定性增强策略,显著降低未知声学环境导致的性能下降
- 适合需要高鲁棒性语音验证系统的研发人员参考
本文介绍布科大学(BUT)参与野生语音欺骗挑战赛(WildSpoof Challenge)在抗欺骗自动说话人验证(SASV)赛道的方案。提出一种SASV框架,旨在弥合通用音频理解与专用语音分析之间的差距。系统集成多种自监督学习前端,涵盖通用音频模型(如Dasheng)和语音专用编码器(如WavLM),并通过轻量级多头因子化注意力后端对不同子任务进行特征聚合。此外,设计基于分布不确定性的特征域增强策略,显式建模并缓解由未见神经声码器及录音环境引起的领域偏移问题。通过融合这些鲁棒的CM分数与当前最先进的ASV系统,本方法在a-DCFs和EER指标上均取得更优表现。
原文摘要 · Abstract (English)
This paper presents the BUT submission to the WildSpoof Challenge, focusing on the Spoofing-robust Automatic Speaker Verification (SASV) track. We propose a SASV framework designed to bridge the gap between general audio understanding and specialized speech analysis. Our subsystem integrates diverse Self-Supervised Learning front-ends ranging from general audio models (e.g., Dasheng) to speech-specific encoders (e.g., WavLM). These representations are aggregated via a lightweight Multi-Head Factorized Attention back-end for corresponding subtasks. Furthermore, we introduce a feature domain augmentation strategy based on Distribution Uncertainty to explicitly model and mitigate the domain shift caused by unseen neural vocoders and recording environments. By fusing these robust CM scores with state-of-the-art ASV systems, our approach achieves superior minimization of the a-DCFs and EERs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。