CUHK提出新语音音色属性检测系统,兼顾泛化与精度
CUHK-EE Systems for the vTAD Challenge at NCMMSC 2025
- 用WavLM-Large+注意力统计池化提取鲁棒声纹表征
- 不同网络结构在未见/已见说话人上分别达77.96%与94.42%准确率
- 揭示模型复杂度与泛化能力的权衡,适合声纹建模研究者
本文介绍香港中文大学电子工程系数字信号处理与语音技术实验室(DSP&STL)为第20届全国人机语音通信会议(NCMMSC 2025)语音音色属性检测(vTAD)挑战赛所开发的系统。系统采用WavLM-Large嵌入结合注意力统计池化(ASTP)提取稳健的说话人表征,随后使用两种Diff-Net变体——前馈神经网络(FFN)和挤压-激励增强残差前馈网络(SE-ResFFN),比较话语对间的音色属性强度。实验表明,WavLM-Large+FFN系统在未见说话人上泛化更好,准确率达77.96%,等错误率(EER)为21.79%;而WavLM-Large+SE-ResFFN在‘已见’设置下表现更优,准确率为94.42%,EER为5.49%。结果凸显了模型复杂度与泛化能力之间的权衡,强调架构选择在细粒度说话人建模中的重要性。分析还揭示了说话人身份、标注主观性和数据不平衡对系统性能的影响,指明未来提升音色属性检测鲁棒性与公平性的方向。
原文摘要 · Abstract (English)
This paper presents the Voice Timbre Attribute Detection (vTAD) systems developed by the Digital Signal Processing & Speech Technology Laboratory (DSP&STL) of the Department of Electronic Engineering (EE) at The Chinese University of Hong Kong (CUHK) for the 20th National Conference on Human-Computer Speech Communication (NCMMSC 2025) vTAD Challenge. The proposed systems leverage WavLM-Large embeddings with attentive statistical pooling (ASTP) to extract robust speaker representations, followed by two variants of Diff-Net, i.e., Feed-Forward Neural Network (FFN) and Squeeze-and-Excitation-enhanced Residual FFN (SE-ResFFN), to compare timbre attribute intensities between utterance pairs. Experimental results demonstrate that the WavLM-Large+FFN system generalises better to unseen speakers, achieving 77.96% accuracy and 21.79% equal error rate (EER), while the WavLM-Large+SE-ResFFN model excels in the 'Seen' setting with 94.42% accuracy and 5.49% EER. These findings highlight a trade-off between model complexity and generalisation, and underscore the importance of architectural choices in fine-grained speaker modelling. Our analysis also reveals the impact of speaker identity, annotation subjectivity, and data imbalance on system performance, pointing to future directions for improving robustness and fairness in timbre attribute detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。