通过正则化与归一化提升无标签语音验证性能
Pushing the Frontiers of Self-Distillation Prototypes Network with Dimension Regularization and Score Normalization
- 引入维度正则化防止说话人嵌入坍缩
- 融合监督式评分归一化,使性能逼近有标签方法
- 在VoxCeleb1上达到1.29%的错误率,刷新自监督记录
无需说话人标签的鲁棒语音验证系统构建长期面临挑战。以往研究显示自监督与全监督方法间存在显著性能差距。本文改进非对比性自监督框架——自蒸馏原型网络(SDPN),通过在说话人嵌入上添加维度正则化项,明确缓解嵌入坍缩问题;同时引入全监督语音验证中的评分归一化技术,进一步缩小与监督方法的差距。结合维度正则化与评分归一化的SDPN在VoxCeleb1语音验证基准上取得新最优结果,试用集VoxCeleb1-{O,E,H}的等错误率分别为1.29%、1.60%和2.80%,相较当前最佳自监督方法分别提升28.3%、19.6%和22.6%,显著推进了语音验证技术边界。
原文摘要 · Abstract (English)
Developing robust speaker verification (SV) systems without speaker labels has been a longstanding challenge. Earlier research has highlighted a considerable performance gap between self-supervised and fully supervised approaches. In this paper, we enhance the non-contrastive self-supervised framework, Self-Distillation Prototypes Network (SDPN), by introducing dimension regularization that explicitly addresses the collapse problem through the application of regularization terms to speaker embeddings. Moreover, we integrate score normalization techniques from fully supervised SV to further bridge the gap toward supervised verification performance. SDPN with dimension regularization and score normalization sets a new state-of-the-art on the VoxCeleb1 speaker verification evaluation benchmark, achieving Equal Error Rate 1.29%, 1.60%, and 2.80% for trial VoxCeleb1-{O,E,H} respectively. These results demonstrate relative improvements of 28.3%, 19.6%, and 22.6% over the current best self-supervised methods, thereby advancing the frontiers of SV technology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。