提升开放集说话人识别鲁棒性,显著降低错误率。
SpeakerRPL v2: Robust Open-set Speaker Identification through Enhanced Few-shot Foundation Tuning and Model Fusion

- 结合互点学习与对数归一化,增强目标说话人表征约束。
- 在新测试集上将错误率从1.28%降至0.09%,相对降低93%。
- 适合需要高可靠少样本说话人识别的场景,如安全验证。
本文提出一种基于预训练说话人基础模型的改进开放集说话人识别方法。在前期互点学习框架(V1)基础上,引入增强的开放集学习目标:融合互点学习与对数归一化(LogitNorm),并加入自适应锚点学习,更好约束目标说话人表征,提升鲁棒性。其次,提出模型融合策略,稳定并强化少样本微调过程,有效降低结果随机性,提升泛化能力。此外,设计模型选择方法以保障融合性能最优。在VoxCeleb、ESD和3D-Speaker数据集上的实验表明,该方法在多种条件下均具有效性和鲁棒性。在新提出的Vox1-O类测试集上,错误率从1.28%降至0.09%,相对降低约93%。
原文摘要 · Abstract (English)
This paper proposes an improved approach for open-set speaker identification based on pretrained speaker foundation models. Building upon the previous Speaker Reciprocal Points Learning framework (V1), we first introduce an enhanced open-set learning objective by integrating reciprocal points learning with logit normalization (LogitNorm) and incorporating adaptive anchor learning to better constrain target speaker representations and improve robustness. Second, we propose a model fusion strategy to stabilize and enhance the few-shot tuning process, effectively reducing result randomness and improving generalization. Furthermore, we introduce a model selection method to ensure optimal performance in model fusion. Experimental evaluations on the VoxCeleb, ESD and 3D-Speaker datasets demonstrate the effectiveness and robustness of the proposed method under diverse conditions. On a newly proposed Vox1-O-like test set, our method reduces the EER from 1.28% to 0.09%, achieving a relative reduction of approximately 93%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。