用简单i向量模型启动自训练,就能达到顶尖说话人识别效果
Speaker-IPL: Unsupervised Learning of Speaker Characteristics with i-Vector based Pseudo-Labels
- 用i向量生成伪标签,迭代优化说话人表示
- 即使初始模型弱,仍达顶尖验证性能(EER 1.69%)
- 适合想低成本训练说话人模型的研究者
迭代自训练(IPL)通过用当前模型改进后的结果作为下一阶段的伪标签,显著提升了说话人表示的质量。现有方法通常从复杂的自监督模型(如DINO)提取初始表示,但这些模型训练复杂且可能泛化性差。本文发现,简单的、成熟的i向量生成模型已足够启动IPL流程,用于无监督学习说话人表示。我们系统研究了初始模型、编码器、数据增强、聚类数量和聚类算法对IPL的影响。结果表明,即使使用较弱的i向量模型,IPL仍能实现与最先进方法相当的性能,尤其在VoxCeleb2数据集上达到EER 1.69%,证明了其有效性与鲁棒性。
原文摘要 · Abstract (English)
Iterative self-training, or iterative pseudo-labeling (IPL) -- using an improved model from the current iteration to provide pseudo-labels for the next iteration -- has proven to be a powerful approach to enhance the quality of speaker representations. Recent applications of IPL in unsupervised speaker recognition start with representations extracted from very elaborate self-supervised methods (e.g., DINO). However, training such strong self-supervised models is not straightforward (they require hyper-parameter tuning and may not generalize to out-of-domain data) and, moreover, may not be needed at all. To this end, we show that the simple, well-studied, and established i-vector generative model is enough to bootstrap the IPL process for the unsupervised learning of speaker representations. We also systematically study the impact of other components on the IPL process, which includes the initial model, the encoder, augmentations, the number of clusters, and the clustering algorithm. Remarkably, we find that even with a simple and significantly weaker initial model like i-vector, IPL can still achieve speaker verification performance that rivals state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。