arXiv:2602.10829eess.AScs.LG2026-02中稿 · publication in Spe…综述被引 2

对比三种自监督学习方法在说话人识别中的表现与稳定性。

Self-Supervised Learning for Speaker Recognition: A study and review

  • 借鉴计算机视觉框架,将自监督学习用于说话人识别。
  • DINO性能最佳但对超参数敏感,SimCLR/MoCo更稳健。
  • 适合研究自监督学习在语音任务中应用的学者。

监督学习模型虽已革新音频与语音处理,但其性能高度依赖人工标注数据,导致扩展成本高且在未见条件下泛化能力差。自监督学习(SSL)通过利用大量无标签数据学习有效表征,成为解决该问题的有前景范式。尽管SSL在自动语音识别(ASR)中研究充分,但在其他下游任务如说话人识别(SR)中仍处于早期阶段。本文综述了主要的SSL实例不变性框架(如SimCLR、MoCo、DINO),并探讨其在SR中的适配。系统分析了这些方法:(1)主超参数的影响;(2)组件作用(如数据增强、投影器、正样本采样);(3)在域内与域外数据上的一致实验设置下评估性能,并提供文献中方法的全面比较。结果表明,DINO在下游任务中表现最优,能有效建模说话人内部变异性,但对超参数和训练条件高度敏感;而SimCLR和MoCo提供更稳健的替代方案,有效捕捉说话人间差异,且不易陷入坍缩。本文旨在揭示当前趋势与挑战。

原文摘要 · Abstract (English)

Deep learning models trained in a supervised setting have revolutionized audio and speech processing. However, their performance inherently depends on the quantity of human-annotated data, making them costly to scale and prone to poor generalization under unseen conditions. To address these challenges, Self-Supervised Learning (SSL) has emerged as a promising paradigm, leveraging vast amounts of unlabeled data to learn relevant representations. The application of SSL for Automatic Speech Recognition (ASR) has been extensively studied, but research on other downstream tasks, notably Speaker Recognition (SR), remains in its early stages. This work describes major SSL instance-invariance frameworks (e.g., SimCLR, MoCo, and DINO), initially developed for computer vision, along with their adaptation to SR. Various SSL methods for SR, proposed in the literature and built upon these frameworks, are also presented. An extensive review of these approaches is then conducted: (1) the effect of the main hyperparameters of SSL frameworks is investigated; (2) the role of SSL components is studied (e.g., data-augmentation, projector, positive sampling); and (3) SSL frameworks are evaluated on SR with in-domain and out-of-domain data, using a consistent experimental setup, and a comprehensive comparison of SSL methods from the literature is provided. Specifically, DINO achieves the best downstream performance and effectively models intra-speaker variability, although it is highly sensitive to hyperparameters and training conditions, while SimCLR and MoCo provide robust alternatives that effectively capture inter-speaker variability and are less prone to collapse. This work aims to highlight recent trends and advancements, identifying current challenges in the field.

自监督学习说话人识别深度学习语音处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。