arXiv:2410.17033eess.AScs.SD2024-10中稿 · ISCSLP 2024被引 1

通过双层对比学习提升说话人验证在跨域场景下的泛化能力

Prototype and Instance Contrastive Learning for Unsupervised Domain Adaptation in Speaker Verification

  • 用聚类生成伪标签构建动态原型,实现类别级对比对齐
  • 对同一语音的不同增强版本进行实例对比,提升鲁棒性
  • 在多个数据集上表现优于现有方法,适合跨域语音识别任务

基于单一领域训练的说话人验证系统在应用到新领域时性能通常下降。为解决此问题,研究者常采用基于特征分布匹配的无监督域适应方法,但这些方法在多种不匹配情况下泛化能力有限且提升效果不足。本文提出原型与实例对比学习(PICL),一种通过双层对比学习实现说话人验证无监督域适应的新方法。原型对比学习利用聚类生成伪标签,构建动态更新的原型表示,将样本与其对应类别或簇原型对齐;实例对比学习则最小化同一实例不同视图或增强版本间的距离,确保表示对噪声等变化具有鲁棒性。该双层机制同时提供高层和低层监督,显著提升模型的泛化性和鲁棒性。与以往仅在单一不匹配情境下评估的研究不同,本文在多个数据集上进行了探索,取得了当前最优性能,验证了方法的广泛适用性。

原文摘要 · Abstract (English)

Speaker verification system trained on one domain usually suffers performance degradation when applied to another domain. To address this challenge, researchers commonly use feature distribution matching-based methods in unsupervised domain adaptation scenarios where some unlabeled target domain data is available. However, these methods often have limited performance improvement and lack generalization in various mismatch situations. In this paper, we propose Prototype and Instance Contrastive Learning (PICL), a novel method for unsupervised domain adaptation in speaker verification through dual-level contrastive learning. For prototype contrastive learning, we generate pseudo labels via clustering to create dynamically updated prototype representations, aligning instances with their corresponding class or cluster prototypes. For instance contrastive learning, we minimize the distance between different views or augmentations of the same instance, ensuring robust and invariant representations resilient to variations like noise. This dual-level approach provides both high-level and low-level supervision, leading to improved generalization and robustness of the speaker verification model. Unlike previous studies that only evaluated mismatches in one situation, we have conducted relevant explorations on various datasets and achieved state-of-the-art performance currently, which also proves the generalization of our method.

说话人验证域适应对比学习无监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。