arXiv:2507.17540eess.AScs.LG2025-07中稿 · INTERSPEECH 2025

用聚类筛选难负例,提升语音验证模型性能

Clustering-based hard negative sampling for supervised contrastive speaker verification

  • 通过聚类相似说话人嵌入,动态调整批次中难负例比例
  • 在VoxCeleb数据集上相对误差率降低18%,minDCF更优
  • 适合追求轻量级高精度语音验证的工程师和研究者

在语音验证任务中,对比学习正逐步取代传统分类方法。有效的难负例(即不同说话人但特征相似的样本)能显著提升模型表现。本文提出一种基于聚类的难负例采样方法CHNS,专用于监督对比学习中的说话人表征。该方法对相似说话人的嵌入进行聚类,并动态调整批次组成,以优化对比损失计算时难负例与易负例的比例。实验表明,CHNS在使用两种轻量级模型架构的情况下,优于基线监督对比方法(含无损负例采样)及当前最优分类方法,在VoxCeleb数据集上的相对EER降低18%,minDCF也显著改善。

原文摘要 · Abstract (English)

In speaker verification, contrastive learning is gaining popularity as an alternative to the traditionally used classification-based approaches. Contrastive methods can benefit from an effective use of hard negative pairs, which are different-class samples particularly challenging for a verification model due to their similarity. In this paper, we propose CHNS - a clustering-based hard negative sampling method, dedicated for supervised contrastive speaker representation learning. Our approach clusters embeddings of similar speakers, and adjusts batch composition to obtain an optimal ratio of hard and easy negatives during contrastive loss calculation. Experimental evaluation shows that CHNS outperforms a baseline supervised contrastive approach with and without loss-based hard negative sampling, as well as a state-of-the-art classification-based approach to speaker verification by as much as 18 % relative EER and minDCF on the VoxCeleb dataset using two lightweight model architectures.

语音验证对比学习难负例采样聚类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。