arXiv:2501.17772eess.AScs.LG2025-01中稿 · publication in IEE…被引 4

通过自举正样本提升语音验证的自监督学习效果

Self-Supervised Frameworks for Speaker Verification via Bootstrapped Positive Sampling

  • 用表示空间相近的伪正样本替代原始正样本,增强多样性
  • 在VoxCeleb1-O上,SimCLR与DINO分别达2.57%和2.53%的EER
  • 无需数据增强即可降低通道信息,适合实际部署场景

自监督学习(SSL)在语音验证(SV)中展现出巨大潜力,但与有监督系统仍存在性能差距。现有SSL框架依赖同一语音片段生成锚点-正样本对,导致正样本与锚点具有相似的录音通道特征,即使经过大量数据增强也无法完全消除这一问题。本文提出自监督正样本采样(SSPS),一种基于自举的策略,从表示空间中选取与锚点相近的伪正样本,这些样本属于同一说话人身份但对应不同录音条件。该方法在多种主流SSL框架(如SimCLR、SwAV、VICReg、DINO)上均显著提升SV性能。在VoxCeleb1-O测试集上,使用SSPS的SimCLR和DINO分别达到2.57%和2.53%的等错误率(EER)。SimCLR实现58%相对误差减少,且在更简单的训练框架下表现接近DINO。此外,SSPS有效降低类内方差,减少说话人表征中的通道信息,并在无数据增强条件下表现出更强鲁棒性。

原文摘要 · Abstract (English)

Recent developments in Self-Supervised Learning (SSL) have demonstrated significant potential for Speaker Verification (SV), but closing the performance gap with supervised systems remains an ongoing challenge. SSL frameworks rely on anchor-positive pairs, constructed from segments of the same audio utterance. Hence, positives have channel characteristics similar to those of their corresponding anchors, even with extensive data-augmentation. Therefore, this positive sampling strategy is a fundamental limitation as it encodes too much information regarding the recording source in the learned representations. This article introduces Self-Supervised Positive Sampling (SSPS), a bootstrapped technique for sampling appropriate and diverse positives in SSL frameworks for SV. SSPS samples positives close to their anchor in the representation space, assuming that these pseudo-positives belong to the same speaker identity but correspond to different recording conditions. This method consistently demonstrates improvements in SV performance on VoxCeleb benchmarks when applied to major SSL frameworks, including SimCLR, SwAV, VICReg, and DINO. Using SSPS, SimCLR and DINO achieve 2.57% and 2.53% EER on VoxCeleb1-O, respectively. SimCLR yields a 58% relative reduction in EER, getting comparable performance to DINO with a simpler training framework. Furthermore, SSPS lowers intra-class variance and reduces channel information in speaker representations while exhibiting greater robustness without data-augmentation.

语音验证自监督学习表征学习说话人识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。