arXiv:2505.14561eess.AScs.AI2025-05中稿 · Interspeech 2025被引 3

提出新采样方法,让语音验证更抗录音条件干扰。

SSPS: Self-Supervised Positive Sampling for Robust Self-Supervised Speaker Verification

  • 在隐空间中找同人异条件正样本,避免通道信息干扰。
  • 在VoxCeleb1-O上将误识率降至2.57%,比现有方法更好。
  • 适合做鲁棒语音验证的研究者和工程师参考。

自监督学习(SSL)在语音验证(SV)中取得显著进展。标准框架采用同一语句内的正样本采样与数据增强生成同说话人锚点-正样本对,但此策略主要编码了由锚点和正样本共享的录音条件中的通道信息,成为主要瓶颈。本文提出新的正样本采样技术——自监督正样本采样(SSPS):对于给定锚点,利用聚类分配与正样本嵌入的记忆队列,在隐空间中寻找同说话人身份但不同录音条件的正样本。SSPS提升了SimCLR和DINO的性能,在VoxCeleb1-O上分别达到2.57%和2.53%的等错误率(EER),优于当前最先进方法。尤其地,SimCLR-SSPS通过降低说话人内方差,实现58%的误识率下降,性能可媲美DINO-SSPS。

原文摘要 · Abstract (English)

Self-Supervised Learning (SSL) has led to considerable progress in Speaker Verification (SV). The standard framework uses same-utterance positive sampling and data-augmentation to generate anchor-positive pairs of the same speaker. This is a major limitation, as this strategy primarily encodes channel information from the recording condition, shared by the anchor and positive. We propose a new positive sampling technique to address this bottleneck: Self-Supervised Positive Sampling (SSPS). For a given anchor, SSPS aims to find an appropriate positive, i.e., of the same speaker identity but a different recording condition, in the latent space using clustering assignments and a memory queue of positive embeddings. SSPS improves SV performance for both SimCLR and DINO, reaching 2.57% and 2.53% EER, outperforming SOTA SSL methods on VoxCeleb1-O. In particular, SimCLR-SSPS achieves a 58% EER reduction by lowering intra-speaker variance, providing comparable performance to DINO-SSPS.

语音验证自监督学习说话人识别正样本采样

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。