用最优传输和伪标签缓解语音识别中的通道差异问题
Channel Adaptation for Speaker Verification Using Optimal Transport with Pseudo Label
- 基于最优传输对齐分布,引入伪标签增强判别能力
- 在VoxCeleb数据集上使错误率降低超过10%
- 适合解决训练与实际测试数据不匹配的语音验证场景
域差距常导致说话人验证(SV)系统在真实语音测试时性能下降,其中通道变化是主要因素之一,但远未得到充分关注。尽管已有多种域自适应算法可用于缓解该问题,多数方法难以处理域对齐中复杂的分布结构与判别学习的结合。本文提出一种新的无监督域自适应方法——联合部分最优传输与伪标签(JPOT-PL),利用最优传输的几何感知距离度量进行分布对齐,并设计基于伪标签的判别学习机制,其中伪标签由最优耦合生成,可视为一种新型软说话人标签。在以VoxCeleb为基准语料的说话人验证通道自适应任务上进行实验,结果表明,相比若干先进通道自适应算法,本方法将等错误率(EER)降低了超过10%。
原文摘要 · Abstract (English)
Domain gap often degrades the performance of speaker verification (SV) systems when the statistical distributions of training data and real-world test speech are mismatched. Channel variation, a primary factor causing this gap, is less addressed than other issues (e.g., noise). Although various domain adaptation algorithms could be applied to handle this domain gap problem, most algorithms could not take the complex distribution structure in domain alignment with discriminative learning. In this paper, we propose a novel unsupervised domain adaptation method, i.e., Joint Partial Optimal Transport with Pseudo Label (JPOT-PL), to alleviate the channel mismatch problem. Leveraging the geometric-aware distance metric of optimal transport in distribution alignment, we further design a pseudo label-based discriminative learning where the pseudo label can be regarded as a new type of soft speaker label derived from the optimal coupling. With the JPOT-PL, we carry out experiments on the SV channel adaptation task with VoxCeleb as the basis corpus. Experiments show our method reduces EER by over 10% compared with several state-of-the-art channel adaptation algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。