通过最小化互信息,分离语音中的年龄与身份特征,提升跨年龄说话人识别效果。
Disentangling Age and Identity with a Mutual Information Minimization Approach for Cross-Age Speaker Verification
- 用互信息最小化方法分离身份与年龄相关嵌入
- 在Vox-CA多个跨年龄测试集上性能优于现有方法
- 引入年龄差感知损失,更关注大年龄跨度的语音变化
跨年龄说话人验证(CASV)受到越来越多关注。然而,由于年龄带来的语音个体差异巨大,现有系统在CASV任务中表现不佳。本文提出一种基于互信息(MI)最小化的解耦表示学习框架。该方法通过主干模型从说话人信息中解耦出身份与年龄相关的嵌入,并训练一个互信息估计器,通过最小化年龄与身份相关嵌入之间的相关性,得到年龄不变的说话人嵌入。此外,我们利用正负样本间的年龄差,设计了一种年龄感知的互信息最小化损失函数,使主干模型更关注大年龄跨度下的声学变化。实验结果表明,所提方法在Vox-CA多个跨年龄测试集上均优于其他方法。
原文摘要 · Abstract (English)
There has been an increasing research interest in cross-age speaker verification~(CASV). However, existing speaker verification systems perform poorly in CASV due to the great individual differences in voice caused by aging. In this paper, we propose a disentangled representation learning framework for CASV based on mutual information~(MI) minimization. In our method, a backbone model is trained to disentangle the identity- and age-related embeddings from speaker information, and an MI estimator is trained to minimize the correlation between age- and identity-related embeddings via MI minimization, resulting in age-invariant speaker embeddings. Furthermore, by using the age gaps between positive and negative samples, we propose an aging-aware MI minimization loss function that allows the backbone model to focus more on the vocal changes with large age gaps. Experimental results show that the proposed method outperforms other methods on multiple Cross-Age test sets of Vox-CA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。