让语音识别更抗情绪干扰,提升说话人验证准确率。
Learning Emotion-Invariant Speaker Representations for Speaker Verification
- 用数据增强合成同一人不同情绪的语音对
- 通过余弦相似度损失降低情绪对语音特征的影响
- 结合能量感知掩码强化不变特征,适合高鲁棒性场景
近年来,基于深度学习的说话人验证技术快速发展,但其提取的说话人表征仍易受情绪变化影响。为此,本文提出多项改进以提升说话人编码器对情绪的鲁棒性:首先,采用基于CopyPaste的数据增强方法生成包含同一说话人不同情绪表达的平行数据;其次,引入余弦相似度损失,约束平行样本对的表征差异,减少表征与情绪信息的相关性;最后,基于语音信号能量设计情绪感知掩码(EM),在输入端进一步强化说话人表征的不变性。通过全面消融实验验证各组件有效性。结果表明,所提方法相较基线系统在EER上相对降低19.29%。
原文摘要 · Abstract (English)
In recent years, the rapid progress in speaker verification (SV) technology has been driven by the extraction of speaker representations based on deep learning. However, such representations are still vulnerable to emotion variability. To address this issue, we propose multiple improvements to train speaker encoders to increase emotion robustness. Firstly, we utilize CopyPaste-based data augmentation to gather additional parallel data, which includes different emotional expressions from the same speaker. Secondly, we apply cosine similarity loss to restrict parallel sample pairs and minimize intra-class variation of speaker representations to reduce their correlation with emotional information. Finally, we use emotion-aware masking (EM) based on the speech signal energy on the input parallel samples to further strengthen the speaker representation and make it emotion-invariant. We conduct a comprehensive ablation study to demonstrate the effectiveness of these various components. Experimental results show that our proposed method achieves a relative 19.29\% drop in EER compared to the baseline system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。