arXiv:2410.05037cs.SDeess.AS2024-10被引 3

用对比学习提升多尺度语音特征,显著改善说话人验证效果

Improving Speaker Representations Using Contrastive Losses on Multi-scale Features

  • 对网络各层中间特征图应用对比损失,强化特征判别性
  • 在VoxCeleb1-O上使等错误率降低9.05%(绝对值)
  • 适合做说话人验证模型优化的研究者和工程师

说话人验证系统随着多尺度特征聚合(MFA)架构(如MFA-Conformer和ECAPA-TDNN)的出现取得了显著进展。这些模型通过在池化与投影层前拼接不同网络深度的中间特征图,利用深层与浅层特征,证明了即使较浅层的特征也包含有价值的说话人信息。在此基础上,我们提出多尺度特征对比(MFCon)损失,直接增强这些中间表示的质量。该损失将对比学习应用于网络中所有特征图,促使模型在中间阶段即学习更具判别性的表示。通过强化特征学习,我们发现生成的说话人嵌入具有更强的判别能力。在VoxCeleb-1O测试集上,相比标准的MFA-Conformer,本方法实现了9.05%的等错误率(EER)改进。

原文摘要 · Abstract (English)

Speaker verification systems have seen significant advancements with the introduction of Multi-scale Feature Aggregation (MFA) architectures, such as MFA-Conformer and ECAPA-TDNN. These models leverage information from various network depths by concatenating intermediate feature maps before the pooling and projection layers, demonstrating that even shallower feature maps encode valuable speaker-specific information. Building upon this foundation, we propose a Multi-scale Feature Contrastive (MFCon) loss that directly enhances the quality of these intermediate representations. Our MFCon loss applies contrastive learning to all feature maps within the network, encouraging the model to learn more discriminative representations at the intermediate stage itself. By enforcing better feature map learning, we show that the resulting speaker embeddings exhibit increased discriminative power. Our method achieves a 9.05% improvement in equal error rate (EER) compared to the standard MFA-Conformer on the VoxCeleb-1O test set.

说话人验证对比学习多尺度特征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。