用音调循环和边界损失提升无文本语音识别准确率
Text-Independent Speaker Identification Using Audio Looping With Margin Based Loss Functions
- 基于VGG16改造模型,输入可变长度梅尔频谱图
- CosFace与ArcFace损失使识别准确率显著优于Softmax
- 适合语音识别、安全系统等需要鲁棒声纹识别的场景
说话人识别在安防系统、虚拟助手和个性化体验中日益重要。本文研究基于VGG16架构的卷积神经网络在无文本说话人识别中的表现,该模型经修改以处理来自VoxCeleb1数据集的可变尺寸梅尔频谱图。通过对比软最大值损失(Softmax loss)基线,引入了CosFace损失和ArcFace损失,并分析其对模型精度与鲁棒性的影响。实验结果表明,采用边界感知损失函数的模型在识别准确率上显著优于传统方法。此外,研究还探讨了梅尔频谱图尺寸及时间长度变化对性能的影响,为未来研究提供参考。
原文摘要 · Abstract (English)
Speaker identification has become a crucial component in various applications, including security systems, virtual assistants, and personalized user experiences. In this paper, we investigate the effectiveness of CosFace Loss and ArcFace Loss for text-independent speaker identification using a Convolutional Neural Network architecture based on the VGG16 model, modified to accommodate mel spectrogram inputs of variable sizes generated from the Voxceleb1 dataset. Our approach involves implementing both loss functions to analyze their effects on model accuracy and robustness, where the Softmax loss function was employed as a comparative baseline. Additionally, we examine how the sizes of mel spectrograms and their varying time lengths influence model performance. The experimental results demonstrate superior identification accuracy compared to traditional Softmax loss methods. Furthermore, we discuss the implications of these findings for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。