用Zipformer模型提升双语儿童语音中的语言识别准确率
Leveraging Zipformer Model for Effective Language Identification in Code-Switched Child-Directed Speech
- 采用Zipformer内部层提取语音语言特征嵌入
- 在不平衡数据下实现81.89%的平衡准确率,提升15.47%
- 适合处理双语儿童语音场景的语言识别研究者
在双语环境下的儿童语音中,代码转换与语言识别面临显著挑战。本文利用Zipformer模型处理包含中文和英文两种语言且不均衡的语音片段。研究表明,Zipformer内部各层能有效编码语言特征,可被用于语言识别任务。本文提出内层选择方法以提取嵌入,并对比了不同后端模块的表现。分析显示,该模型在各类后端下均表现稳健。所提方法有效应对数据不平衡问题,在测试集上达到81.89%的平衡准确率(BAC),较语言识别基线提升15.47%。结果表明,该架构在真实场景中具备应用潜力。
原文摘要 · Abstract (English)
Code-switching and language identification in child-directed scenarios present significant challenges, particularly in bilingual environments. This paper addresses this challenge by using Zipformer to handle the nuances of speech, which contains two imbalanced languages, Mandarin and English, in an utterance. This work demonstrates that the internal layers of the Zipformer effectively encode the language characteristics, which can be leveraged in language identification. We present the selection methodology of the inner layers to extract the embeddings and make a comparison with different back-ends. Our analysis shows that Zipformer is robust across these backends. Our approach effectively handles imbalanced data, achieving a Balanced Accuracy (BAC) of 81.89%, a 15.47% improvement over the language identification baseline. These findings highlight the potential of the transformer encoder architecture model in real scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。