通过实证研究给出通用语音识别的最优训练方案。
An Empirical Recipe for Universal Phone Recognition
- 基于大规模多语言数据,优化自监督表征与损失函数
- 在100+语言上实现17.7%的错误率,英式口音下达10.6%
- 公开数据与代码,适合多语种语音研究者使用
语音识别(PR)是多语言和低资源语音处理的关键技术,但鲁棒性能仍难以实现。现有以英语为主的高性能模型难以跨语言泛化,而多语言模型又未能充分使用预训练表示。当前对数据规模、模型架构和训练目标的影响尚不明确。本文提出PhoneticXEUS——在大规模多语言数据上训练,实现多语言语音识别(17.7% PFER)和带口音英语语音识别(10.6% PFER)的最新性能。通过在统一框架下对100+语言进行受控消融实验,我们实证验证了训练方案,并量化了自监督表示、数据规模和损失目标的影响。此外,还分析了不同语系、口音及发音特征下的错误模式。所有数据与代码已开源至https://github.com/changelinglab/PhoneticXeus。
原文摘要 · Abstract (English)
Phone recognition (PR) is a key enabler of multilingual and low-resource speech processing tasks, yet robust performance remains elusive. Highly performant English-focused models do not generalize across languages, while multilingual models underutilize pretrained representations. It also remains unclear how data scale, architecture, and training objective contribute to multilingual PR. We present PhoneticXEUS -- trained on large-scale multilingual data and achieving state-of-the-art performance on both multilingual (17.7% PFER) and accented English speech (10.6% PFER). Through controlled ablations with evaluations across 100+ languages under a unified scheme, we empirically establish our training recipe and quantify the impact of SSL representations, data scale, and loss objectives. In addition, we analyze error patterns across language families, accented speech, and articulatory features. All data and code are released openly at https://github.com/changelinglab/PhoneticXeus
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。