用排序模型选最佳语种迁移,提升低资源语音识别效果
DonorRank: Donor Language Selection for Low-Resource Cross-Lingual Speech Recognition

- 基于学习排序框架预测跨语言迁移的有效语种
- 在印地与非洲语系数据上显著优于按亲缘关系选源语言的常规方法
- 可分析迁移规律,指导多语言低资源语音识别实践
低资源自动语音识别常依赖跨语言迁移,即从高资源语种迁移到目标语种。然而,面对资源匮乏语言社区的口语数据,因语言差异、书写系统演变和资源分布不均,选择合适的源语言仍具挑战。本文提出 DonorRank,一种用于零样本语音识别的语种迁移学习排序框架。我们在印地语与非洲语系的两个多语语音语料库上评估该方法,结果表明其能准确预测源语言的优劣排序,并显著优于基于语系亲缘性或高资源语种的启发式策略。此外,该框架还具备分析源语言选择机制的能力:源语言组合决定了哪些语言特征有助于成功迁移。我们进一步揭示了迁移模式,为低资源多语言语音识别提供了实用指导。
原文摘要 · Abstract (English)
Low-resource automatic speech recognition (ASR) commonly relies on cross-lingual transfer, where models are adapted from higher-resource donor languages. However, selecting donors remains challenging for spontaneous speech from under-resourced language communities, due to linguistic variation, evolving orthographic conventions, and uneven resource availability. We present DonorRank, a learning-to-rank framework for predicting effective donor languages for zero-shot ASR. We evaluate DonorRank on two multilingual speech corpora of Indic and African language families. It accurately predicts donor language rankings and improves donor selection over common heuristics based on genetic similarity or high-resource languages. Beyond improving transfer, we show how DonorRank is a general framework for analyzing donor language selection itself. Our analyses show that the composition of the donor set determines which linguistic cues are useful in predicting successful transfer. We also identify transfer patterns that provide practical guidance for multilingual ASR in low-resource settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。