无需微调,用少量数据即可让语音识别模型跨语言适应低资源语言。
SMILE: Speech Meta In-Context Learning for Low-Resource Language Automatic Speech Recognition
- 基于元学习和上下文学习,从高资源语言迁移能力到低资源语言。
- 在无训练场景下实现少样本多语言语音识别,字符与词错误率显著降低。
- 适合低资源语言语音识别研究者及工业界快速部署场景。
自动语音识别(ASR)模型在高资源语言上表现优异,但在低资源语言上因训练数据有限和跨语言泛化能力不足而面临挑战。现有适配方法如浅层融合、数据增强和直接微调,或依赖外部资源,或计算效率低下,或无法在测试时适应。为此,我们提出语音元上下文学习(SMILE),将元学习与语音上下文学习(SICL)结合。SMILE通过高资源语言的元训练,实现对低资源语言的鲁棒少样本泛化,无需在目标领域显式微调。在ML-SUPERB基准上的大量实验表明,SMILE持续优于基线方法,在无训练的少样本多语言ASR任务中显著降低字符错误率(CER)和词错误率(WER)。
原文摘要 · Abstract (English)
Automatic Speech Recognition (ASR) models demonstrate outstanding performance on high-resource languages but face significant challenges when applied to low-resource languages due to limited training data and insufficient cross-lingual generalization. Existing adaptation strategies, such as shallow fusion, data augmentation, and direct fine-tuning, either rely on external resources, suffer computational inefficiencies, or fail in test-time adaptation scenarios. To address these limitations, we introduce Speech Meta In-Context LEarning (SMILE), an innovative framework that combines meta-learning with speech in-context learning (SICL). SMILE leverages meta-training from high-resource languages to enable robust, few-shot generalization to low-resource languages without explicit fine-tuning on the target domain. Extensive experiments on the ML-SUPERB benchmark show that SMILE consistently outperforms baseline methods, significantly reducing character and word error rates in training-free few-shot multilingual ASR tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。