arXiv:2410.13445cs.CLcs.AI2024-10

用少量数据提升低资源语音识别效果,结合文本适配与高效微调。

Parameter-efficient Adaptation of Multilingual Multimodal Models for Low-resource ASR

  • 利用多语言多模态模型,通过无标签文本进行文本适配。
  • 在零样本场景下实现17%相对词错误率降低。
  • 适合低资源语言语音识别研究者快速部署模型。

低资源语言的自动语音识别(ASR)因标注数据稀缺而面临挑战。参数高效微调和仅文本适配是两种常用方法。本文研究如何将二者结合,使用类似SeamlessM4T的多语言多模态模型,在无标签文本支持下进行文本适配,并进一步通过参数高效微调提升ASR性能。实验表明,跨语言迁移可在零样本设置下实现相对17%的词错误率(WER)降低,且无需任何标注语音数据。

原文摘要 · Abstract (English)

Automatic speech recognition (ASR) for low-resource languages remains a challenge due to the scarcity of labeled training data. Parameter-efficient fine-tuning and text-only adaptation are two popular methods that have been used to address such low-resource settings. In this work, we investigate how these techniques can be effectively combined using a multilingual multimodal model like SeamlessM4T. Multimodal models are able to leverage unlabeled text via text-only adaptation with further parameter-efficient ASR fine-tuning, thus boosting ASR performance. We also show cross-lingual transfer from a high-resource language, achieving up to a relative 17% WER reduction over a baseline in a zero-shot setting without any labeled speech.

语音识别低资源多模态高效微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。