arXiv:2509.15362cs.CLcs.SD2025-09被引 2

首个为沃洛夫语训练的语音大模型,提升听写与翻译能力。

Speech Language Models for Under-Represented Languages: Insights from Wolof

  • 用自发口语数据持续预训练HuBERT,提升语音识别性能。
  • 首次构建沃洛夫语语音大模型,支持语音翻译等任务。
  • 引入多步思维链,让模型先推理再转录,效果更优。

本文介绍了为西非沃洛夫语训练语音语言模型的历程,并分享关键洞察。首先强调收集大规模、自然、高质量无监督语音数据的重要性,结果显示在该数据集上持续预训练HuBERT,优于基础模型和非洲语种专用模型的自动语音识别表现。随后,将该语音编码器集成到沃洛夫语语言模型中,训练出首个针对该语言的语音大模型,拓展其能力至语音翻译等任务。此外,探索在转录或翻译前让语音大模型执行多步思维链推理,结果表明该模型不仅提升语音识别效果,也在语音翻译任务中表现良好。模型与代码将公开共享。

原文摘要 · Abstract (English)

We present our journey in training a speech language model for Wolof, an underrepresented language spoken in West Africa, and share key insights. We first emphasize the importance of collecting large-scale, spontaneous, high-quality unsupervised speech data, and show that continued pretraining HuBERT on this dataset outperforms both the base model and African-centric models on ASR. We then integrate this speech encoder into a Wolof LLM to train the first Speech LLM for this language, extending its capabilities to tasks such as speech translation. Furthermore, we explore training the Speech LLM to perform multi-step Chain-of-Thought before transcribing or translating. Our results show that the Speech LLM not only improves speech recognition but also performs well in speech translation. The models and the code will be openly shared.

语音大模型低资源语言沃洛夫语多步推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。