arXiv:2509.14804cs.SDeess.AS2025-09被引 2

首个泰语自监督语音编码器,助力低资源语言语音大模型多任务理解。

Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages

  • 用3.6万小时泰语数据训练出首个泰语自监督语音编码器XLSR-Thai
  • 提出更高效多任务的U-Align对齐方法,降低计算成本
  • 构建超1000小时泰语语音理解数据集,推动低资源语言研究

基于语音编码器、适配器和大语言模型的语音大模型(SLLM)在英语、中文等高资源语言中表现出色,但在泰语等低资源语言中性能显著下降。原因有三:(1)现有语音编码器如Whisper家族在低资源语言中表现不佳,且不支持广泛的语言理解任务;(2)基于ASR的对齐范式需全模型微调,计算开销大;(3)低资源语言中配对语音-文本数据稀缺。为解决泰语中的问题,我们提出XLSR-Thai,首个基于36,000小时泰语语音数据持续训练的自监督学习语音编码器。同时提出U-Align方法,比传统ASR对齐更高效且支持多任务。最后构建Thai-SUP数据生成流水线,生成首个超过1,000小时的泰语语音理解数据集。实验验证了方法的有效性。我们开源XLSR-Thai与Thai-SUP以促进后续研究。

原文摘要 · Abstract (English)

Speech large language models (SLLMs) built on speech encoders, adapters, and LLMs demonstrate remarkable multitask understanding performance in high-resource languages such as English and Chinese. However, their effectiveness substantially degrades in low-resource languages such as Thai. This limitation arises from three factors: (1) existing commonly used speech encoders, like the Whisper family, underperform in low-resource languages and lack support for broader spoken language understanding tasks; (2) the ASR-based alignment paradigm requires training the entire SLLM, leading to high computational cost; (3) paired speech-text data in low-resource languages is scarce. To overcome these challenges in the low-resource language Thai, we introduce XLSR-Thai, the first self-supervised learning (SSL) speech encoder for Thai. It is obtained by continuously training the standard SSL XLSR model on 36,000 hours of Thai speech data. Furthermore, we propose U-Align, a speech-text alignment method that is more resource-efficient and multitask-effective than typical ASR-based alignment. Finally, we present Thai-SUP, a pipeline for generating Thai spoken language understanding data from high-resource languages, yielding the first Thai spoken language understanding dataset of over 1,000 hours. Multiple experiments demonstrate the effectiveness of our methods in building a Thai multitask-understanding SLLM. We open-source XLSR-Thai and Thai-SUP to facilitate future research.

语音大模型低资源语言自监督学习泰语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。