arXiv:2409.10999cs.CLcs.AI2024-09被引 13

提升音频语言模型在泰语等低资源语言的指令跟随能力

Enhancing Low-Resource Language and Instruction Following Capabilities of Audio Language Models

  • 用混合数据训练模型,兼顾泰语与英语的指令理解
  • 新模型Typhoon-Audio在泰语任务上超越现有开源模型
  • 适合需要多语言语音交互系统的开发者使用

音频语言模型通过文本提示处理音频输入,用于语音识别和音频字幕等任务。尽管基于多语言预训练组件构建,但多数模型主要在英语上训练,限制了其他语言的应用。本文评估了音频语言模型在泰语(低资源语言)上的表现,发现其缺乏涌现的跨语言能力。为解决此问题,我们探索了融合目标语言与英语数据的训练策略,将音频理解与语音指令遵循统一建模。实验表明,通过平衡语言特定与多语言训练数据,可有效提升低资源语言的指令遵循能力。提出的模型Typhoon-Audio在英语和泰语上均显著优于现有开源模型,性能接近最先进的Gemini-1.5-Pro。

原文摘要 · Abstract (English)

Audio language models process audio inputs using textual prompts for tasks like speech recognition and audio captioning. Although built on multilingual pre-trained components, most are trained primarily on English, limiting their usability for other languages. This paper evaluates audio language models on Thai, a low-resource language, and finds that they lack emergent cross-lingual abilities despite their multilingual foundations. To address this, we explore data mixtures that optimize audio language models for both a target language and English while integrating audio comprehension and speech instruction-following into a unified model. Our experiments provide insights into improving instruction-following in low-resource languages by balancing language-specific and multilingual training data. The proposed model, Typhoon-Audio, significantly outperforms existing open-source models and achieves performance comparable to state-of-the-art Gemini-1.5-Pro in both English and Thai.

音频语言模型低资源语言指令跟随多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。