首个面向东南亚多语言的音频大模型,支持语音理解与对话交互。
SeaLLMs-Audio: Large Audio-Language Models for Southeast Asia
- 构建多语言音频-文本联合模型,支持5种语言输入与混合模态处理。
- 在语音识别、情感分析、语音问答等任务上表现优异,超越多数现有模型。
- 适合研究者与开发者用于东南亚语音应用开发,推动本地化智能语音系统落地。
我们提出 SeaLLMs-Audio,首个专为东南亚多语言(印尼语、泰语、越南语、英语、中文)设计的大规模音频-语言模型(LALM)。该模型基于大规模音频语料训练,在细粒度音频理解与语音交互任务中表现强劲。其核心特性包括:1)多语言支持,覆盖印尼语(id)、泰语(th)、越南语(vi)、英语(en)和中文(zh);2)多模态输入,支持纯音频、纯文本及音频+文本组合输入;3)多任务能力,涵盖音频描述生成、自动语音识别、语音翻译、语音情绪识别、语音问答与语音摘要等任务,并支持基于语音的事实、数学及通用知识问答。为自动化评估东南亚场景下的音频大模型,我们构建了 SeaBench-Audio 基准测试集。实验表明,SeaLLMs-Audio 在东南亚语言上的表现优于或媲美现有 LALM 模型。
原文摘要 · Abstract (English)
We introduce SeaLLMs-Audio, the first large audio-language model (LALM) tailored for multiple Southeast Asian (SEA) languages-Indonesian (id), Thai (th), and Vietnamese (vi)-alongside English (en) and Chinese (zh). Trained on a large-scale audio corpus, SeaLLMs-Audio exhibits strong performance across diverse audio-centric tasks, spanning fine-grained audio understanding and voice-based interaction. Its key features include: 1) Multilingual: the model primarily supports 5 languages, namely Indonesian, Thai, Vietnamese, English, and Chinese; 2) Multimodal: the model accepts flexible input modalities, including audio only, text only, as well as audio with text; 3) Multi-task: the model supports a wide range of tasks, including audio analysis tasks such as Audio Captioning, Automatic Speech Recognition, Speech-to-Text Translation, Speech Emotion Recognition, Speech Question Answering, and Speech Summarization. It also enables voice-based dialogue, including answering factual, mathematical, and general knowledge queries. As a significant step towards advancing audio LLMs in Southeast Asia, we expect SeaLLMs-Audio to benefit both the regional research community and industry. To automate LALM evaluation for Southeast Asia, we introduce SeaBench-Audio, a benchmark spanning multiple tasks. Experiments show that SeaLLMs-Audio achieves competitive performance compared with other LALMs on SEA languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。