针对阿拉伯语语音大模型,提出多任务训练调度策略提升低资源下性能。
Multi-Task Instruction Tuning via Data Scheduling for Low-Resource Arabic SpeechLLMs
- 设计分阶段训练策略,动态调度生成与判别任务
- 新数据集 AraMega-SSum 支持端到端阿拉伯语语音摘要
- 两阶段策略在判别任务上超越商用模型 Gemini-2.5-Pro
语音大语言模型(Audio LLMs)实现了统一的语音理解与生成,但在阿拉伯语-英语这类语言复杂、方言丰富的场景中仍具挑战。本文针对资源受限环境下以阿拉伯语为核心的语音大模型,开展多任务指令微调的受控研究,涵盖生成性任务(如自动语音识别 ASR 及语音/文本摘要)和判别性任务(如方言识别 DID 与语音情感识别 SER)。为支持端到端阿拉伯语语音摘要,我们构建了首个专用于训练与评估阿拉伯语中心语音大模型的语音摘要数据集 AraMega-SSum。比较四种训练策略:(i) 均匀混合(UM)、(ii) 任务渐进式课程学习(TPC)、(iii) 基于对齐器的多样化采样(ADS)用于训练批次构建,以及 (iv) 两阶段策略 TPC→ADS。结果表明存在明显的效率-鲁棒性权衡:TPC 在生成任务(如 ASR 与摘要)中表现最佳;ADS 单独使用可提升副语言任务性能,但降低生成稳定性;而两阶段策略在整体表现上最优,兼具最强的 DID 与 SER 性能,并在判别任务上优于 Gemini-2.5-Pro 等大型商用模型。所有实验资源与 AraMega-SSum 数据集将公开发布,以推动阿拉伯语语音理解研究。
原文摘要 · Abstract (English)
Audio large language models (LLMs) enable unified speech understanding and generation, but adapting them to linguistically complex and dialect-rich settings such as Arabic-English remains challenging. We present a controlled study of multi-task instruction tuning for an Arabic-centric audio LLM across generative tasks, including automatic speech recognition (ASR) and speech and text summarization, as well as discriminative tasks, including dialect identification (DID) and speech emotion recognition (SER), in a resource-constrained setting. To support end-to-end Arabic speech summarization, we introduce AraMega-SSum, the first Arabic speech summarization dataset designed for training and benchmarking Arabic-centric audio LLMs. We compare four training strategies: (i) Uniform Mixing (UM), (ii) Task-Progressive Curriculum (TPC), (iii) Aligner-Based Diverse Sampling (ADS) for training-time batch construction, and (iv) a two-stage TPC->ADS strategy. Our results reveal a clear efficiency-robustness trade-off. TPC achieves the strongest performance on generative tasks, including ASR and summarization. ADS improves paralinguistic tasks but reduces generative stability when used alone. The two-stage TPC->ADS strategy provides the best overall balance, achieving the strongest DID and SER performance while outperforming large proprietary models such as Gemini-2.5-Pro on discriminative tasks. We will publicly release AraMega-SSum together with all experimental resources to support future research in Arabic speech understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。