小样本语音数据下,大模型微调可高效提升口语理解能力
Exploring Fine-Tuning of Large Audio Language Models for Spoken Language Understanding under Limited Speech Data
- 仅用文本微调已表现优异,体现强泛化能力
- 加入2-5%语音数据即显著提升性能,课程学习更优
- 跨语言场景中少量目标语音+大量文本可有效适配
大型音频语言模型(LALMs)在语音任务中表现出强大能力,但其微调研究仍不充分,尤其是在语音数据有限的情况下。为填补这一空白,我们系统考察了文本微调、直接混合与课程学习等不同微调策略对口语理解(SLU)的影响,重点关注文本标签对丰富而语音-标签配对数据稀缺的场景。结果表明,仅使用文本微调的LALMs已达到具有竞争力的性能,凸显其强大的泛化能力。即使加入少量语音数据(2%-5%),性能也有显著提升,其中课程学习在数据稀缺时尤为有效。在跨语言SLU中,结合源语言语音数据、目标语言文本及少量目标语言语音数据,可实现有效迁移。整体而言,本研究为实际数据受限条件下的LALM微调提供了实用洞见。
原文摘要 · Abstract (English)
Large Audio Language Models (LALMs) have emerged as powerful tools for speech-related tasks but remain underexplored for fine-tuning, especially with limited speech data. To bridge this gap, we systematically examine how different fine-tuning schemes including text-only, direct mixing, and curriculum learning affect spoken language understanding (SLU), focusing on scenarios where text-label pairs are abundant while paired speech-label data are limited. Results show that LALMs already achieve competitive performance with text-only fine-tuning, highlighting their strong generalization ability. Adding even small amounts of speech data (2-5%) yields substantial further gains, with curriculum learning particularly effective under scarce data. In cross-lingual SLU, combining source-language speech data with target-language text and minimal target-language speech data enables effective adaptation. Overall, this study provides practical insights into the LALM fine-tuning under realistic data constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。