arXiv:2505.07731cs.CLcs.LG2025-05被引 3

让语音模型零样本理解新任务,无需标注数据

Spoken Language Understanding on Unseen Tasks With In-Context Learning

  • 用随机类别标签做无任务依赖微调,提升泛化能力
  • 在未见过的任务上性能显著优于传统方法
  • 适合缺乏标注数据的语音理解场景

语音语言理解(SLU)任务涵盖多种技能,检验模型的信息抽取、分类和/或生成能力。在该场景下,特定任务的训练数据可能无法获得。传统任务专用的SLU模型难以满足此需求,而语音-文本大语言模型(LLMs)展现出潜在替代方案,具备涌现能力。然而,我们评估发现,主流开源语音-文本LLMs在零样本/少样本设置下的SLU性能尚不理想。本文提出一种新颖的鲁棒性任务无关微调方法,采用随机类别标签。经此微调后,语音-文本LLMs在未见任务上的表现显著优于标准方法。关键在于,该方法无需任务特定数据标注即可实现新任务支持。

原文摘要 · Abstract (English)

Spoken language understanding (SLU) tasks involve diverse skills that probe the information extraction, classification and/or generation capabilities of models. In this setting, task-specific training data may not always be available. While traditional task-specific SLU models are unable to cater to such requirements, the speech-text large language models (LLMs) offer a promising alternative with emergent abilities. However, out of-the-box, our evaluations indicate that the zero/few-shot performance of prominent open-source speech-text LLMs on SLU tasks are not up to the mark. In this paper, we introduce a novel approach to robust task-agnostic fine-tuning using randomized class labels. With this proposed fine-tuning, we illustrate that the performance of the speech-text LLMs on an unseen task is significantly improved over standard approaches. Critically, the proposed approach avoids the requirement of task-specific data annotations for enabling new tasks in speech-text LLMs.

语音理解零样本学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。