arXiv:2512.01603cs.CLcs.MM2025-12被引 3

构建汽车舱内多意图语音理解数据集,评测大模型表现。

MAC-SLU: Multi-Intent Automotive Cabin Spoken Language Understanding Benchmark

  • 设计真实复杂的多意图语音数据集,提升任务难度。
  • 大模型微调后性能远超零样本提示,端到端模型避免识别错误传播。
  • 适合语音交互、车载系统研究者参考。

语音理解(SLU)旨在提取用户语义以执行下游任务,是任务导向对话系统的关键组件。现有SLU数据集普遍缺乏多样性和复杂性,且缺少针对最新大语言模型(LLMs)和大音频语言模型(LALMs)的统一基准。本文提出MAC-SLU,一个新型的汽车舱内多意图语音理解数据集,通过引入真实、复杂的多意图数据提升了任务难度。基于该数据集,我们对主流开源LLMs和LALMs进行了全面评估,涵盖上下文学习、监督微调(SFT)、端到端(E2E)与流水线范式。实验表明,尽管LLMs和LALMs在上下文学习下具备完成SLU任务的潜力,但性能仍显著落后于微调方法;而端到端LALMs表现接近流水线方法,并有效避免了语音识别错误传播。代码与数据集已公开于GitHub和Hugging Face。

原文摘要 · Abstract (English)

Spoken Language Understanding (SLU), which aims to extract user semantics to execute downstream tasks, is a crucial component of task-oriented dialog systems. Existing SLU datasets generally lack sufficient diversity and complexity, and there is an absence of a unified benchmark for the latest Large Language Models (LLMs) and Large Audio Language Models (LALMs). This work introduces MAC-SLU, a novel Multi-Intent Automotive Cabin Spoken Language Understanding Dataset, which increases the difficulty of the SLU task by incorporating authentic and complex multi-intent data. Based on MAC-SLU, we conducted a comprehensive benchmark of leading open-source LLMs and LALMs, covering methods like in-context learning, supervised fine-tuning (SFT), and end-to-end (E2E) and pipeline paradigms. Our experiments show that while LLMs and LALMs have the potential to complete SLU tasks through in-context learning, their performance still lags significantly behind SFT. Meanwhile, E2E LALMs demonstrate performance comparable to pipeline approaches and effectively avoid error propagation from speech recognition. Code\footnote{https://github.com/Gatsby-web/MAC\_SLU} and datasets\footnote{huggingface.co/datasets/Gatsby1984/MAC\_SLU} are released publicly.

语音理解车载系统大模型多意图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。