提出语音函数调用新范式,提升大模型在开放域任务中的语义理解能力
Spoken Function Calling: A New Perspective on Spoken Language Understanding for Large Audio Language Models

- 设计结构化规则的语音函数调用框架,替代传统模糊规则
- 在SFC-Bench数据集上,大语言模型与音频大模型性能显著优于传统SLU
- 适合研究开放域对话系统、多模态大模型语义理解的学者使用
语音理解(SLU)是任务导向对话系统的核心,也是实现人机无缝交互的关键。传统SLU在领域内监督微调后可有效提取封闭域任务的用户语义,但在开放域任务中因规则定义模糊,难以利用上下文学习。本文提出语音函数调用(SFC),一种基于结构化规则定义的新语义理解范式,以突破传统封闭域SLU的局限。我们基于传统SLU数据集构建并扩展了一套语音函数,通过多智能体系统合成SFC-Bench数据集,评估大语言模型(LLMs)和大音频语言模型(LALMs)的表现,并通过后训练增强LALMs的SFC能力。实验表明,SFC在语义提取准确率上显著优于传统SLU,大幅提升LLMs与LALMs在开放域任务中的表现。
原文摘要 · Abstract (English)
Spoken Language Understanding (SLU) is the core component of task-oriented dialogue systems and a pivotal link in achieving seamless human-agent interaction. While traditional SLU can effectively extract user semantics for closed-set tasks after in-domain supervised fine-tuning, it faces significant challenges in leveraging in-context learning for open-domain tasks due to its ambiguous rule definitions. This work proposes Spoken Function Calling (SFC), a novel semantic understanding perspective that optimizes semantic understanding with structured rule definitions, to evolve beyond traditional closed-set SLU. Specifically, we curate and extend a suite of spoken functions based on traditional SLU datasets, construct a multi-agent system to synthesize the SFC-Bench dataset, evaluate the performance of Large Language Models (LLMs) and Large Audio Language Models (LALMs), and enhance the SFC capabilities of LALMs through post-training. Experiments demonstrate that SFC outperforms traditional SLU, substantially enhancing the semantic extraction accuracy for LLMs and LALMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。