让语音大模型像人一样思考,提升对话槽位填充准确率。
Slot Filling as a Reasoning Task for SpeechLLMs
- 用思维链分解语音槽位填充任务,分步推理提升性能。
- 混合模式微调的语音大模型比单一模式表现更好。
- 专门用于数学逻辑的文本模型不适合作为语音推理基础。
我们提出将推理能力融入语音大语言模型(speechLLMs),以实现端到端的槽位填充任务。受近期推理型大语言模型启发,我们采用思维链框架,将槽位填充任务分解为多个推理步骤,构建推理数据集,并对语音大模型进行监督微调。区分普通语音大模型与推理型语音大模型,实验不同类型和规模的LLM作为其文本基础模型。结果表明引入中间推理步骤可显著提升性能。然而,主要面向数学、逻辑和编程领域设计的推理型文本大模型,作为推理型语音大模型的基础时表现反而较差。进一步发现,基于混合文本基础模型并微调以同时保留直接响应与推理两种操作模式的混合语音大模型,性能优于仅使用单一模式微调的模型。
原文摘要 · Abstract (English)
We propose integration of reasoning into speech large language models (speechLLMs) for the end-to-end slot-filling task. Inspired by the recent development of reasoning LLMs, we use a chain-of-thought framework to decompose the slot-filling task into multiple reasoning steps, create a reasoning dataset and apply the supervised fine-tuning strategy to a speechLLM. We distinguish between regular and reasoning speechLLMs and experiment with different types and sizes of LLMs as their text foundation models. We demonstrate performance improvements by introducing reasoning (intermediate) steps. However, we show that a reasoning textual LLM developed mainly for math, logic and coding domains might be inferior as a foundation model for a reasoning speechLLM. We further show that hybrid speechLLMs, built on a hybrid text foundation LLM and fine-tuned to preserve both direct and reasoning modes of operation, have better performance than those fine-tuned employing only one mode of operation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。