用完整语音对话历史提升对话状态追踪效果
The Speech-LLM Takes It All: A Truly Fully End-to-End Spoken Dialogue State Tracking Approach
- 直接输入完整语音对话历史,替代传统文本+语音混合方式
- 在SpokenWOZ数据集上性能超越已有方法,准确率显著提升
- 语音历史压缩后仍保持高精度,适合资源受限场景
本文针对基于语音大模型(Speech-LLM)的端到端语音对话状态追踪,系统评估了三种上下文管理策略:传统多模态上下文(文本历史+当前语音回合)、完整语音历史和压缩语音历史。在SpokenWOZ数据集上的实验表明,将完整语音对话作为输入能显著提升模型性能,优于同规模模型的先前方法。此外,基于注意力池化的语音历史压缩方法在大幅减少上下文长度的同时,仍保持了竞争力的准确率。详细分析证实,性能提升源于更高效的上下文利用。
原文摘要 · Abstract (English)
This paper presents a comparative study of context management strategies for end-to-end Spoken Dialog State Tracking using Speech-LLMs. We systematically evaluate traditional multimodal context (combining text history and spoken current turn), full spoken history, and compressed spoken history approaches. Our experiments on the SpokenWOZ corpus demonstrate that providing the full spoken conversation as input yields the highest performance among models of similar size, significantly surpassing prior methods. Furthermore, we show that attention-pooling-based compression of the spoken history offers a strong trade-off, maintaining competitive accuracy with reduced context size. Detailed analysis confirms that improvements stem from more effective context utilization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。