用大模型生成自然语言描述对话状态,更准确且易懂。
Interpretable and Robust Dialogue State Tracking via Natural Language Summarization with LLMs
- 用大模型直接生成人类可读的状态描述,替代传统槽值对
- 在MultiWOZ 2.1和Taskmaster-1上联合目标准确率显著更高
- 对噪声鲁棒、结果可解释,适合需要透明性的对话系统
本文提出一种新型对话状态追踪方法(NL-DST),利用大语言模型(LLMs)生成自然语言形式的对话状态描述,突破传统槽值对表示的局限。传统方法在开放域对话和噪声输入下表现不佳。NL-DST框架通过训练大模型直接合成可读的状态描述。在MultiWOZ 2.1和Taskmaster-1数据集上的实验表明,该方法在联合目标准确率和槽位准确率上均显著优于基于规则、BERT判别式及GPT-2生成式槽填充的基线模型。消融研究与人工评估进一步验证了自然语言状态生成的有效性,凸显其对噪声的鲁棒性与更强的可解释性。结果表明,NL-DST为任务导向对话系统提供了更灵活、准确且人类可理解的新范式。
原文摘要 · Abstract (English)
This paper introduces a novel approach to Dialogue State Tracking (DST) that leverages Large Language Models (LLMs) to generate natural language descriptions of dialogue states, moving beyond traditional slot-value representations. Conventional DST methods struggle with open-domain dialogues and noisy inputs. Motivated by the generative capabilities of LLMs, our Natural Language DST (NL-DST) framework trains an LLM to directly synthesize human-readable state descriptions. We demonstrate through extensive experiments on MultiWOZ 2.1 and Taskmaster-1 datasets that NL-DST significantly outperforms rule-based and discriminative BERT-based DST baselines, as well as generative slot-filling GPT-2 DST models, in both Joint Goal Accuracy and Slot Accuracy. Ablation studies and human evaluations further validate the effectiveness of natural language state generation, highlighting its robustness to noise and enhanced interpretability. Our findings suggest that NL-DST offers a more flexible, accurate, and human-understandable approach to dialogue state tracking, paving the way for more robust and adaptable task-oriented dialogue systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。