用生成式方法流式构建并优化对话槽位模式,提升自动识别能力。
Generative Induction of Dialogue Task Schemas with Streaming Refinement and Simulated Interactions
- 将槽位诱导视为文本生成任务,逐条对话流中增量构建和优化槽位模式。
- 在新领域上实现当前最佳性能,关键指标显著优于以往方法。
- 适合对话系统研发者和需自动化设计槽位的场景使用。
在面向任务的对话(TOD)系统中,槽位模式诱导(SSI)对于无需人工干预地从对话数据中自动识别关键信息槽位至关重要。本文提出一种新型最先进的(SoTA)方法,将SSI建模为文本生成任务,使语言模型在流式对话数据中逐步构建并精炼槽位模式。为开发该方法,我们提出一种完全自动化的基于大模型的TOD仿真方法,可生成带有高质量状态标签的新领域对话数据。此外,我们发现现有SSI评估存在数据泄露及指标与人类判断不一致的问题,通过结合人工指导与修正的仿真数据以及改进的评估指标加以解决。这些贡献为未来SSI研究奠定基础,并推动了对话理解与系统开发的最新进展。
原文摘要 · Abstract (English)
In task-oriented dialogue (TOD) systems, Slot Schema Induction (SSI) is essential for automatically identifying key information slots from dialogue data without manual intervention. This paper presents a novel state-of-the-art (SoTA) approach that formulates SSI as a text generation task, where a language model incrementally constructs and refines a slot schema over a stream of dialogue data. To develop this approach, we present a fully automatic LLM-based TOD simulation method that creates data with high-quality state labels for novel task domains. Furthermore, we identify issues in SSI evaluation due to data leakage and poor metric alignment with human judgment. We resolve these by creating new evaluation data using our simulation method with human guidance and correction, as well as designing improved evaluation metrics. These contributions establish a foundation for future SSI research and advance the SoTA in dialogue understanding and system development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。