通过增强对话状态追踪的槽位名称,提升模型在未见领域中的零样本适应能力。
Schema Augmentation for Zero-Shot Domain Adaptation in Dialogue State Tracking
- 在提示中引入槽位名称的变异形式,增强模型泛化能力。
- 在MultiWOZ和SpokenWOZ上实现未见领域准确率超两倍提升。
- 适合需要快速适配新领域的对话系统研发者使用。
对话状态追踪(DST)的零样本领域自适应仍是任务导向对话系统中的难题,要求模型在训练时未见的目标领域上具备泛化能力。当前基于大语言模型的方法依赖提示工程引入目标领域知识,但效果高度依赖提示设计和底层模型的零样本能力。本文提出一种名为Schema Augmentation的新数据增强方法,通过在提示中引入槽位名称的变体,对语言模型进行微调以提升其零样本领域自适应性能。在MultiWOZ和SpokenWOZ上的实验表明,该方法显著优于基线,在部分实验中使未见领域的准确率提升超过两倍,同时在所有领域上保持或超越原有性能。
原文摘要 · Abstract (English)
Zero-shot domain adaptation for dialogue state tracking (DST) remains a challenging problem in task-oriented dialogue (TOD) systems, where models must generalize to target domains unseen at training time. Current large language model approaches for zero-shot domain adaptation rely on prompting to introduce knowledge pertaining to the target domains. However, their efficacy strongly depends on prompt engineering, as well as the zero-shot ability of the underlying language model. In this work, we devise a novel data augmentation approach, Schema Augmentation, that improves the zero-shot domain adaptation of language models through fine-tuning. Schema Augmentation is a simple but effective technique that enhances generalization by introducing variations of slot names within the schema provided in the prompt. Experiments on MultiWOZ and SpokenWOZ showed that the proposed approach resulted in a substantial improvement over the baseline, in some experiments achieving over a twofold accuracy gain over unseen domains while maintaining equal or superior performance over all domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。