通过迭代学习成功对话中的关键子目标,提升任务型对话系统性能。
Learning from Relevant Subgoals in Successful Dialogs using Iterative Training for Task-oriented Dialog Systems
- 从模型生成的对话中自动识别对成功有贡献的子目标
- 在主流基准上达到新的最好水平,优于传统微调和偏好学习
- 适合希望改进对话系统策略的开发者与研究者
任务型对话系统需完成多个子目标才能达成用户目标,但通常只能在对话结束时获得反馈。本文提出SUIT(子目标感知的迭代训练)方法,通过采样待优化模型生成的对话,利用远距离监督识别对对话成功有贡献的子目标,从而生成高质量训练数据。该方法可迭代生成更多数据,无需依赖固定静态数据集,并显著提升监督微调或偏好学习的效果。SUIT在主流任务型对话基准上达到新最优性能。
原文摘要 · Abstract (English)
Task-oriented Dialog (ToD) systems have to solve multiple subgoals to accomplish user goals, whereas feedback is often obtained only at the end of the dialog. In this work, we propose SUIT (SUbgoal-aware ITerative Training), an iterative training approach for improving ToD systems. We sample dialogs from the model we aim to improve and determine subgoals that contribute to dialog success using distant supervision to obtain high quality training samples. We show how this data improves supervised fine-tuning or, alternatively, preference learning results. SUIT is able to iteratively generate more data instead of relying on fixed static sets. SUIT reaches new state-of-the-art performance on a popular ToD benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。