用大模型生成真实对话中的沟通错误,提升基准数据集的现实性
CoPrUS: Consistency Preserving Utterance Synthesis towards more realistic benchmark dialogues
- 基于语言学理论设计三类沟通错误,用大模型分两步生成错误与修复语句
- 在MultiWOZ数据集上修改近1900条对话,生成的错误语句经人工评估质量良好
- 适合研究对话系统鲁棒性、错误容忍能力的研究者使用
大规模的Wizard-Of-Oz对话数据集已推动基于深度学习的对话系统训练。尽管它们作为基准数据集表现良好,但缺乏某些真实对话中常见的语句类型,导致其真实性不足。本文研究在自动流程中生成合成通信错误。基于语言学理论,提出并遵循一个简单的错误分类体系,聚焦三类真实对话中存在但基准数据集中被低估的误沟通:误解、无法理解及模糊相关问题。采用两步法,先用先进大语言模型(LLM)生成错误语句,再生成修复语句。通过语言模型评估确保生成语句质量。将方法应用于MultiWOZ数据集,进行定性和实证评估,并由人工评判。结果表明,当前大语言模型可有效为基准数据集添加后置误沟通,实现数据增强。我们发布改进后的数据集CoPrUS-MultiWOZ,其中近1900条对话已被修改,以促进未来对话系统研究。
原文摘要 · Abstract (English)
Large-scale Wizard-Of-Oz dialogue datasets have enabled the training of deep learning-based dialogue systems. While they are successful as benchmark datasets, they lack certain types of utterances, which would make them more realistic. In this work, we investigate the creation of synthetic communication errors in an automatic pipeline. Based on linguistic theory, we propose and follow a simple error taxonomy. We focus on three types of miscommunications that could happen in real-world dialogues but are underrepresented in the benchmark dataset: misunderstandings, non-understandings and vaguely related questions. Our two-step approach uses a state-of-the-art Large Language Model (LLM) to first create the error and secondly the repairing utterance. We perform Language Model-based evaluation to ensure the quality of the generated utterances. We apply the method to the MultiWOZ dataset and evaluate it both qualitatively and empirically as well as with human judges. Our results indicate that current LLMs can aid in adding post-hoc miscommunications to benchmark datasets as a form of data augmentation. We publish the resulting dataset, in which nearly 1900 dialogues have been modified, as CoPrUS-MultiWOZ to facilitate future work on dialogue systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。