arXiv:2502.17268cs.CL2025-02NAACL

用德语独白生成对话数据,解决任务型对话系统训练数据少的问题。

MonoTODia: Translating Monologue Requests to Task-Oriented Dialogues

  • 将邮件独白改写为对话格式,用大模型自动转换并标注。
  • 人工评估确认生成对话质量高,适合用于训练任务型对话系统。
  • 公开数据集,助力后续研究,尤其适合资源有限的公司使用。

基于Transformer的模型在实际应用中面临数据稀缺问题,尤其是任务型对话(TOD)系统需要专门标注的数据集,而这类数据通常难以获取。本研究提出一种新方法:从现有德语独白材料中提取并转换为适用于训练TOD系统的对话数据。以一家通过电子邮件提供旅行预订服务的公司为例,我们微调了先进的大语言模型,将邮件内容重写为对话形式并进行标注。为保证数据质量和有效性,我们邀请众包工作者对生成的对话进行多维度评估,并建立测试集的黄金标准标注。实验表明,生成的对话和标注质量良好,可作为训练TOD系统的可靠起点。最后,该标注数据集已公开,以促进未来研究。

原文摘要 · Abstract (English)

Data scarcity is one of the main problems when it comes to real-world applications of transformer-based models. This is especially evident for task-oriented dialogue (TOD) systems, which require specialized datasets, that are usually not readily available. This can hinder companies from adding TOD systems to their services. This study therefore investigates a novel approach to sourcing annotated dialogues from existing German monologue material. Focusing on a real-world example, we investigate whether these monologues can be transformed into dialogue formats suitable for training TOD systems. We show the approach with the concrete example of a company specializing in travel bookings via e-mail. We fine-tune state-of-the-art Large Language Models for the task of rewriting e-mails as dialogues and annotating them. To ensure the quality and validity of the generated data, we employ crowd workers to evaluate the dialogues across multiple criteria and to provide gold-standard annotations for the test dataset. We further evaluate the usefulness of the dialogues for training TOD systems. Our evaluation shows that the dialogues and annotations are of high quality and can serve as a valuable starting point for training TOD systems. Finally, we make the annotated dataset publicly available to foster future research.

任务型对话数据生成大模型应用德国语料

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。