arXiv:2603.18361cs.CL2026-03ACL被引 1

用合成数据提升对话模型的常识推理多样性与质量

Synthetic Data Generation for Training Diversified Commonsense Reasoning Models

  • 分两阶段生成合成数据,解决高质量多样数据稀缺问题
  • 在不同规模大模型上,生成多样性和质量均显著提升
  • 适合需要多样化对话输出的智能客服、聊天机器人场景

对话系统不仅需生成具备常识性的高质量回复,还需考虑多种合理情境以体现响应多样性。尽管对多样化常识生成模型的需求日益增长,但该领域进展受限于缺乏大规模高质量的多样化训练数据。现有生成式常识推理(GCR)数据集因标注成本高,仅由少量人工标注者创建,覆盖场景有限。为此,我们提出一种两阶段方法,构建首个合成数据集CommonSyn,用于训练多样化常识推理模型。在不同规模的大语言模型上微调后,该模型在生成多样性与质量方面均优于基线模型及基于人工数据集训练的模型。

原文摘要 · Abstract (English)

Conversational agents are required to respond to their users not only with high quality (i.e. commonsense bearing) responses, but also considering multiple plausible alternative scenarios, reflecting the diversity in their responses. Despite the growing need to train diverse commonsense generators, the progress of this line of work has been significantly hindered by the lack of large-scale high-quality diverse commonsense training datasets. Due to the high annotation costs, existing Generative Commonsense Reasoning (GCR) datasets are created using a small number of human annotators, covering only a narrow set of commonsense scenarios. To address this training resource gap, we propose a two-stage method to create the first-ever synthetic dataset CommonSyn for diversified (GCR). The model fine-tuned on our synthetic data jointly increase both generation diversity and quality compared with vanilla models and the model fine-tuned on human-crafted dataset across different size Large Language Models (LLMs)

常识推理合成数据对话系统多样性生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。