arXiv:2504.14375cs.LG2025-04NAACL被引 5

用分步生成法提升对话数据质量,更可控更真实。

Bottom-Up Synthesis of Knowledge-Grounded Task-Oriented Dialogues with Iteratively Self-Refined Prompts

  • 先生成问答对再拼成对话,分步控制内容
  • 人工与自动评估均显示对话更真实高质量
  • 适合需要精准控制的对话系统训练场景

训练对话式问答系统需大量领域内数据,但实际中常稀缺。传统方法多采用自上而下的大模型生成多轮对话,虽连贯但难以精细控制且易幻觉。本文提出自下而上的对话合成方法:先生成问答对,再组合成连贯对话。该方法将流程分为两步,实现指令优化与验证分离,提升控制精度。同时,在不涉及专有知识的阶段可使用非本地模型,进一步提高生成数据质量。人工与自动化评估均表明,该方法生成的对话比自上而下方法更真实、更高质量。

原文摘要 · Abstract (English)

Training conversational question-answering (QA) systems requires a substantial amount of in-domain data, which is often scarce in practice. A common solution to this challenge is to generate synthetic data. Traditional methods typically follow a top-down approach, where a large language model (LLM) generates multi-turn dialogues from a broad prompt. Although this method produces coherent conversations, it offers limited fine-grained control over the content and is susceptible to hallucinations. We introduce a bottom-up conversation synthesis approach, where QA pairs are generated first and then combined into a coherent dialogue. This method offers greater control and precision by dividing the process into two distinct steps, allowing refined instructions and validations to be handled separately. Additionally, this structure allows the use of non-local models in stages that do not involve proprietary knowledge, enhancing the overall quality of the generated data. Both human and automated evaluations demonstrate that our approach produces more realistic and higher-quality dialogues compared to top-down methods.

对话生成数据合成LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。