arXiv:2409.11500cs.CLcs.AI2024-09被引 8

用合成对话提升多文档问答模型性能,效果优于真实数据训练。

Multi-Document Grounded Multi-Turn Synthetic Dialog Generation

  • 用思维链生成有逻辑的用户问题,控制对话流程
  • 每轮对话后更新检索文档,确保信息准确
  • 用大模型自评过滤错误答案,保证数据质量

我们提出一种多文档接地的多轮合成对话生成技术,包含三个核心思想:首先,通过思维链提示生成受分类体系引导的用户问题,以控制对话整体流向;其次,模拟真实检索器使用方式,在每轮用户提问后更新接地文档,支持多文档信息融合;第三,利用大模型作为评判者筛选出答案错误的问题。人类评估表明,合成对话数据具有多样性、连贯性,且多数答案正确。对可回答问题的人类与自动评估均显示,基于合成对话数据微调的模型,在四个公开的多轮文档接地基准测试集上,表现持续优于基于现有真人标注数据微调的模型。

原文摘要 · Abstract (English)

We introduce a technique for multi-document grounded multi-turn synthetic dialog generation that incorporates three main ideas. First, we control the overall dialog flow using taxonomy-driven user queries that are generated with Chain-of-Thought (CoT) prompting. Second, we support the generation of multi-document grounded dialogs by mimicking real-world use of retrievers to update the grounding documents after every user-turn in the dialog. Third, we apply LLM-as-a-Judge to filter out queries with incorrect answers. Human evaluation of the synthetic dialog data suggests that the data is diverse, coherent, and includes mostly correct answers. Both human and automatic evaluations of answerable queries indicate that models fine-tuned on synthetic dialogs consistently out-perform those fine-tuned on existing human generated training data across four publicly available multi-turn document grounded benchmark test sets.

对话生成多文档问答合成数据大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。