arXiv:2505.15087cs.CL2025-05ACL被引 2

自动生成高质量多跳问答题,无需人工干预。

HopWeaver: Cross-Document Synthesis of High-Quality and Authentic Multi-Hop Questions

  • 跨文档识别互补信息,构建真实推理路径
  • 生成问题质量接近人工标注,成本更低
  • 适合研究者快速构建评测基准或训练数据

多跳问答(MHQA)对评估模型整合异源信息的能力至关重要。然而,构建大规模高质量的MHQA数据集面临挑战:一是人工标注成本高,二是现有合成方法常产生简单问题或需大量人工指导。本文提出HopWeaver,首个无需人工干预的跨文档多跳问题生成框架。该框架通过创新管道识别互补文档,并构建真实推理路径,以生成桥接型与对比型问题。我们还构建了全面的评估体系。实证表明,合成问题在质量上可媲美甚至超越人工标注数据,且成本更低。该框架可从任意原始语料库自动构建挑战性评测基准,为研究社区提供高效工具,尤其适用于资源稀缺领域,推动问答模型推理能力的评估与训练。

原文摘要 · Abstract (English)

Multi-Hop Question Answering (MHQA) is crucial for evaluating the model's capability to integrate information from diverse sources. However, creating extensive and high-quality MHQA datasets is challenging: (i) manual annotation is expensive, and (ii) current synthesis methods often produce simplistic questions or require extensive manual guidance. This paper introduces HopWeaver, the first cross-document framework synthesizing authentic multi-hop questions without human intervention. HopWeaver synthesizes bridge and comparison questions through an innovative pipeline that identifies complementary documents and constructs authentic reasoning paths to ensure true multi-hop reasoning. We further present a comprehensive system for evaluating the synthesized multi-hop questions. Empirical evaluations demonstrate that the synthesized questions achieve comparable or superior quality to human-annotated datasets at a lower cost. Our framework provides a valuable tool for the research community: it can automatically generate challenging benchmarks from any raw corpus, which opens new avenues for both evaluation and targeted training to improve the reasoning capabilities of advanced question answering models, especially in domains with scarce resources.

多跳问答自动构建推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。