构建可调控难度的多跳问答数据集,提升RAG评估可靠性
MHTS: Multi-Hop Tree Structure Framework for Generating Difficulty-Controllable QA Datasets for RAG Evaluation
- 用多跳树结构生成逻辑连贯的多段落问题
- 提出细粒度难度公式,与RAG性能强相关
- 适合需要真实评估推理能力的研究者
现有RAG评测基准常忽视查询难度,导致在简单问题上表现虚高,评估不可靠。一个可靠的评测数据集需满足质量、多样性和难度三个关键标准,以准确捕捉基于多跳推理的复杂性及支撑证据分布。本文提出MHTS(Multi-Hop Tree Structure)框架,通过多跳树结构系统控制多跳推理复杂度,生成逻辑关联、多片段的查询。其细粒度难度估计公式与RAG系统的整体性能指标呈现强相关性,验证了其在评估检索与答案生成能力方面的有效性。该方法确保了高质量、多样化且难度可控的查询,显著提升了RAG的评估与基准测试能力。
原文摘要 · Abstract (English)
Existing RAG benchmarks often overlook query difficulty, leading to inflated performance on simpler questions and unreliable evaluations. A robust benchmark dataset must satisfy three key criteria: quality, diversity, and difficulty, which capturing the complexity of reasoning based on hops and the distribution of supporting evidence. In this paper, we propose MHTS (Multi-Hop Tree Structure), a novel dataset synthesis framework that systematically controls multi-hop reasoning complexity by leveraging a multi-hop tree structure to generate logically connected, multi-chunk queries. Our fine-grained difficulty estimation formula exhibits a strong correlation with the overall performance metrics of a RAG system, validating its effectiveness in assessing both retrieval and answer generation capabilities. By ensuring high-quality, diverse, and difficulty-controlled queries, our approach enhances RAG evaluation and benchmarking capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。