构建首个针对图增强生成的多学科推理评测基准。
GraphRAG-Bench: Challenging Domain-Specific Reasoning for Evaluating Graph Retrieval-Augmented Generation
- 设计大学水平多跳推理题,避免简单检索即可作答。
- 覆盖16个学科20本教材,含多种题型与完整流程评估。
- 适合研究图结构对大模型推理能力提升的学者使用。
图增强生成(GraphRAG)通过结构化领域语料提升大语言模型的复杂推理能力,但现有评估多依赖传统问答数据集,难以全面检验其推理优势。为此,我们提出GraphRAG-Bench,一个大规模、领域特定的评测基准。该基准具备三大优势:(i) 高挑战性题目设计——包含需数学或编程推理的大学级问题,仅靠内容检索无法解答;(ii) 多样化任务覆盖——涵盖选择题、判断题、多选题、开放题和填空题,覆盖16个学科的20本核心教材;(iii) 全流程评估框架——从图构建、知识检索到答案生成,不仅评估最终答案正确率,还分析推理过程的逻辑一致性。我们将九种主流GraphRAG方法应用于该基准,量化了图结构对推理能力的提升效果,并揭示了图架构设计、检索效率与推理表现之间的关键关系,为研究社区提供可操作的改进方向。
原文摘要 · Abstract (English)
Graph Retrieval Augmented Generation (GraphRAG) has garnered increasing recognition for its potential to enhance large language models (LLMs) by structurally organizing domain-specific corpora and facilitating complex reasoning. However, current evaluations of GraphRAG models predominantly rely on traditional question-answering datasets. Their limited scope in questions and evaluation metrics fails to comprehensively assess the reasoning capacity improvements enabled by GraphRAG models. To address this gap, we introduce GraphRAG-Bench, a large-scale, domain-specific benchmark designed to rigorously evaluate GraphRAG models. Our benchmark offers three key superiorities: \((i)\) Challenging question design. Featuring college-level, domain-specific questions that demand multi-hop reasoning, the benchmark ensures that simple content retrieval is insufficient for problem-solving. For example, some questions require mathematical reasoning or programming. \((ii)\) Diverse task coverage. The dataset includes a broad spectrum of reasoning tasks, multiple-choice, true/false, multi-select, open-ended, and fill-in-the-blank. It spans 16 disciplines in twenty core textbooks. \((iii)\) Holistic evaluation framework. GraphRAG-Bench provides comprehensive assessment across the entire GraphRAG pipeline, including graph construction, knowledge retrieval, and answer generation. Beyond final-answer correctness, it evaluates the logical coherence of the reasoning process. By applying nine contemporary GraphRAG methods to GraphRAG-Bench, we demonstrate its utility in quantifying how graph-based structuring improves model reasoning capabilities. Our analysis reveals critical insights about graph architectures, retrieval efficacy, and reasoning capabilities, offering actionable guidance for the research community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。