提出无偏评估框架,揭示现有GraphRAG性能提升被高估。
How Significant Are the Real Performance Gains? An Unbiased Evaluation Framework for GraphRAG
- 用图-文本-基础的问题生成法,确保问题与数据相关
- 实测3种GraphRAG方法性能提升远低于先前报告
- 适合关注真实性能评估的研究者和开发者
通过从知识图谱中检索上下文,基于图的检索增强生成(GraphRAG)提升了大语言模型(LLMs)回答用户问题的质量。尽管已有多种GraphRAG方法被提出并报告了优异的问答质量表现,但我们发现当前的GraphRAG评估框架存在两个关键缺陷:问题无关性与评估偏差,可能导致性能结论失真。为此,我们提出一个无偏评估框架,采用图-文本-基础的问题生成方法生成更贴近底层数据集的问题,并通过无偏评估流程消除基于LLM的答案评估偏差。我们将该框架应用于评估3种代表性GraphRAG方法,结果表明其性能增益远低于此前报道。尽管本框架仍可能存在局限,但呼吁开展科学化评估,为GraphRAG研究奠定坚实基础。
原文摘要 · Abstract (English)
By retrieving contexts from knowledge graphs, graph-based retrieval-augmented generation (GraphRAG) enhances large language models (LLMs) to generate quality answers for user questions. Many GraphRAG methods have been proposed and reported inspiring performance in answer quality. However, we observe that the current answer evaluation framework for GraphRAG has two critical flaws, i.e., unrelated questions and evaluation biases, which may lead to biased or even wrong conclusions on performance. To tackle the two flaws, we propose an unbiased evaluation framework that uses graph-text-grounded question generation to produce questions that are more related to the underlying dataset and an unbiased evaluation procedure to eliminate the biases in LLM-based answer assessment. We apply our unbiased framework to evaluate 3 representative GraphRAG methods and find that their performance gains are much more moderate than reported previously. Although our evaluation framework may still have flaws, it calls for scientific evaluations to lay solid foundations for GraphRAG research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。