提出模块化框架,系统拆解和优化图增强生成流程。
LEGO-GraphRAG: Modularizing Graph-based Retrieval-Augmented Generation for Design Space Exploration
- 将图增强生成流程细分为可插拔模块,支持灵活组合
- 在真实大规模图上验证,揭示推理质量与效率的权衡关系
- 适合研究图神经网络与大模型融合的开发者使用
GraphRAG 将知识图谱与大语言模型结合,提升推理准确性和上下文相关性。尽管应用前景广阔且跨数据库与自然语言处理多个领域,现有研究仍缺乏模块化工作流分析、系统性解决方案框架及深入的实证研究。为此,我们提出 LEGO-GraphRAG,一个模块化框架,实现:1)对 GraphRAG 流程的细粒度分解;2)现有技术与已实现 GraphRAG 实例的系统分类;3)新 GraphRAG 实例的构建。该框架支持在大规模真实图数据和多样查询集上的全面实证研究,揭示了推理质量、运行效率、令牌开销或 GPU 成本之间的平衡规律,为构建先进 GraphRAG 系统提供关键指导。
原文摘要 · Abstract (English)
GraphRAG integrates (knowledge) graphs with large language models (LLMs) to improve reasoning accuracy and contextual relevance. Despite its promising applications and strong relevance to multiple research communities, such as databases and natural language processing, GraphRAG currently lacks modular workflow analysis, systematic solution frameworks, and insightful empirical studies. To bridge these gaps, we propose LEGO-GraphRAG, a modular framework that enables: 1) fine-grained decomposition of the GraphRAG workflow, 2) systematic classification of existing techniques and implemented GraphRAG instances, and 3) creation of new GraphRAG instances. Our framework facilitates comprehensive empirical studies of GraphRAG on large-scale real-world graphs and diverse query sets, revealing insights into balancing reasoning quality, runtime efficiency, and token or GPU cost, that are essential for building advanced GraphRAG systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。