arXiv:2507.17399cs.CLcs.AI2025-07中稿 · SIGIR 2025 LiveRAG…

将GeAR图检索增强生成扩展至百万级文档,验证其通用性。

Millions of $\text{GeAR}$-s: Extending GraphRAG to Millions of Documents

  • 基于实体关系构建文档图谱,提升跨文档检索能力
  • 在SIGIR 2025 LiveRAG挑战中验证百万文档场景下的性能
  • 为图结构RAG提供可扩展范例,适合大规模知识系统开发者

近期研究探索了基于图的检索增强生成方法,利用从文档中提取的实体及其关系等结构化信息来提升检索效果。然而,这些方法通常针对特定任务(如多跳问答、查询聚焦摘要)设计,缺乏在更广泛数据集上的通用性证据。本文旨在适配最先进的图基RAG方案GeAR,探究其在SIGIR 2025 LiveRAG挑战中的表现与局限性,评估其在百万级文档场景下的适用性。

原文摘要 · Abstract (English)

Recent studies have explored graph-based approaches to retrieval-augmented generation, leveraging structured or semi-structured information -- such as entities and their relations extracted from documents -- to enhance retrieval. However, these methods are typically designed to address specific tasks, such as multi-hop question answering and query-focused summarisation, and therefore, there is limited evidence of their general applicability across broader datasets. In this paper, we aim to adapt a state-of-the-art graph-based RAG solution: $\text{GeAR}$ and explore its performance and limitations on the SIGIR 2025 LiveRAG Challenge.

图神经网络RAG信息检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。