arXiv:2603.22633cs.AIcs.IR2026-03被引 1

提出新框架,让医学文献问答系统不仅能找对段落,还能跨多章节找证据。

Graph-Aware Late Chunking for Retrieval-Augmented Generation in Biomedical Literature

  • 用图结构智能识别文献分段边界,融合医学知识图谱增强检索
  • 在2033个跨章节问题上,结构感知方法覆盖15.6倍更多章节
  • 适合需要全面整合文献证据的科研人员和精准医疗应用

针对生物医学文献的检索增强生成(RAG)系统,传统评估依赖排名指标(如MRR),仅衡量能否找到最相关的一个文本块。我们指出这一范式不完整:它只奖励检索精度,忽略检索广度——即从文档各结构部分(如引言、方法、结果)中获取证据的能力。为此提出GraLC-RAG框架,结合延迟分块与图结构感知,实现结构敏感的分块边界检测、UMLS知识图谱融合及图引导的混合检索。在2,359篇经IMRaD筛选的PubMed Central文章上,使用2,033个跨章节问题进行评估,采用标准排名指标(MRR、Recall@k)和结构覆盖指标(SecCov@k、CS Recall)。结果表明:基于内容相似性的方法获得最高MRR(0.517),但始终仅从单一章节检索;而结构感知方法可覆盖最多15.6倍的章节。生成实验显示,知识图谱注入的检索使答案质量差距缩小至delta-F1 = 0.009,同时保持4.6倍的章节多样性。这说明标准指标系统性低估结构检索价值,且多章节信息整合仍是生物医学RAG的关键挑战。

原文摘要 · Abstract (English)

Retrieval-Augmented Generation (RAG) systems for biomedical literature are typically evaluated using ranking metrics like Mean Reciprocal Rank (MRR), which measure how well the system identifies the single most relevant chunk. We argue that for full-text scientific documents, this paradigm is incomplete: it rewards retrieval precision while ignoring retrieval breadth -- the ability to surface evidence from across a document's structural sections. We propose GraLC-RAG, a framework that unifies late chunking with graph-aware structural intelligence, introducing structure-aware chunk boundary detection, UMLS knowledge graph infusion, and graph-guided hybrid retrieval. We evaluate six strategies on 2,359 IMRaD-filtered PubMed Central articles using 2,033 cross-section questions and two metric families: standard ranking metrics (MRR, Recall@k) and structural coverage metrics (SecCov@k, CS Recall). Our results expose a sharp divergence: content-similarity methods achieve the highest MRR (0.517) but always retrieve from a single section, while structure-aware methods retrieve from up to 15.6x more sections. Generation experiments show that KG-infused retrieval narrows the answer-quality gap to delta-F1 = 0.009 while maintaining 4.6x section diversity. These findings demonstrate that standard metrics systematically undervalue structural retrieval and that closing the multi-section synthesis gap is a key open problem for biomedical RAG.

医学AI检索增强知识图谱文本分块

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。