构建可追溯证据的图谱框架,提升长期临床事件验证准确性。
Search Broadly, Seek Evidence on Both Sides, Decide Narrowly: Evidence-Admissible GraphRAG for Longitudinal Clinical Event Verification

- 用患者图谱关联事件与多源证据,支持正反面信息检索。
- 在多个数据集上准确率最高达96.8%,较基线提升超26点。
- 适合需要可解释、高可信医疗决策的临床研究与系统开发。
纵向临床事件关系验证旨在判断患者记录是否支持特定临床事件间的关系。该任务挑战大,因证据分散于结构化记录、病历文本、检验轨迹、就诊记录及时间维度中,且存在否定、时间错配、重复记录和矛盾发现等问题,易导致误判。本文提出MedEventGraph-RAG,一种可接纳证据的框架:将患者个体事件发生构建为图谱,每个事件关联结构化行、文本片段、时间戳与数值轨迹等来源证据。给定包含事件、关系与临床范围的查询,图谱引导候选事件链发现,并从支持与反对两方面检索证据。通过查询定制的证据契约,过滤身份、范围、发生绑定与来源可追溯性后,由独立评估者判定支持、冲突、反驳或证据不足。在i2b2、n2c2、MIMIC-IV和LUNGUAGE上的十项协议测试中,该方法在时间、药物不良反应与记录顺序验证任务上分别达到78.6、67.3、96.8的平衡准确率,优于最强基线26.9、4.9、30.4个百分点。在证据遮蔽下,平衡准确率达92.2且无虚假支持预测;当中间事件隐藏时,可在57.9% i2b2与70.0% LUNGUAGE案例中恢复完整可追溯事件链。结果表明,分离广度证据发现与窄范围证据可接受评估,能显著提升纵向临床验证性能并减少无依据结论。
原文摘要 · Abstract (English)
Longitudinal clinical event-relation verification determines whether a patient record supports a specified relation among two or more clinical events. This task is challenging because evidence is distributed across structured records, notes, laboratory trajectories, encounters, and time, while negation, temporal mismatch, repeated documentation, and conflicting findings can make retrieved information appear relevant without establishing the relation. We present MedEventGraph-RAG, an evidence-admissible framework that represents event occurrences in a patient-specific graph and links each occurrence to source evidence, including structured rows, note spans, timestamps, and numerical trajectories. Given a verification query specifying events, relation, and clinical scope, the graph guides discovery of candidate event chains and retrieves evidence from both supporting and contradicting sides. A query-specific evidence contract filters information by patient identity, scope, occurrence binding, and source traceability before a separate assessor determines supported, conflicting, refuted, or insufficient outcomes. Across ten protocols on i2b2, n2c2, MIMIC-IV, and LUNGUAGE, MedEventGraph-RAG achieves balanced accuracies of 78.6, 67.3, and 96.8 on temporal, medication-adverse-event, and recorded-order verification, improving over the strongest matched baselines by 26.9, 4.9, and 30.4 points. Under evidence masking, it reaches 92.2 balanced accuracy with no false-support predictions. When intermediate events are hidden, it recovers complete source-traceable event chains in 57.9% of i2b2 and 70.0% of LUNGUAGE cases. These results show that separating broad evidence discovery from narrow evidence-admissible assessment improves longitudinal clinical verification and reduces unsupported conclusions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。