基于循证医学构建运动康复知识图谱,提升医疗问答准确性
From Evidence-Based Medicine to Knowledge Graph: Retrieval-Augmented Generation for Sports Rehabilitation and a Domain Benchmark
- 将PICO框架融入知识图谱构建与检索,实现问题与证据精准对齐
- 提出贝叶斯证据等级重排算法,无需预设权重即可校准证据优先级
- 在运动康复领域发布大规模知识图谱与评测数据集,适合临床辅助系统研发
当前医疗检索增强生成方法忽视循证医学(EBM)原则,存在两大缺陷:(1)查询与检索证据间缺乏PICO框架对齐;(2)重排阶段未考虑证据等级。本文提出SR-RAG,一种适配EBM的GraphRAG框架,将PICO融入知识图谱构建与检索,并提出贝叶斯证据等级重排(BETR),通过证据等级动态校准排序分数,无需预设权重。在运动康复领域验证,发布包含357,844个节点、371,226条边的知识图谱及1,637组问答对的基准数据集。SR-RAG在证据召回率@10达0.812,关键信息覆盖率达0.830,答案忠实度0.819,语义相似度0.882,PICOT匹配准确率0.788,显著优于五种基线模型。五位专家临床医生评分4.66–4.84(5分制),系统排名在人工验证黄金子集(n=80)上保持稳定。
原文摘要 · Abstract (English)
Current medical retrieval-augmented generation (RAG) approaches overlook evidence-based medicine (EBM) principles, leading to two key gaps: (1) the lack of PICO alignment between queries and retrieved evidence, and (2) the absence of evidence hierarchy considerations during reranking. We present SR-RAG, an EBM-adapted GraphRAG framework that integrates the PICO framework into knowledge graph construction and retrieval, and proposes Bayesian Evidence Tier Reranking (BETR) to calibrate ranking scores by evidence grade without predefined weights. Validated in sports rehabilitation, we release a knowledge graph (357,844 nodes, 371,226 edges) and a benchmark of 1,637 QA pairs. SR-RAG achieves 0.812 evidence recall@10, 0.830 nugget coverage, 0.819 answer faithfulness, 0.882 semantic similarity, and 0.788 PICOT match accuracy, substantially outperforming five baselines. Five expert clinicians rated the system 4.66--4.84 on a 5-point Likert scale, and system rankings are preserved on a human-verified gold subset (n=80).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。