arXiv:2604.17114cs.CL2026-04被引 1

构建可追溯证据的时序知识图谱,让罕见病AI推理结果可验证。

The Provenance Gap in Clinical AI: Evidence-Traceable Temporal Knowledge Graphs for Rare Disease Reasoning

论文配图:The Provenance Gap in Clinical AI: Evidence-Traceable Temporal Knowledge Graphs for Rare Disease Reasoning
图 1 · 摘自论文原文
  • 用4512篇文献构建分层证据图谱,标注1280个疾病轨迹节点。
  • 在36个临床场景中实现100%引用可验证,远超模型自动生成的15.3%。
  • 系统本地部署,保护患者数据隐私,适合医院临床决策支持。

前沿大语言模型虽能生成临床准确内容,但常虚构参考文献,形成“溯源鸿沟”。我们在36个由医生验证的罕见神经肌肉疾病场景中测试了五种前沿LLM,结果发现:无提示时所有模型均无法生成有效PubMed标识符;即使要求引用,表现最佳者仅15.3%的引用与临床相关,多数指向无关领域的真实论文。为此,我们提出HEG-TKG(分层证据锚定时间知识图谱),基于4512篇PubMed文献与人工校验资料,构建含1280个疾病轨迹里程碑的图谱。在相同合成模型下,HEG-TKG在保持原有临床特征覆盖度的同时,实现100%可验证引用,含203个内联引用。相比之下,仅提供原始文本的Guideline-RAG产生零可验证引用。独立临床评估显示,其可验证性优势显著(Cohen's d = 1.81, p < 0.001),且不影响安全性和完整性。反事实实验表明,该系统对注入的临床错误具有80%抵抗能力,并能100%通过引用追踪检测。系统采用开源模型本地部署,确保患者数据不出机构。

原文摘要 · Abstract (English)

Frontier large language models generate clinically accurate outputs, but their citations are often fabricated. We term this the Provenance Gap. We tested five frontier LLMs across 36 clinician-validated scenarios for three rare neuromuscular disease pairs. No model produced a clinically relevant PubMed identifier without prompting. When explicitly asked to cite, the best model achieved 15.3% relevant PMIDs; the majority resolved to real publications in unrelated fields. We present HEG-TKG (Hierarchical Evidence-Grounded Temporal Knowledge Graphs), a system that grounds clinical claims in temporal knowledge graphs built from 4,512 PubMed records and curated sources with quality-tier stratification and 1,280 disease-trajectory milestones. In a controlled three-arm comparison using the same synthesis model, HEG-TKG matches baseline clinical feature coverage while achieving 100% evidence verifiability with 203 inline citations. Guideline-RAG, given overlapping source documents as raw text, produces zero verifiable citations. LLM judges cannot distinguish fabricated from verified citations without PubMed audit data. Independent clinician evaluation confirms the verifiability advantage (Cohen's d = 1.81, p < 0.001) with no degradation on safety or completeness. A counterfactual experiment shows 80% resistance to injected clinical errors with 100% detectability via citation trace. The system deploys on-premise via open-source models so patient data never leaves institutional infrastructure.

临床AI知识图谱可解释性罕见病

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。