arXiv:2605.22734cs.CL2026-05

构建首个带时间轴的医学知识图谱,助力临床诊断时序推理。

ChronoMedKG: A Temporally-Grounded Biomedical Knowledge Graph and Benchmark for Clinical Reasoning

论文配图:ChronoMedKG: A Temporally-Grounded Biomedical Knowledge Graph and Benchmark for Clinical Reasoning
图 1 · 摘自论文原文
  • 用多代理大模型从文献中提取带时间信息的疾病关联
  • 覆盖1.3万种疾病,6250种新增疾病有时间标注
  • 适合开发精准时序医疗问答与辅助决策系统

现有生物医学知识图谱将疾病关联视为静态事实,但临床推理需考虑时间因素——例如3岁出现的症状可能提示一种疾病,而13岁时则指向另一种。现有知识图谱如PrimeKG、Hetionet和iKraph未记录发现的临床相关时间节点,限制了其在纵向推理与检索增强中的应用。本文提出ChronoMedKG,一个包含460,497条经证据验证三元组(源自1300万条原始提取)的时序医学知识图谱,覆盖13,431种疾病。每条关联均附带发病窗口或进展阶段等时间信息,并有可追溯至PMID的证据和多信号可信度评分。图谱通过疾病自主的多代理流水线构建,多个前沿大模型独立从PubMed和PMC提取知识,仅保留多模型共识且通过可信度过滤与本体对齐的关系。ChronoMedKG与Orphadata相比达成92.7%一致性,为6,250种原无时间标注的疾病添加时序信息,包括1,657种罕见病(Orphanet编码)。我们还推出ChronoTQA基准,包含3,341个问题,涵盖八类任务(六类时序+两类静态对照),另有12题附加探测题。前沿大模型在从静态到时序问题上性能下降约30分,而使用ChronoMedKG检索可挽回47-65%的长尾失败,优于HPOA-RAG的17-29%。因此,ChronoMedKG为检索增强型临床系统提供了此前缺失的时间维度。

原文摘要 · Abstract (English)

Biomedical knowledge graphs (KGs) treat disease associations as static facts, but temporal information is crucial for clinical reasoning, e.g., a symptom diagnostic of one disease at age 3 may imply a different disease at age 13. Existing KGs such as PrimeKG, Hetionet, and iKraph do not encode when a finding becomes clinically relevant over the course of a disease. This limits their usefulness for longitudinal clinical reasoning and retrieval augmentation. We introduce ChronoMedKG, a temporal biomedical knowledge graph that contains 460,497 evidence-linked triples (filtered from 13M raw extractions) covering 13,431 diseases. Each association is tied to temporal components like onset window or progression stage, which are backed by PMID-traceable evidence and a multi-signal credibility score. The graph is constructed through a disease-autonomous multi-agent pipeline in which multiple frontier LLMs independently extract knowledge from PubMed and PMC literature. Only those relations are kept that are supported by multi-model consensus, survive credibility filtering, as well as ontology alignment. ChronoMedKG scored 92.7% agreement against Orphadata and adds temporal grounding for 6,250 diseases absent from HPOA, Orphadata, and Phenopackets, including 1,657 Orphanet-coded rare diseases. We further introduce ChronoTQA, a benchmark of 3,341 questions across eight task types (six temporal plus two static controls), with a 12-question supplementary probe. Frontier LLMs lose roughly 30 points moving from static to temporal questions; ChronoMedKG retrieval rescues 47-65% of their long-tail failures, against 17-29% for HPOA-RAG. As such, ChronoMedKG provides a crucial temporal axis for retrieval-augmented clinical systems that was previously absent.

知识图谱临床推理时序建模医学AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。