用大模型补全医学知识图谱中的治疗关系,发现存在严重不靠谱风险。
Can LLMs Support Medical Knowledge Imputation? An Evaluation-Based Perspective
- 用大模型自动补全疾病与治疗的缺失关联
- 结果与临床指南不符,存在误诊隐患
- 适合关注医疗AI安全的研究者参考
医学知识图谱在临床决策支持和生物医学研究中至关重要,但常因知识空白和编码系统结构性缺陷导致不完整,尤其在治疗映射方面,ICD、Mondo 和 ATC 等编码系统覆盖不全,造成疾病与潜在治疗间关联缺失或不一致。本研究探索使用大语言模型(LLMs)进行治疗关系补全。尽管大模型具备知识增强潜力,其在医学知识补全中的应用仍存在重大风险,包括事实错误、虚构关联以及模型间与模型内结果不稳定。我们通过基准对比系统评估了大模型驱动的治疗映射可靠性,发现其结果与既定临床指南存在显著不一致,可能危及患者安全。本研究为研究人员和实践者提供警示,强调在利用大模型增强医学知识图谱治疗映射时,必须进行严格评估并采用混合方法。
原文摘要 · Abstract (English)
Medical knowledge graphs (KGs) are essential for clinical decision support and biomedical research, yet they often exhibit incompleteness due to knowledge gaps and structural limitations in medical coding systems. This issue is particularly evident in treatment mapping, where coding systems such as ICD, Mondo, and ATC lack comprehensive coverage, resulting in missing or inconsistent associations between diseases and their potential treatments. To address this issue, we have explored the use of Large Language Models (LLMs) for imputing missing treatment relationships. Although LLMs offer promising capabilities in knowledge augmentation, their application in medical knowledge imputation presents significant risks, including factual inaccuracies, hallucinated associations, and instability between and within LLMs. In this study, we systematically evaluate LLM-driven treatment mapping, assessing its reliability through benchmark comparisons. Our findings highlight critical limitations, including inconsistencies with established clinical guidelines and potential risks to patient safety. This study serves as a cautionary guide for researchers and practitioners, underscoring the importance of critical evaluation and hybrid approaches when leveraging LLMs to enhance treatment mappings on medical knowledge graphs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。