对比七种方法发现疾病关联,发现大模型未必能发现新关系。
Revealing Interconnections between Diseases: from Statistical Methods to Large Language Models
- 用真实病历和医学文本数据,测试七种算法找疾病关联。
- 大模型生成的疾病关联种类最少,可能难发现新联系。
- 结果可作医疗知识图谱基础,适合研究者与AI医生参考。
通过人工分析大规模临床数据识别疾病关联耗时费力,且易受主观判断影响。尽管机器学习有潜力,仍面临三大挑战:(1) 从海量方法中选择最优方案;(2) 判断真实临床数据(如MIMIC-IV电子健康记录)或结构化疾病描述哪个更可靠;(3) 缺乏“真实标准”,部分疾病关联在医学上尚未被探索。大语言模型虽具广泛适用性,但常缺乏专业医学知识。为此,我们系统评估了七种方法,基于两类数据:(i) MIMIC-IV EHR中的ICD-10编码序列;(ii) 全量ICD-10编码及其文本描述。框架包括:(i) 统计共现分析与基于真实数据的掩码语言建模(MLM);(ii) 领域专用BERT(Med-BERT、BioClinicalBERT);(iii) 通用BERT与文档检索;(iv) 四个LLMs(Mistral、DeepSeek、Qwen、YandexGPT)。基于图的比较显示,基于LLM的方法产生的疾病关联矩阵多样性最低,相较于其他方法(包括文本与领域专用方法),提示其发现新关联的能力有限。由于尚无权威的疾病间关联数据库,本研究结果构成一个有价值的医学疾病本体资源,可为未来临床研究及医疗人工智能应用提供基础支持。
原文摘要 · Abstract (English)
Identifying disease interconnections through manual analysis of large-scale clinical data is labor-intensive, subjective, and prone to expert disagreement. While machine learning (ML) shows promise, three critical challenges remain: (1) selecting optimal methods from the vast ML landscape, (2) determining whether real-world clinical data (e.g., electronic health records, EHRs) or structured disease descriptions yield more reliable insights, (3) the lack of "ground truth," as some disease interconnections remain unexplored in medicine. Large language models (LLMs) demonstrate broad utility, yet they often lack specialized medical knowledge. To address these gaps, we conduct a systematic evaluation of seven approaches for uncovering disease relationships based on two data sources: (i) sequences of ICD-10 codes from MIMIC-IV EHRs and (ii) the full set of ICD-10 codes, both with and without textual descriptions. Our framework integrates the following: (i) a statistical co-occurrence analysis and a masked language modeling (MLM) approach using real clinical data; (ii) domain-specific BERT variants (Med-BERT and BioClinicalBERT); (iii) a general-purpose BERT and document retrieval; and (iv) four LLMs (Mistral, DeepSeek, Qwen, and YandexGPT). Our graph-based comparison of the obtained interconnection matrices shows that the LLM-based approach produces interconnections with the lowest diversity of ICD code connections to different diseases compared to other methods, including text-based and domain-based approaches. This suggests an important implication: LLMs have limited potential for discovering new interconnections. In the absence of ground truth databases for medical interconnections between ICD codes, our results constitute a valuable medical disease ontology that can serve as a foundational resource for future clinical research and artificial intelligence applications in healthcare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。