用真实病历测试大模型诊罕见病,发现表现远未达标
MIMIC-RD: Can LLMs differentially diagnose rare diseases in real-world clinical settings?
- 从真实病历中提取实体并映射到罕见病数据库Orphanet构建新数据集
- 145例患者测试显示当前顶级大模型诊断准确率很低
- 研究揭示临床真实场景下大模型在罕见病诊断上的巨大短板
尽管罕见病影响美国每10人中有1人,其鉴别诊断仍具挑战。由于具备出色的回忆能力,大语言模型(LLMs)近期被用于鉴别诊断。现有评估方法存在两大局限:依赖理想化临床案例,无法反映真实临床复杂性;或使用ICD编码作为疾病标签,严重低估罕见病数量,因许多罕见病在综合数据库如Orphanet中无直接映射。为解决这些问题,我们构建了MIMIC-RD——一个通过将临床文本实体直接映射至Orphanet的罕见病鉴别诊断基准。方法包括初始的LLM挖掘流程及四位医学标注者验证以确认实体真实性。我们在包含145名患者的数据集上评估多种模型,发现当前最先进大模型在罕见病鉴别诊断任务中表现不佳,凸显现有能力与临床需求之间的显著差距。基于研究结果,我们提出了若干未来改进方向。
原文摘要 · Abstract (English)
Despite rare diseases affecting 1 in 10 Americans, their differential diagnosis remains challenging. Due to their impressive recall abilities, large language models (LLMs) have been recently explored for differential diagnosis. Existing approaches to evaluating LLM-based rare disease diagnosis suffer from two critical limitations: they rely on idealized clinical case studies that fail to capture real-world clinical complexity, or they use ICD codes as disease labels, which significantly undercounts rare diseases since many lack direct mappings to comprehensive rare disease databases like Orphanet. To address these limitations, we explore MIMIC-RD, a rare disease differential diagnosis benchmark constructed by directly mapping clinical text entities to Orphanet. Our methodology involved an initial LLM-based mining process followed by validation from four medical annotators to confirm identified entities were genuine rare diseases. We evaluated various models on our dataset of 145 patients and found that current state-of-the-art LLMs perform poorly on rare disease differential diagnosis, highlighting the substantial gap between existing capabilities and clinical needs. From our findings, we outline several future steps towards improving differential diagnosis of rare diseases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。