构建真实罕见病案例数据集,评估大模型临床诊断能力。
CUPCase: Clinically Uncommon Patient Cases and Diagnoses Dataset
- 基于3562个真实病例构建罕见病诊断数据集
- GPT-4o在多选和开放式任务中分别达87.9%与0.764的性能
- 仅用20%信息仍保持87%以上准确率,适合临床早期辅助
医疗基准数据集对发展用于医学知识提取、诊断和摘要的大语言模型(LLMs)至关重要。然而,现有基准多来自医学生考试题或文献描述,缺乏真实世界中偏离经典教科书范式的复杂病例,如罕见病、常见病非典型表现及意外治疗反应。本文基于BMC收录的3562个真实病例报告,构建了临床罕见患者病例与诊断数据集(CUPCase),包含开放式文本诊断和带干扰项的多选题。利用该数据集,评估了包括通用和临床专用模型在内的多种先进大模型的诊断能力,并测试了部分信息下的表现。结果显示,通用模型GPT-4o在多选任务(平均准确率87.9%)和开放式任务(BERTScore F1 0.764)中均优于MedLM-Large等专业医学模型;且在仅提供病例前20%文本时,仍能保持多选任务87%和自由文本任务88%的性能,展现出大模型在真实临床早期诊断中的潜力。CUPCase为临床决策支持的开放、可复现评估提供了新工具。
原文摘要 · Abstract (English)
Medical benchmark datasets significantly contribute to developing Large Language Models (LLMs) for medical knowledge extraction, diagnosis, summarization, and other uses. Yet, current benchmarks are mainly derived from exam questions given to medical students or cases described in the medical literature, lacking the complexity of real-world patient cases that deviate from classic textbook abstractions. These include rare diseases, uncommon presentations of common diseases, and unexpected treatment responses. Here, we construct Clinically Uncommon Patient Cases and Diagnosis Dataset (CUPCase) based on 3,562 real-world case reports from BMC, including diagnoses in open-ended textual format and as multiple-choice options with distractors. Using this dataset, we evaluate the ability of state-of-the-art LLMs, including both general-purpose and Clinical LLMs, to identify and correctly diagnose a patient case, and test models' performance when only partial information about cases is available. Our findings show that general-purpose GPT-4o attains the best performance in both the multiple-choice task (average accuracy of 87.9%) and the open-ended task (BERTScore F1 of 0.764), outperforming several LLMs with a focus on the medical domain such as Meditron-70B and MedLM-Large. Moreover, GPT-4o was able to maintain 87% and 88% of its performance with only the first 20% of tokens of the case presentation in multiple-choice and free text, respectively, highlighting the potential of LLMs to aid in early diagnosis in real-world cases. CUPCase expands our ability to evaluate LLMs for clinical decision support in an open and reproducible manner.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。