测试了检验结果对大模型临床诊断建议的影响,发现加了化验单后准确率显著提升。
Evaluating the Impact of Lab Test Results on Large Language Models Generated Differential Diagnoses from Clinical Case Vignettes
- 用50个真实病例构建带化验数据的临床情景,对比模型有无检验信息时的诊断表现
- 加了化验数据后GPT-4的顶级诊断准确率达55%,前十项达60%,最宽松标准下80%
- 肝功能、代谢毒理和免疫检测结果多数被正确理解,适合医疗AI评估与临床辅助研究
鉴别诊断在医学中至关重要,有助于医护人员系统区分症状相似的疾病。本研究评估了实验室检查结果对大语言模型(LLMs)生成临床病例鉴别诊断的影响。从PubMed Central选取50个病例报告,构建包含患者人口统计、症状及实验室结果的临床情景。测试了GPT-4、GPT-3.5、Llama-2-70b、Claude-2和Mixtral-8x7B共五种模型,在有无实验室数据条件下分别生成前10、前5及前1个鉴别诊断。通过GPT-4、知识图谱和临床医生进行综合评估。GPT-4表现最佳,使用实验室数据时前1个诊断准确率为55%,前10个为60%,宽松标准下最高达80%。实验室结果显著提升诊断准确率,尤其在GPT-4和Mixtral上表现突出,但完全匹配率仍较低。肝功能、代谢/毒理学面板以及血清学/免疫学检查结果总体上被模型正确解读用于鉴别诊断。
原文摘要 · Abstract (English)
Differential diagnosis is crucial for medicine as it helps healthcare providers systematically distinguish between conditions that share similar symptoms. This study assesses the impact of lab test results on differential diagnoses (DDx) made by large language models (LLMs). Clinical vignettes from 50 case reports from PubMed Central were created incorporating patient demographics, symptoms, and lab results. Five LLMs GPT-4, GPT-3.5, Llama-2-70b, Claude-2, and Mixtral-8x7B were tested to generate Top 10, Top 5, and Top 1 DDx with and without lab data. A comprehensive evaluation involving GPT-4, a knowledge graph, and clinicians was conducted. GPT-4 performed best, achieving 55% accuracy for Top 1 diagnoses and 60% for Top 10 with lab data, with lenient accuracy up to 80%. Lab results significantly improved accuracy, with GPT-4 and Mixtral excelling, though exact match rates were low. Lab tests, including liver function, metabolic/toxicology panels, and serology/immune tests, were generally interpreted correctly by LLMs for differential diagnosis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。