arXiv:2604.16980cs.LGcs.AI2026-04被引 1

十款多模态大模型在南非真实病房数据中表现相近,均优于人工诊断。

Evaluating Multimodal LLMs for Inpatient Diagnosis: Real-World Performance, Safety, and Cost Across Ten Frontier Models

论文配图:Evaluating Multimodal LLMs for Inpatient Diagnosis: Real-World Performance, Safety, and Cost Across Ten Frontier Models
图 1 · 摘自论文原文
  • 用真实病患的影像、检验和病历数据评估十款模型零样本诊断能力
  • 所有模型平均诊断准确率和安全得分均显著高于临床医生常规诊疗
  • 低成本模型表现与顶级模型相当,适合资源有限地区部署

背景:大语言模型(LLMs)被提出用于辅助诊断,但很少使用真实世界多模态住院数据进行评估,尤其在低收入和中等收入国家(LMIC)公立医院。方法:我们开展了VALID研究,回顾性评估了南非一家三级公立医院的539例多模态住院病例,输入包括放射影像(CT、MRI、CXR)及报告、实验室结果、临床记录和生命体征。专家小组对300例病例(平衡与分歧子集)进行裁定,确立真实诊断、鉴别诊断和推理依据。十款多模态LLM生成零样本输出,由校准的三模型LLM评审团对所有输出及常规病房诊断进行评分,涵盖诊断准确性、鉴别诊断质量、推理能力和患者安全(>10,000次评估)。主要结局为综合评分($S_3$、$S_4$)和胜率。结果:(i) 尽管成本差异大,各模型性能高度集中(<15%波动),低成本模型表现与顶级模型相当;(ii) 所有模型平均诊断和安全得分均显著优于常规病房诊断;(iii) GPT-5.1表现最佳,其次为Gemini系列;(iv) 加入放射报告使性能提升6%;(v) 诊断与推理得分高度相关($ρ=0.85$);(vi) 输出率因输入限制而异(65%-100%)。结果在各子集和评估设计中均稳健。结论:在真实世界LMIC数据集上,多模态LLM性能相似,且普遍优于常规医疗实践,性价比、鲁棒性和部署约束可能比微小性能差异更具实际意义。

原文摘要 · Abstract (English)

Background: Large language models (LLMs) are increasingly proposed for diagnostic support, but few evaluations use real-world multimodal inpatient data, particularly in low and middle-income country (LMIC) public hospitals. Methods: We conducted VALID, a retrospective evaluation of 539 multimodal inpatient cases from a tertiary public hospital in South Africa. Inputs included radiology imaging (CT, MRI, CXR) and reports, laboratory results, clinical notes, and vital signs. Expert panels adjudicated 300 cases (balanced and discordant subsets) to establish ground truth diagnoses, differentials, and reasoning. Ten multimodal LLMs generated zero-shot outputs. A calibrated three-model LLM Jury scored all outputs and routine ward diagnoses across diagnostic accuracy, differential quality, reasoning, and patient safety (>10,000 evaluations). Primary outcomes were composite scores ($S_3$, $S_4$) and win rates. Results: (i) LLM performance was tightly clustered (<15% variation) despite large cost differences; low-cost models performed comparably to top models. (ii) All LLMs significantly outperformed routine ward diagnoses on average diagnostic and safety scores. (iii) Top performance was achieved by GPT-5.1, followed by Gemini models. (vi) Adding radiology reports improved performance by 6%. (v) Diagnostic and reasoning scores were highly correlated ($ρ= 0.85$). (vi) Output rates varied (65-100%) due to input constraints. Results were robust across subsets and evaluation design. Conclusions: Across a real-world LMIC dataset, multimodal LLMs showed similar diagnostic performance despite large cost differences and outperformed routine care on average safety metrics. Affordability, robustness, and deployment constraints may outweigh marginal performance differences in LMIC settings.

多模态大模型医疗诊断低成本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。