构建多模态医学影像评估框架,发现大模型在多数病例中优于人类医生。
Comprehensive Evaluation of Multimodal AI Models in Medical Imaging Diagnosis: From Data Augmentation to Preference-Based Comparison
- 融合影像与临床数据,通过可控增强扩展至3000例
- Llama 3.2-90B在85.27%病例中超越人类诊断
- 适合医疗AI研发者和临床决策系统评估者
本研究提出一个用于多模态医学影像诊断的评估框架。我们构建了包含数据预处理、模型推理与偏好比较的全流程,将初始500个临床病例通过受控增强扩展至3000例。方法结合医学影像与临床观察生成诊断建议,并使用Claude 3.5 Sonnet进行独立评估,对比医师撰写诊断。结果表明模型表现各异:Llama 3.2-90B在85.27%的病例中优于人类诊断;而专用视觉模型如BLIP2和Llava分别在41.36%和46.77%的案例中表现出偏好。该框架揭示了大型多模态模型在特定任务中超越人类诊断的潜力。
原文摘要 · Abstract (English)
This study introduces an evaluation framework for multimodal models in medical imaging diagnostics. We developed a pipeline incorporating data preprocessing, model inference, and preference-based evaluation, expanding an initial set of 500 clinical cases to 3,000 through controlled augmentation. Our method combined medical images with clinical observations to generate assessments, using Claude 3.5 Sonnet for independent evaluation against physician-authored diagnoses. The results indicated varying performance across models, with Llama 3.2-90B outperforming human diagnoses in 85.27% of cases. In contrast, specialized vision models like BLIP2 and Llava showed preferences in 41.36% and 46.77% of cases, respectively. This framework highlights the potential of large multimodal models to outperform human diagnostics in certain tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。