对比不同架构与训练策略,找出适合医学多模态任务的最优模型方案。
MedM-VL: What Makes a Good Medical LVLM?
- 基于LLaVA框架构建2D与3D医学多模态模型,系统测试多种配置
- 提出两个预训练模型:MedM-VL-2D和MedM-VL-CT-Chest,支持临床应用
- 开源代码与模型,推动医学视觉语言研究可复现与扩展
医学图像分析在现代医疗中至关重要。深度学习已将研究重点转向复杂医学多模态任务,如报告生成与视觉问答。传统任务专用模型难以应对这些挑战。大型视觉语言模型(LVLM)为此提供了新解决方案。本研究基于流行的LLaVA框架,系统探索了2D与3D医学LVLM的模型架构与训练策略。我们呈现了广泛的实证发现与实用指导。为支持可复现性与未来研究,我们发布模块化代码库MedM-VL,以及两个预训练模型:用于2D医学图像分析的MedM-VL-2D,以及面向3D CT应用的MedM-VL-CT-Chest。代码可在https://github.com/MSIIP/MedM-VL获取。
原文摘要 · Abstract (English)
Medical image analysis is essential in modern healthcare. Deep learning has redirected research focus toward complex medical multimodal tasks, including report generation and visual question answering. Traditional task-specific models often fall short in handling these challenges. Large vision-language models (LVLMs) offer new solutions for solving such tasks. In this study, we build on the popular LLaVA framework to systematically explore model architectures and training strategies for both 2D and 3D medical LVLMs. We present extensive empirical findings and practical guidance. To support reproducibility and future research, we release a modular codebase, MedM-VL, and two pre-trained models: MedM-VL-2D for 2D medical image analysis and MedM-VL-CT-Chest for 3D CT-based applications. The code is available at: https://github.com/MSIIP/MedM-VL
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。