探究医学影像多模态模型的组合泛化能力,发现其是跨任务泛化的关键。
Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging
- 用组合泛化框架分析多任务学习中的知识迁移机制。
- 在106个医学数据集上验证模型能理解未见过的图像组合。
- 适合关注医疗AI泛化能力与少样本学习的研究者。
医学影像为诊断提供关键视觉信息,多模态大语言模型(MLLMs)因其强大泛化能力被广泛用于分析;然而其泛化机制尚不明确。现有研究认为多任务训练优于单任务,因不同任务间可相互促进,但常忽视任务内部关联。为此,我们采用组合泛化(CG)作为分析框架,即模型通过重组已学元素理解新组合的能力。由于医学图像可由成像模态、解剖部位和任务类型精确定义,天然适合探索CG,因此我们构建了包含106个医学数据集的Med-MAT数据集以支持全面实验。实验结果表明,MLLMs能够利用组合泛化理解未见医学图像,并确认组合泛化是多任务训练中观察到泛化现象的主要驱动因素之一。进一步研究表明,组合泛化在数据量有限的场景下仍有效,并证实MLLMs可在分类与检测任务间实现组合泛化,凸显其更广泛的泛化潜力。Med-MAT开源地址:https://github.com/FreedomIntelligence/Med-MAT。
原文摘要 · Abstract (English)
Medical imaging provides essential visual insights for diagnosis, and multimodal large language models (MLLMs) are increasingly utilized for its analysis due to their strong generalization capabilities; however, the underlying factors driving this generalization remain unclear. Current research suggests that multi-task training outperforms single-task as different tasks can benefit each other, but they often overlook the internal relationships within these tasks. To analyze this phenomenon, we attempted to employ compositional generalization (CG), which refers to the models' ability to understand novel combinations by recombining learned elements, as a guiding framework. Since medical images can be precisely defined by Modality, Anatomical area, and Task, naturally providing an environment for exploring CG, we assembled 106 medical datasets to create Med-MAT for comprehensive experiments. The experiments confirmed that MLLMs can use CG to understand unseen medical images and identified CG as one of the main drivers of the generalization observed in multi-task training. Additionally, further studies demonstrated that CG effectively supports datasets with limited data and confirmed that MLLMs can achieve CG across classification and detection tasks, underscoring its broader generalization potential. Med-MAT is available at https://github.com/FreedomIntelligence/Med-MAT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。