用2D大模型提升3D医学影像理解,无需人工标注
Enhancing 3D Medical Image Understanding with Pretraining Aided by 2D Multimodal Large Language Models
- 将3D图像编码器与2D多模态大模型通过切片感知模块结合
- 在CT/MRI数据上实现分割与分类的顶尖性能,优于现有自监督方法
- 适合希望提升3D医学图像模型性能的研究者和开发者
3D医学图像理解对医疗领域至关重要,但现有的3D卷积与基于Transformer的自监督学习方法往往缺乏深层语义理解。近期多模态大语言模型(MLLMs)通过文本描述增强了图像理解能力。为利用2D MLLMs提升3D医学图像理解,我们提出Med3DInsight,一种新颖的预训练框架,通过专门设计的平面切片感知变换器模块将3D图像编码器与2D MLLMs融合。此外,模型采用基于部分最优传输的对齐策略,对LLM生成内容中潜在噪声具有更强容忍度。Med3DInsight开创了一种无需人工标注的可扩展多模态3D医学表征学习新范式。大量实验表明,在分割与分类两个下游任务上,其在多种公开数据集(涵盖CT与MRI模态)上均达到当前最佳表现,显著优于现有自监督方法。Med3DInsight可无缝集成至现有3D医学图像理解网络,有望提升其性能。代码、生成数据集及预训练模型将开源于https://github.com/Qybc/Med3DInsight。
原文摘要 · Abstract (English)
Understanding 3D medical image volumes is critical in the medical field, yet existing 3D medical convolution and transformer-based self-supervised learning (SSL) methods often lack deep semantic comprehension. Recent advancements in multimodal large language models (MLLMs) provide a promising approach to enhance image understanding through text descriptions. To leverage these 2D MLLMs for improved 3D medical image understanding, we propose Med3DInsight, a novel pretraining framework that integrates 3D image encoders with 2D MLLMs via a specially designed plane-slice-aware transformer module. Additionally, our model employs a partial optimal transport based alignment, demonstrating greater tolerance to noise introduced by potential noises in LLM-generated content. Med3DInsight introduces a new paradigm for scalable multimodal 3D medical representation learning without requiring human annotations. Extensive experiments demonstrate our state-of-the-art performance on two downstream tasks, i.e., segmentation and classification, across various public datasets with CT and MRI modalities, outperforming current SSL methods. Med3DInsight can be seamlessly integrated into existing 3D medical image understanding networks, potentially enhancing their performance. Our source code, generated datasets, and pre-trained models will be available at https://github.com/Qybc/Med3DInsight.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。