arXiv:2604.10233cs.CVcs.AI2026-04

将2D大模型适配3D医学影像,提升报告生成与问答性能

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis

  • 用2D预训练模型迁移支持3D医学体数据输入
  • 引入文本引导的分层MoE框架,实现任务定制特征提取
  • 两阶段训练兼顾通用与特定任务特征,适合临床辅助诊断场景

3D医学图像分析在疾病诊断与治疗中至关重要。近年来,多模态大语言模型(MLLM)展现出强大的感知能力、跨模态对齐能力和良好的泛化性,有望显著提升医学报告生成(MRG)和医学视觉问答(MVQA)性能。然而,由于3D医学图像数据稀缺,现有3D医学MLLM存在视觉编码器预训练不足、难以为不同任务提取定制化特征的问题。本文提出首先将经过2D自然图像充分预训练的2D MLLM迁移到支持3D医学体数据输入,并复用全部预训练参数。为进一步实现任务定制化特征提取,设计了文本引导的分层混合专家(TGH-MoE)框架,在文本提示指导下区分不同任务。同时提出两阶段训练策略,学习共享与特定任务的图像特征。实验证明,该方法在MRG与MVQA任务上均优于现有3D医学MLLM。代码将在论文录用后公开。

原文摘要 · Abstract (English)

3D medical image analysis is of great importance in disease diagnosis and treatment. Recently, multimodal large language models (MLLMs) have exhibited robust perceptual capacity, strong cross-modal alignment, and promising generalizability. Therefore, they have great potential to improve the performance of medical report generation (MRG) and medical visual question answering (MVQA), which serve as two important tasks in clinical scenarios. However, due to the scarcity of 3D medical images, existing 3D medical MLLMs suffer from insufficiently pretrained vision encoder and inability to extract customized image features for different kinds of tasks. In this paper, we propose to first transfer a 2D MLLM, which is well trained with 2D natural images, to support 3D medical volumetric inputs while reusing all of its pre-trained parameters. To enable the vision encoder to extract tailored image features for various tasks, we then design a Text-Guided Hierarchical MoE (TGH-MoE) framework, which can distinguish tasks under the guidance of the text prompt. Furthermore, we propose a two-stage training strategy to learn both task-shared and task-specific image features. As demonstrated empirically, our method outperforms existing 3D medical MLLMs in both MRG and MVQA tasks. Our code will be released once this paper is accepted.

3D医学影像多模态大模型报告生成视觉问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。