3D医学多模态模型实现报告生成与精准分割统一,支持语言/点/框提示。
MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation
- 构建统一架构融合3D影像与文本理解,支持多粒度空间推理。
- 在3D CT数据上训练,实现报告生成、VQA和语义/指代/交互式分割全任务领先。
- 支持语言、点、框多种提示,可实现可控的精确三维定位与跨模态推理。
近期医学视觉语言模型(VLMs)在图像级文本主导任务如报告生成和视觉问答(VQA)上取得显著进展。然而,在3D医学VLM中实现细粒度视觉定位与体积分层空间推理仍具挑战,尤其在单一通用框架内整合这些能力。为此,我们提出MedVL-SAM2,一种统一的3D医学多模态模型,可同时支持报告生成、VQA及多范式分割(包括语义、指代和交互分割)。该模型通过融合图像级推理与像素级感知的协同架构,结合基于SAM2的体积分割模块,实现精确的多粒度空间推理。模型采用多阶段训练:首先在大规模3D CT图像-文本对语料库上预训练,对齐体积分层视觉特征与放射科语言嵌入;随后在综合性3D CT分割数据集上联合优化语言理解与分割目标。这种联合训练使模型能通过语言、点或框提示灵活交互,从而统一高层视觉推理与空间精确定位。实验表明,该统一架构在报告生成、VQA及多项3D分割任务上均达到当前最优性能。深入分析显示,模型具备可靠的3D视觉定位能力、可控的交互分割效果与稳健的跨模态推理能力,证明在统一的3D医学VLM中,高层语义推理与精确三维定位可协同实现。
原文摘要 · Abstract (English)
Recent progress in medical vision-language models (VLMs) has achieved strong performance on image-level text-centric tasks such as report generation and visual question answering (VQA). However, achieving fine-grained visual grounding and volumetric spatial reasoning in 3D medical VLMs remains challenging, particularly when aiming to unify these capabilities within a single, generalizable framework. To address this challenge, we proposed MedVL-SAM2, a unified 3D medical multimodal model that concurrently supports report generation, VQA, and multi-paradigm segmentation, including semantic, referring, and interactive segmentation. MedVL-SAM2 integrates image-level reasoning and pixel-level perception through a cohesive architecture tailored for 3D medical imaging, and incorporates a SAM2-based volumetric segmentation module to enable precise multi-granular spatial reasoning. The model is trained in a multi-stage pipeline: it is first pre-trained on a large-scale corpus of 3D CT image-text pairs to align volumetric visual features with radiology-language embeddings. It is then jointly optimized with both language-understanding and segmentation objectives using a comprehensive 3D CT segmentation dataset. This joint training enables flexible interaction via language, point, or box prompts, thereby unifying high-level visual reasoning with spatially precise localization. Our unified architecture delivers state-of-the-art performance across report generation, VQA, and multiple 3D segmentation tasks. Extensive analyses further show that the model provides reliable 3D visual grounding, controllable interactive segmentation, and robust cross-modal reasoning, demonstrating that high-level semantic reasoning and precise 3D localization can be jointly achieved within a unified 3D medical VLM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。