首个面向多参数3D MRI的视觉语言模型,实现跨模态精准解读。
Mr3D-VL: A generalist vision language foundation model for Multiparametric 3D Magnetic Resonance Imaging

- 构建3D体积感知的共享编码器与4D旋转位置嵌入,融合多模态空间信息。
- 报告生成BERTScore达0.856,问答准确率71.3%,多选正确率达91.2%。
- 专为脑肿瘤诊断设计,适合医学影像与多模态大模型研究者使用。
多参数磁共振成像(mpMRI)是脑肿瘤诊疗的核心,但现有AI模型缺乏自然语言交互与可解释性,难以整合空间信息并进行跨模态推理。主要挑战包括各模态间物理意义差异大、扫描时间间隔导致的空间错位,以及胶质瘤分级等任务中复杂的多特征解析需求。尽管视觉语言模型(VLMs)在跨模态理解上展现潜力,但现有方法主要集中于2D图像建模,忽视了对3D体数据的直接感知。虽已有3D VLM用于3D CT的报告生成与特征对齐,但mpMRI应用需在多种成像模态间协同推理,现有方案尚无法满足。为此,我们提出Mr3D-VL——一个针对多参数3D MRI的专用视觉语言基础模型。该模型含40亿参数,采用无监督预训练的共享3D编码器与4D旋转位置嵌入,实现双模态-空间融合。其跨模态投影层采用多分辨率特征嵌入策略,提升多尺度特征感知能力。实验表明,在文本生成任务中显著优于现有4B/7B/30B领域专用与通用模型:报告生成的BERTScore达0.856,问答准确率为0.713,多选准确率达0.912。
原文摘要 · Abstract (English)
Multi-parametric magnetic resonance imaging (mpMRI) is a cornerstone for brain tumor diagnosis and treatment, yet current AI models face critical limitations: their lack of natural language interaction and interpretability impedes spatial information integration and cross-modal reasoning required clinically. Key challenges arise from significant physical meaning differences across modalities, spatial misalignment due to scan intervals, and the need for complex multi-feature interpretation in tasks like glioma grading. While visual-language models (VLMs) show promise in cross-modal understanding, existing methods focus mainly on 2D image modeling, neglecting direct perception of 3D volumetric space. Although 3D VLMs have been proposed for report generation and feature alignment in 3D CT imaging, mpMRI applications demand collaborative inference across multiple imaging modalities-a requirement unmet by current solutions. To address this, we introduce Mr3D-VL, a dedicated visual-language foundation model for multi-parametric 3D MRI. With 4 billion parameters, it employs an unsupervised pre-trained shared 3D encoder and 4D rotational positional embedding for dual modality-spatial integration. Its cross-modal projection layer uses a multi-resolution feature implantation strategy to enhance feature perception across resolutions. Experimental results show significant improvements over existing 4B/7B/30B domain-specific and general-purpose models in text generation tasks, achieving a BERTScore of 0.856 for report generation, with question-answering accuracy at 0.713 and multiple-choice accuracy at 0.912.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。