arXiv:2605.09151cs.CV2026-05

统一处理2D与3D医学影像,提升诊断效率与模型泛化能力

MultiMedVision: Multi-Modal Medical Vision Framework

论文配图:MultiMedVision: Multi-Modal Medical Vision Framework
图 1 · 摘自论文原文
  • 基于稀疏视觉Transformer,原生支持混合模态批量处理
  • 仅用5倍少数据即达到2D/3D任务顶尖性能(最高AUROC 0.85)
  • 适合多模态医学影像分析、跨维度特征学习的研究者使用

多模态医学影像有助于全面诊断,但现有基础模型对2D(如X光)和3D(如CT)数据采用独立的、维度特定的架构。我们提出MultiMedVision,一种基于稀疏视觉Transformer的统一框架,用于联合2D/3D表示学习。该模型采用3D旋转位置嵌入与可变长度序列打包技术,在共享潜在空间中原生处理混合模态批次,无需模态专用适配器或将3D体积视为2D切片序列。在胸部X光(MIMIC-CXR)和CT扫描(CT-RATE)上通过自监督目标训练,使用单一共享编码器且数据量仅为之前的1/5,模型在2D基准(MIMIC宏AUROC 0.82,CheXpert 0.84)和3D任务(CT-RATE 0.85)上均表现优异。对学习表征的分析显示,存在共存的模态特异与共享特征子空间,证明了统一跨维度表示学习的可行性,且不牺牲模态特异性性能。

原文摘要 · Abstract (English)

Multi-modal medical imaging enables comprehensive diagnostics, yet current foundation models process 2D (e.g. X-ray) and 3D (e.g. CT) data with separate, dimensionality-specific architectures. We present MultiMedVision, a unified framework for joint 2D/3D representation learning built on a Sparse Vision Transformer. Our model uses 3D Rotary Positional Embeddings and variable-length sequence packing to process mixed-modality batches natively within a shared latent space, without modality-specific adapters or treating 3D volumes as 2D slice sequences. Trained with a self-supervised objective on chest X-rays (MIMIC-CXR) and CT scans (CT-RATE), and using a single shared encoder with 5x less data, MultiMedVision achieves competitive performance on both 2D benchmarks (Macro AUROC 0.82 on MIMIC, 0.84 on CheXpert) and 3D tasks (0.85 on CT-RATE). Analysis of the learned representations reveals coexisting modality-specific and shared feature subspaces, demonstrating that unified cross-dimensional representation learning is feasible without sacrificing modality-specific performance.

多模态医学影像视觉Transformer跨维度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。