Medverse统一处理3D医学影像分割、变换与增强,无需重训练即可全分辨率输出。
Medverse: A Universal Model for Full-Resolution 3D Medical Image Segmentation, Transformation and Enhancement
- 采用分层自回归框架,从粗到细逐步优化,实现多尺度解剖理解。
- 在22个数据集上验证,跨中心、器官、物种和模态表现显著优于现有方法。
- 适合需要通用、高保真3D医学图像处理的研究者与临床应用开发人员。
上下文学习(ICL)为通用医学图像分析提供了新范式,使模型无需重新训练即可执行多种图像处理任务。然而,当前医学影像的ICL模型仍存在两大局限:难以同时实现高保真预测与全局解剖理解,且缺乏覆盖分割、增强等多任务及多解剖区域的统一模型。为此,我们提出 extbf{Medverse},一个面向3D医学影像的通用ICL模型,基于22个数据集训练,涵盖多种任务、器官、成像模态与临床中心。Medverse采用下一代自回归上下文学习框架,逐级细化预测,生成一致的全分辨率体数据输出,并具备多尺度解剖感知能力。我们还设计了块状交叉注意力模块,在保持空间稀疏性以控制计算开销的同时,促进上下文与目标输入间的长程交互。在涵盖未见临床中心、器官、物种和成像模态的多个测试集上,结果表明Medverse显著优于现有ICL基线,确立了新的上下文学习范式。代码与模型权重将公开提供,详见https://github.com/jiesihu/Medverse。
原文摘要 · Abstract (English)
In-context learning (ICL) offers a promising paradigm for universal medical image analysis, enabling models to perform diverse image processing tasks without retraining. However, current ICL models for medical imaging remain limited in two critical aspects: they cannot simultaneously achieve high-fidelity predictions and global anatomical understanding, and there is no unified model trained across diverse medical imaging tasks (e.g., segmentation and enhancement) and anatomical regions. As a result, the full potential of ICL in medical imaging remains underexplored. Thus, we present \textbf{Medverse}, a universal ICL model for 3D medical imaging, trained on 22 datasets covering diverse tasks in universal image segmentation, transformation, and enhancement across multiple organs, imaging modalities, and clinical centers. Medverse employs a next-scale autoregressive in-context learning framework that progressively refines predictions from coarse to fine, generating consistent, full-resolution volumetric outputs and enabling multi-scale anatomical awareness. We further propose a blockwise cross-attention module that facilitates long-range interactions between context and target inputs while preserving computational efficiency through spatial sparsity. Medverse is extensively evaluated on a broad collection of held-out datasets covering previously unseen clinical centers, organs, species, and imaging modalities. Results demonstrate that Medverse substantially outperforms existing ICL baselines and establishes a novel paradigm for in-context learning. Code and model weights will be made publicly available. Our model are publicly available at https://github.com/jiesihu/Medverse.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。