用2D基础模型轻松实现多任务3D医学图像分类,效果顶尖且省资源。
Revisiting 2D Foundation Models for Scalable 3D Medical Image Classification
- 基于冻结的2D模型加轻量插件,每任务仅增100万参数。
- 在12个不同疾病/器官/模态任务上达顶尖性能,部分超越专用3D模型。
- 支持多视角输入、像素级监督和可解释热力图,适合临床部署。
3D医学图像分类对现代临床工作至关重要。医学基础模型(FMs)虽有望扩展至新任务,但现有研究存在三大缺陷:数据范式偏差、适配不充分、任务覆盖不足。本文提出AnyMC3D,一种从2D基础模型演化而来的可扩展3D分类器。该方法通过在单一冻结主干网络上添加轻量级插件(每任务约100万参数),高效适应新任务。框架支持多视角输入、辅助像素级监督及可解释热力图生成。我们构建了涵盖12项任务的综合性基准,涵盖多种病灶、解剖结构与成像模态,并系统评估主流3D分类技术。分析揭示:(1)有效适配是释放基础模型潜力的关键;(2)若适配得当,通用基础模型可媲美医学专用模型;(3)基于2D的方法在3D分类中优于专用3D架构。首次证明单一可扩展框架可在多样应用中达到最先进水平(包括在VLM3D挑战赛中获得第一名),无需为每项任务单独建模。
原文摘要 · Abstract (English)
3D medical image classification is essential for modern clinical workflows. Medical foundation models (FMs) have emerged as a promising approach for scaling to new tasks, yet current research suffers from three critical pitfalls: data-regime bias, suboptimal adaptation, and insufficient task coverage. In this paper, we address these pitfalls and introduce AnyMC3D, a scalable 3D classifier adapted from 2D FMs. Our method scales efficiently to new tasks by adding only lightweight plugins (about 1M parameters per task) on top of a single frozen backbone. This versatile framework also supports multi-view inputs, auxiliary pixel-level supervision, and interpretable heatmap generation. We establish a comprehensive benchmark of 12 tasks covering diverse pathologies, anatomies, and modalities, and systematically analyze state-of-the-art 3D classification techniques. Our analysis reveals key insights: (1) effective adaptation is essential to unlock FM potential, (2) general-purpose FMs can match medical-specific FMs if properly adapted, and (3) 2D-based methods surpass 3D architectures for 3D classification. For the first time, we demonstrate the feasibility of achieving state-of-the-art performance across diverse applications using a single scalable framework (including 1st place in the VLM3D challenge), eliminating the need for separate task-specific models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。