构建3D CT自监督模型,提升医学影像迁移能力
CoralBay: A Self-Supervised CT Foundation Model

- 用3D Swin结构+多尺度特征自蒸馏,实现高效自监督训练
- 在多个解剖部位上表现稳定,下游任务性能显著提升
- 开源3D放射学基准,推动体积化表示学习评估标准化
自监督学习已在2D自然图像上实现大规模预训练,生成可跨任务迁移的通用视觉表征。然而,许多医学影像模态(如CT扫描)本质上是三维的,其结构与语义与自然图像存在根本差异。体数据捕捉空间连续性、器官解剖结构及基于强度的组织特性(如亨氏单位),而2D预训练难以充分建模这些特征。为此,我们提出CoralBay,一种基于DINO的自蒸馏框架,采用分层3D Swin骨干网络,并对拼接的多尺度特征应用自蒸馏,实现了数据高效的自监督学习,获得富含全局语义与细粒度局部结构信息的空间表征。结果表明,CoralBay在多种下游放射学任务中表现出良好的迁移性能,且在不同解剖目标间保持一致表现。此外,我们通过开源eva框架,引入公开可复现的3D放射学排行榜,整合多个数据集,建立标准化基准以评估体数据表征学习方法。
原文摘要 · Abstract (English)
Self-supervised learning has enabled large-scale pre-training on 2D natural images, producing general-purpose visual representations that transfer effectively across tasks. However, many medical imaging modalities, such as CT scans, are inherently three-dimensional and differ fundamentally from natural images in both structure and semantics. Volumetric modalities capture spatial continuity, organ anatomy, and intensity-based tissue properties (e.g., Hounsfield Units), which are not adequately modeled by 2D pre-training. To bridge this gap, we introduce CoralBay, a self-distillation framework that extends DINO by using a hierarchical 3D Swin backbone and applying self-distillation to concatenated multi-scale features, enabling data-efficient self-supervised learning of rich spatial representations that encode both global semantics and fine-grained local structure. As a result, CoralBay transfers effectively to a wide range of downstream radiological tasks, demonstrating strong and consistent performance across diverse anatomical targets. In addition, we contribute to the open-source \eva framework by introducing a public, reproducible 3D radiology leaderboard that unifies multiple datasets and establishes a standardized benchmark for evaluating volumetric representation learning methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。