用自监督学习提升3D CT图像的空间结构理解能力
Rad-JEPA 3D: Radiology Joint-Embedding Predictive Model for 3D Computed Tomography

- 通过掩码视图预测完整扫描的潜在特征,学习三维CT表征
- 仅40亿参数即在器官识别与空间推理任务中达到顶尖水平
- 适合需要高效精准医疗影像分析的研究者与临床应用
自监督预训练在3D医学图像分析中至关重要,因未标注的CT体积数据丰富而专家标注稀缺。然而现有体积分解器常无法保持下游推理依赖的粗粒度空间与几何结构,限制了其在器官解耦、异常检测和空间理解方面的表现。我们提出Rad-JEPA 3D,一种联合嵌入预测框架,通过从掩码视图预测完整扫描的潜在特征来学习体积分解表示。核心为混合H-Mamba编码器,融合基于Mamba状态空间分支(建模跨切片连续性)与分组查询注意力分支(捕捉跨平面空间上下文),并通过轻量级逐标记路由模块结合。为进一步提升中间表示质量,提出隐状态正交正则化(HSOR),对齐学生-教师隐状态并减少编码器内特征冗余。该逐层正则化生成更一致且判别性强的体积分解表示,在约12万张CT扫描上预训练后,仅4.0B参数即在闭合式VQA任务上取得与顶尖模型相当的表现,并在Spatial-Med基准上获得最佳平均空间推理得分。消融实验表明,混合模块与HSOR贡献互补增益,所诱导的空间结构可替代语言模型规模以实现体积推理性能。
原文摘要 · Abstract (English)
Self-supervised pretraining is central to 3D medical image analysis, where unlabeled CT volumes are abundant but expert annotations are scarce. Yet existing volumetric encoders often fail to preserve the coarse spatial and geometric structure that downstream reasoning depends on, limiting their performance on organ disentanglement, abnormality detection, and spatial understanding when paired with language models. We introduce Rad-JEPA 3D, a joint-embedding predictive framework that learns volumetric CT representations by predicting the latent features of a complete scan from a masked view. At its core is a hybrid H-Mamba encoder that fuses a Mamba state-space branch, which models inter-slice continuity through sequential scanning, with a grouped-query attention branch, which captures cross-plane spatial context, combined through a lightweight per-token router. To improve the quality of intermediate representations, we further propose Hidden States Orthogonal Regularization (HSOR), which aligns student-teacher hidden states and reduces feature redundancy throughout the encoder. This layer-wise regularization produces more consistent and discriminative volumetric representations, leading to improved performance on organ recognition and spatial reasoning tasks. Pretrained on approximately 120,000 CT scans, Rad-JEPA 3D attains state-of-the-art results despite its compact size: with only 4.0B total parameters, it achieves competitive results with state-of-the-art on closed-ended VQA and the best average spatial-reasoning score on the Spatial-Med benchmark. Ablation studies confirm that the hybrid block and HSOR contribute complementary gains, and that the induced spatial structure can substitute for raw language-model scale on volumetric reasoning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。