arXiv:2604.01612cs.CVcs.AI2026-04

用超块设计提升3DCT图像自监督学习效率与精度

NEMESIS: Noise-suppressed Efficient MAE with Enhanced Superpatch Integration Strategy

  • 基于128×128×128超块构建轻量级掩码自编码器,降低内存消耗
  • 在BTCV数据集上达0.9633平均AUROC,10%标签下仍保持0.9075
  • 适用于标注稀缺的3D医学影像任务,尤其适合资源受限场景

体积CT成像对临床诊断至关重要,但三维体数据标注成本高昂,推动了无标签数据的自监督学习(SSL)发展。然而,将SSL应用于3D CT面临全体积变换器内存开销大、以及CT数据各向异性结构难以被传统掩码策略捕捉的挑战。本文提出NEMESIS,一种在局部128×128×128超块上运行的掩码自编码器框架,实现内存高效训练并保留解剖细节。NEMESIS引入三个核心组件:(i) 噪声增强重建作为预训练任务,(ii) 双重掩码的掩码解剖变换器块(MATB),通过并行平面与轴向标记移除建模空间结构,(iii) 跨尺度上下文聚合的NEMESIS Tokens(NT)。在BTCV多器官分类基准上,使用冻结主干+线性分类器的NEMESIS达到0.9633平均AUROC,优于全微调的SuPreM(0.9493)和VoCo(0.9387)。在仅10%标注数据的低标签环境下,仍保持0.9075的AUROC,展现强标签效率。此外,超块设计使前向计算量降至31.0 GFLOPs/次,相较全体积基线的985.8 GFLOPs显著降低,为3D医学影像提供可扩展、鲁棒的基础架构。

原文摘要 · Abstract (English)

Volumetric CT imaging is essential for clinical diagnosis, yet annotating 3D volumes is expensive and time-consuming, motivating self-supervised learning (SSL) from unlabeled data. However, applying SSL to 3D CT remains challenging due to the high memory cost of full-volume transformers and the anisotropic spatial structure of CT data, which is not well captured by conventional masking strategies. We propose NEMESIS, a masked autoencoder (MAE) framework that operates on local 128x128x128 superpatches, enabling memory-efficient training while preserving anatomical detail. NEMESIS introduces three key components: (i) noise-enhanced reconstruction as a pretext task, (ii) Masked Anatomical Transformer Blocks (MATB) that perform dual-masking through parallel plane-wise and axis-wise token removal, and (iii) NEMESIS Tokens (NT) for cross-scale context aggregation. On the BTCV multi-organ classification benchmark, NEMESIS with a frozen backbone and a linear classifier achieves a mean AUROC of 0.9633, surpassing fully fine-tuned SuPreM (0.9493) and VoCo (0.9387). Under a low-label regime with only 10% of available annotations, it retains an AUROC of 0.9075, demonstrating strong label efficiency. Furthermore, the superpatch-based design reduces computational cost to 31.0 GFLOPs per forward pass, compared to 985.8 GFLOPs for the full-volume baseline, providing a scalable and robust foundation for 3D medical imaging.

3D医学影像自监督学习掩码自编码器轻量化模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。