SPECTRE用自监督与多模态预训练,让3D CT影像模型更通用更强。
Scaling Self-Supervised and Cross-Modal Pretraining for Volumetric CT Transformers
- 分层设计局部与全局变压器,解决3D CT的计算难题
- 仅用公开数据训练,在多个任务上超越现有模型
- 支持零样本和微调,适合医疗影像研究与应用
我们提出SPECTRE,一种全基于Transformer的体积计算机断层扫描(CT)基础模型。该方法采用可扩展的3D视觉变换器架构,结合现代自监督与视觉-语言预训练策略,学习通用的CT表征。体积CT存在极端令牌扩展、几何各向异性及弱或噪声临床标注等独特挑战,使标准Transformer与对比学习方法直接应用无效。框架联合优化局部变换器以提取高分辨率体积特征,以及全局变换器以建模整幅扫描上下文,使大规模3D注意力在计算上可行。值得注意的是,SPECTRE仅在公开可用的CT数据集上训练,证明无需私有数据即可获得高性能、可泛化的表征。预训练结合DINO式自蒸馏与基于SigLIP的视觉-语言对齐,利用配对放射科报告,生成既几何一致又临床有意义的特征。在多个CT基准测试中,SPECTRE在零样本与微调设置下持续优于先前的CT基础模型,确立其为可扩展、开放且完全基于Transformer的3D医学影像基础模型。
原文摘要 · Abstract (English)
We introduce SPECTRE, a fully transformer-based foundation model for volumetric computed tomography (CT). Our Self-Supervised & Cross-Modal Pretraining for CT Representation Extraction (SPECTRE) approach utilizes scalable 3D Vision Transformer architectures and modern self-supervised and vision-language pretraining strategies to learn general-purpose CT representations. Volumetric CT poses unique challenges, such as extreme token scaling, geometric anisotropy, and weak or noisy clinical supervision, that make standard transformer and contrastive learning recipes ineffective out of the box. The framework jointly optimizes a local transformer for high-resolution volumetric feature extraction and a global transformer for whole-scan context modeling, making large-scale 3D attention computationally tractable. Notably, SPECTRE is trained exclusively on openly available CT datasets, demonstrating that high-performing, generalizable representations can be achieved without relying on private data. Pretraining combines DINO-style self-distillation with SigLIP-based vision-language alignment using paired radiology reports, yielding features that are both geometrically consistent and clinically meaningful. Across multiple CT benchmarks, SPECTRE consistently outperforms prior CT foundation models in both zero-shot and fine-tuned settings, establishing SPECTRE as a scalable, open, and fully transformer-based foundation model for 3D medical imaging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。