首个多中心全身体积PET/CT基础模型,提升肿瘤分割标注效率。
An Open Multi-Center Whole-Body FDG PET/CT Foundation Model for Tumor Segmentation

- 早期通道拼接融合影像与代谢特征,实现跨模态即时交互。
- 仅需10%标注数据即达全量训练性能,5次样本下仍优于单模态预训练。
- 开源多中心数据集支持,适合医学影像自动化研究与临床落地。
PET与CT的协同解析对肿瘤影像至关重要。现有深度学习方法多为特定任务、依赖单中心数据,或采用双分支融合方式,延迟跨模态交互并低估早期空间对应关系。为此,我们提出一个开源、多中心、全身体积FDG PET/CT基础模型,基于4个公开数据集的4,997例统一扫描。框架采用分层UNet结构与早期通道拼接,使解剖与代谢特征从首层嵌入即开始交互。引入基于零均值插补的掩码自编码目标,并结合加权全局重建损失,避免掩码区域边界出现非物理强度不连续。在下游AutoPET病灶分割任务中,该模型表现出强标签效率:仅用10%标注数据,性能即可媲美全量数据训练的模型;在极端5样本线性探测下,联合模态预训练的Dice分数高于单模态预训练。该多中心基础模型展示了标签效率与跨模态表征学习能力,为自动化肿瘤影像提供稳健、开源的基座,显著降低临床大规模人工标注需求。
原文摘要 · Abstract (English)
The synergistic interpretation of anatomical information from computed tomography (CT) and metabolic information from positron emission tomography (PET) is important to oncologic imaging. However, existing deep learning methods for PET/CT remain largely task-specific, are often trained on single-center cohorts, or adopt dual-branch fusion schemes that delay cross-modal interaction and underutilize early spatial correspondence between PET and CT. To address these limitations, we present an open-source, multi-center, whole-body FDG PET/CT foundation model utilizing 4,997 harmonized scans from four public datasets. Our framework employs hierarchical UNet-shaped backbones with early channel-wise concatenation, enabling anatomical and metabolic features to interact from the first embedding layer onward. We further introduce a masked autoencoding objective based on zero-mean imputation, combined with a weighted global reconstruction loss. This design avoids non-physical intensity discontinuities at masked-region boundaries that arise from learnable mask tokens. On downstream AutoPET lesion segmentation, the proposed models demonstrate strong label efficiency: with only 10\% of the labeled training data, they achieve performance comparable to models trained from scratch on the full dataset. Under extreme 5-shot linear probing, joint PET/CT pretraining also achieves higher Dice scores than separated-modality pretraining. This multi-center foundation model demonstrates label efficiency and cross-modality representation learning for PET/CT tumor segmentation. It provides a robust, open-source basis for advancing automated oncologic imaging, significantly reducing the need for large-scale manual annotations in clinical practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。