arXiv:2503.02824cs.CVcs.AI2025-03被引 9

构建跨模态PET/CT基础模型,提升癌症影像分析通用性

Developing a PET/CT Foundation Model for Cross-Modal Anatomical and Functional Imaging

  • 采用双分支ViT+交叉注意力解码器,联合学习PET与CT图像特征
  • 在多中心数据上预训练后,在肿瘤分割等任务上显著优于专用模型
  • 适合医学影像研究者、临床AI开发人员使用

在肿瘤学中,正电子发射断层扫描-计算机断层扫描(PET/CT)广泛用于癌症诊断、分期及治疗监测,因其结合了CT提供的解剖结构信息与PET反映的代谢活性和分子表达功能信息。然而,现有基于人工智能的PET/CT分析大多依赖于从头训练的任务特定模型或小规模数据集,限制了其泛化能力与鲁棒性。为此,我们提出一种专为多模态PET/CT影像设计的基础模型方法。引入跨模态孪生掩码自编码器(FratMAE),该框架通过独立的视觉变换器(ViT)编码器处理PET和CT图像,并利用交叉注意力解码器实现模态间的协同交互。此外,还融合文本元数据以增强PET表示学习。通过在多个PET/CT数据集上进行预训练,FratMAE捕捉到复杂的跨模态关系与全局摄取模式,在下游任务中表现优异,展现出作为可泛化的基础模型的巨大潜力。

原文摘要 · Abstract (English)

In oncology, Positron Emission Tomography-Computed Tomography (PET/CT) is widely used in cancer diagnosis, staging, and treatment monitoring, as it combines anatomical details from CT with functional metabolic activity and molecular marker expression information from PET. However, existing artificial intelligence-driven PET/CT analyses rely predominantly on task-specific models trained from scratch or on limited datasets, limiting their generalizability and robustness. To address this, we propose a foundation model approach specifically designed for multimodal PET/CT imaging. We introduce the Cross-Fraternal Twin Masked Autoencoder (FratMAE), a novel framework that effectively integrates whole-body anatomical and functional or molecular information. FratMAE employs separate Vision Transformer (ViT) encoders for PET and CT scans, along with cross-attention decoders that enable synergistic interactions between modalities during masked autoencoder training. Additionally, it incorporates textual metadata to enhance PET representation learning. By pre-training on PET/CT datasets, FratMAE captures intricate cross-modal relationships and global uptake patterns, achieving superior performance on downstream tasks and demonstrating its potential as a generalizable foundation model.

PET/CT基础模型多模态医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。