arXiv:2605.21906cs.CV2026-05被引 2

构建通用CT影像基础模型,统一多任务医学分析

Universal CT Representations from Anatomy to Disease Phenotype through Agglomerative Pretraining

论文配图:Universal CT Representations from Anatomy to Disease Phenotype through Agglomerative Pretraining
图 1 · 摘自论文原文
  • 分三阶段聚合持续预训练,融合切片、整体和报告语义
  • 在5类任务中超越或持平专用模型,肿瘤分期特征可被有效捕捉
  • 适合医学影像研究者、临床医生及多任务模型开发者使用

计算机断层扫描(CT)是三维医学影像的核心,但现有基于CT的人工智能模型仍局限于特定任务,如分割、分类、配准和报告分析。本文提出FlexiCT,一种通过聚合持续预训练构建的CT基础模型家族,其在来自56个公开数据集的266,227个CT体积上进行训练,形成大规模公共资源用于CT表征学习。FlexiCT采用三阶段预训练策略:二维轴向预训练、三维解剖预训练以及报告引导的语义对齐,支持切片级、体积分级及视觉-语言分析。在五个下游任务类别(分割、分类、配准、视觉-语言理解与临床检索)中,FlexiCT在多个基准测试上达到或超过先前专用方法的表现。其嵌入向量能沿肿瘤分期梯度有序排列CT扫描,表明该模型可捕捉与疾病表型相关的影像特征。项目主页与代码已公开:https://ricklisz.github.io/flexict.github.io 及 https://github.com/ricklisz/FlexiCT。

原文摘要 · Abstract (English)

Computed tomography (CT) is a central to three-dimensional medical imaging, yet CT-based artificial intelligence remains fragmented across task-specific models for segmentation, classification, registration, and report analysis. Here we present FlexiCT, a family of CT foundation models trained by agglomerative continual pretraining on 266,227 CT volumes from 56 publicly available datasets, forming a large-scale public resource for CT representation learning. FlexiCT uses agglomerative pretraining across three stages: two-dimensional axial pretraining, three-dimensional anatomical pretraining and report-guided semantic alignment. This training strategy supports slice-level, volume-level and vision-language analysis. Across five downstream task families (segmentation, classification, registration, vision-language understanding and clinical retrieval), FlexiCT matches or exceeds prior task-specific approaches on multiple benchmarks. Its embeddings further organize CT scans along gradients associated with various tumor stages, suggesting that CT foundation models can capture imaging features relevant to disease phenotype characterization. Project page and code are available at: https://ricklisz.github.io/flexict.github.io and https://github.com/ricklisz/FlexiCT.

CT基础模型医学影像多任务学习视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。