arXiv:2604.04133cs.CVcs.AI2026-04被引 1

无需语言数据,3D CT模型可高效提取临床任务特征。

Learning Robust Visual Features in Computed Tomography Enables Efficient Transfer Learning for Clinical Tasks

  • 用自蒸馏方法训练3D视觉特征,不依赖图文配对数据
  • 冻结主干仅用轻量探针,在7类任务上表现优于现有模型
  • 适合资源有限团队快速迁移,尤其擅长无监督场景

当前医学影像人工智能多聚焦于构建通用视觉-语言系统,但CT领域缺乏足够规模的图像-文本配对数据。现有方法通常需微调整个网络,计算成本高。本文提出VoxelFM,一个基于DINO框架的3D CT基础模型,通过自蒸馏学习语义丰富的视觉特征,无需语言监督。在七类临床任务(分类、回归、生存分析、实例检索、定位、分割、报告生成)中,使用冻结主干+轻量探针的方式评估,VoxelFM在所有类别上达到或超越四个现有模型。尽管未接受语言对齐训练,其在报告生成任务上仍优于显式进行语言对齐的模型。结果表明,现有CT基础模型作为特征提取器比作为视觉编码器更有效。模型权重与代码已公开。

原文摘要 · Abstract (English)

There is substantial interest in developing artificial intelligence systems to support radiologists across tasks ranging from segmentation to report generation. Existing computed tomography (CT) foundation models have largely focused on building generalist vision-language systems capable of tasks such as question answering and report generation. However, training reliable vision-language systems requires paired image-text data at a scale that remains unavailable in CT. Moreover, adapting the underlying visual representations to downstream tasks typically requires partial or full backbone fine-tuning, a computationally demanding process inaccessible to many research groups. Instead, foundation models should prioritise learning robust visual representations that enable efficient transfer to new tasks with minimal labelled data and without backbone fine-tuning. We present VoxelFM, a 3D CT foundation model trained with self-distillation using the DINO framework, which learns semantically rich features without language supervision. We evaluated VoxelFM across seven categories of clinically relevant downstream tasks using frozen backbone representations with lightweight probes: classification, regression, survival analysis, instance retrieval, localisation, segmentation, and report generation. VoxelFM matched or outperformed four existing CT foundation models across all task categories. Despite receiving no language supervision during pre-training, VoxelFM surpassed models explicitly trained with language-alignment objectives, including on report generation. Our results indicate that current CT foundation models perform significantly better as feature extractors for lightweight probes rather than as vision encoders for vision-language models. Model weights and training code are publicly available.

CT生成特征提取自蒸馏迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。