arXiv:2601.00260cs.CV2026-01被引 5

用器官分离框架训练3DCT基础模型,高效提升医学影像理解能力

TotalFM: An Organ-Separated 3D-CT Foundation Model Leveraging Large-Scale Routine Clinical Radiology Data

  • 按器官分割3DCT数据,结合语言模型生成图文对进行预训练
  • 零样本器官病变分类在6个器官中5个优于对比模型,发现分类83%类别表现更优
  • 适合临床部署的3DCT视觉语言模型研究,开源代码和模型可复现

尽管放射学基础模型有望应用于多种临床任务,但3D-CT数据训练仍面临计算成本高的挑战。本研究提出TotalFM,一种基于器官分离概念的放射学基础模型,利用包含14万例系列的大规模临床数据。通过分割技术与大语言模型(LLM)处理报告,自动构建器官体积与查找句对,并结合VideoMAE自监督预训练与体积-文本对的对比学习,实现计算效率与表征能力的平衡。在零样本器官级病灶分类任务中,该模型在6个器官中的5个优于CT-CLIP,14个器官中9个优于Merlin;在零样本发现级病灶分类任务中,30个发现类别中有25个(83%)的AUROC高于Merlin。此外,在放射学报告生成任务中表现接近现有视觉-语言模型(VLMs)。结果表明,器官分离学习框架可为3D-CT基础模型的实际应用提供可行设计路径。源代码与预训练模型已公开于https://github.com/jichi-labo/TotalFM。

原文摘要 · Abstract (English)

While foundation models in radiology are expected to be applied to various clinical tasks, computational cost constraints remain a major challenge when training on 3D-CT volumetric data. In this study, we propose TotalFM, a radiological foundation model that efficiently learns the correspondence between 3D-CT images and linguistic expressions based on the concept of organ separation, utilizing a large-scale dataset of 140,000 series. By automating the creation of organ volume and finding-sentence pairs through segmentation techniques and Large Language Model (LLM)-based radiology report processing, and by combining self-supervised pre-training via VideoMAE with contrastive learning using volume-text pairs, we aimed to balance computational efficiency and representation capability. In zero-shot organ-wise lesion classification tasks, the proposed model achieved higher F1 scores in 83% (5/6) of organs compared to CT-CLIP and 64% (9/14) of organs compared to Merlin. These results suggest that the proposed model exhibits high generalization performance in a clinical evaluation setting using actual radiology report sentences. Furthermore, in zero-shot finding-wise lesion classification tasks, our model achieved a higher AUROC in 83% (25/30) of finding categories compared to Merlin. We also confirmed performance comparable to existing Vision-Language Models (VLMs) in radiology report generation tasks. Our results demonstrate that the organ-separated learning framework can serve as a realistic and effective design guideline for the practical implementation of 3D-CT foundation models. The source code and pretrained models are publicly available at https://github.com/jichi-labo/TotalFM.

3DCT基础模型器官分离视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。