arXiv:2606.25546cs.CV2026-06

针对3D CT影像设计了疾病导向的视觉语言预训练模型,提升诊断准确率。

Disease-Centric Vision-Language Pretraining with Hybrid Visual Encoding for 3D Computed Tomography

论文配图:Disease-Centric Vision-Language Pretraining with Hybrid Visual Encoding for 3D Computed Tomography
图 1 · 摘自论文原文
  • 采用CNN-ViT混合编码器,兼顾局部解剖细节与全局注意力。
  • 在CT-RATE上达84.4% AUC,60种疾病基准提升9.8% AUC。
  • 适合医学AI研究者,尤其关注3D影像多模态建模与临床部署。

视觉-语言预训练(VLP)通过利用放射科报告作为丰富的文本监督,在通用医疗AI中展现出巨大潜力,但现有方法在3D CT影像上表现不佳,主要受限于低效的视觉主干网络和粗粒度的语义对齐。为此,我们提出一种定制化的VLP框架,包含三个关键组件:(1) 采用CNN-ViT混合编码器,以3D CNN替代ViT的图像块嵌入,高效捕捉局部解剖细节,同时保持全局注意力,并兼容预训练跨模态先验;(2) 基于可学习查询令牌的疾病级对比学习机制,动态提取完整报告中的疾病特异性语义,并与对应视觉特征对齐,从而分离同一解剖区域内的不同疾病;(3) 诊断感知提示策略,使用真实临床短语和聚合疾病原型,弥合预训练与推理间的差距,提升零样本诊断可靠性。模型在CT-RATE上达到84.4% AUC(+5.1%),Rad-ChestCT上达75.4% AUC(+5.4%),在60种疾病的挑战性基准上更取得+9.8% AUC的显著提升,并在放射科报告生成任务中展现强泛化能力,验证了该方法的通用性与临床实用性。

原文摘要 · Abstract (English)

Vision-language pre-training (VLP) holds great promise for general-purpose medical AI by leveraging radiology reports as rich textual supervision, yet existing methods struggle with 3D CT imaging due to inefficient visual backbones and coarse semantic alignment. To address these issues, we propose a tailored VLP framework featuring three key components: (1) a CNN-ViT hybrid encoder that replaces ViT's patch embedding with a 3D CNN backbone to efficiently capture local anatomical details while preserving global attention and compatibility with pre-trained cross-modal priors; (2) a disease-level contrastive learning mechanism using learnable query tokens to dynamically extract disease-specific semantics from full reports and align them with corresponding visual features, thereby disentangling distinct diseases within the same anatomical region; and (3) a diagnosis-aware prompt strategy that employs real clinical phrases and aggregated disease prototypes to bridge the pre-training-inference gap and enhance zero-shot diagnostic reliability. Our model achieves state-of-the-art performance on CT-RATE (84.4% AUC, +5.1%) and Rad-ChestCT (75.4% AUC, +5.4%), with even larger gains (+9.8% AUC) on a challenging 60-disease benchmark, and demonstrates strong transferability to radiology report generation, underscoring the generality and clinical utility of our approach.

3D CT视觉语言预训练疾病识别医学AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。