arXiv:2603.02026cs.CVcs.CL2026-03被引 1

用病灶感知的视觉语言模型,让AI精准定位CT报告中的具体图像位置。

Learning to Read Where to Look: Disease-Aware Vision-Language Pretraining for 3D CT

  • 在9.8万对报告-影像数据上训练,融合疾病提示增强跨模态对齐
  • 文本到图像检索准确率31.5(提升9.3),定位误差降至36.3mm
  • 一个模型同时支持检索、分类和扫描内定位,适合临床辅助诊断

近期的3D CT视觉语言模型通过对比学习对齐影像与报告,但通常依赖有限公共数据且仅提供粗粒度全局监督。本文在单家医院收集的9.8万对报告-体积数据(5万名患者)基础上,结合公开数据集,采用类似SigLIP的对比学习,并在共享视觉-文本嵌入空间中引入基于提示的疾病监督进行训练。在CT-RATE上,模型实现最先进的文本到图像检索性能(R@10 31.5 vs. 22.2),以及具有竞争力的疾病分类能力(AUC 83.8 vs. 83.8),在Rad-ChestCT上同样表现稳定(AUC 77.0 vs. 77.3)。进一步发现放射科医生常在报告中引用特定图像(如“序列X,图像Y”),将文本描述与精确轴向位置关联。我们自动挖掘出26.2万组此类片段-切片对,提出扫描内片段定位任务——预测文本片段所指的轴向深度,使平均绝对误差降至36.3mm(特征分辨率12mm),优于最佳基线的67.0mm。加入该定位目标后,检索与分类性能保持在置信区间内不变,实现了单一统一模型完成检索、分类与扫描内定位。

原文摘要 · Abstract (English)

Recent 3D CT vision-language models align volumes with reports via contrastive pretraining, but typically rely on limited public data and provide only coarse global supervision. We train a 3D CT vision-language model on 98k report-volume pairs (50k patients) collected at a single hospital, combined with public datasets, using SigLIP-style contrastive pretraining together with prompt-based disease supervision in the shared vision-text embedding space. On CT-RATE, our model achieves state-of-the-art text-to-image retrieval (R@10 31.5 vs. 22.2) and competitive disease classification (AUC 83.8 vs. 83.8), with consistent results on Rad-ChestCT (AUC 77.0 vs. 77.3). We further observe that radiologists routinely reference specific images within their reports (e.g., ``series X, image Y''), linking textual descriptions to precise axial locations. We automatically mine 262k such snippet-slice pairs and introduce the task of intra-scan snippet localization -- predicting the axial depth referred to by a text snippet -- reducing mean absolute error to 36.3 mm at 12 mm feature resolution, compared with 67.0 mm for the best baseline. Adding this localization objective leaves retrieval and classification broadly unchanged within confidence bounds, yielding a single unified model for retrieval, classification, and intra-scan grounding.

3D CT视觉语言模型定位医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。