arXiv:2607.13892cs.CV2026-07

用器官特异性影像模式提升CT报告理解的精准度

Fine-Grained Vision-Language Pretraining with Organ-Conditioned Pattern Tokens for CT Understanding

论文配图:Fine-Grained Vision-Language Pretraining with Organ-Conditioned Pattern Tokens for CT Understanding
图 1 · 摘自论文原文
  • 引入器官条件模式令牌,细粒度对齐影像与文本
  • 零样本诊断准确率在两个公开数据集上分别达84.5%和69.9%
  • 适合医学影像与自然语言融合研究者参考

基于配对的CT影像与放射科报告进行视觉-语言预训练是一项可扩展但具挑战的任务。现有方法多采用全局扫描-报告对比学习,虽可扩展但掩盖了不同器官的异质性证据;而直接的器官级对齐又过于粗略,因同一解剖结构可能呈现多种不同的影像表现。因此,预训练需要更细粒度的对齐单元:器官条件的放射学模式。本文提出OCP-CT框架,保留稳定的全局CT-报告对比分支,引入器官模式接口:通过稀疏的Mixture-of-Experts(MoE)路由图像与文本令牌至潜在放射学模式,可学习的槽位将路由后的令牌映射为连续的模式令牌,并通过成对令牌对比学习,与从报告中构建的临床相似性软目标对齐。在公开的CT-RATE和RAD-ChestCT基准上,OCP-CT实现了零样本异常检测的平均AUROC分别为84.5%和69.9%,相比最强先前结果分别提升6.7和0.8个百分点。

原文摘要 · Abstract (English)

Computed tomography (CT) vision-language pretraining from paired volumes and radiology reports is a scalable yet challenging task. Existing methods commonly adopt global scan-report contrast, which is scalable but obscures heterogeneous organ evidence. Meanwhile, direct organ-level alignment remains coarse, since the same anatomy can exhibit multiple distinct radiological appearances. Therefore, pretraining requires a finer alignment unit: the organ-conditioned radiological pattern. In this work, we propose OCP-CT, an organ-conditioned pattern-token alignment framework for CT vision-language pretraining. Specifically, OCP-CT preserves a stable global CT-report contrastive branch and introduces an organ pattern interface: sparse Mixture-of-Experts (MoE) routes image and text tokens according to latent radiological patterns, learnable slots query the routed tokens into continuous pattern tokens, and paired token contrast aligns image-text pattern tokens with structured soft targets built from report-derived clinical similarity. On the publicly available CT-RATE and RAD-ChestCT benchmarks, OCP-CT achieves average AUROCs of 84.5% and 69.9% for zero-shot abnormality diagnosis, respectively. Compared with the strongest prior reported results, these results yield absolute AUROC gains of 6.7 and 0.8 percentage points.

医学影像视觉语言模式识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。