arXiv:2607.27154cs.CVcs.AI2026-07被引 1

轻量级适配器让CT模型精准理解解剖结构,同时保留全局信息。

Anatomy Contextualized Adaptation of CT Foundation Models

论文配图:Anatomy Contextualized Adaptation of CT Foundation Models
图 1 · 摘自论文原文
  • 用TotalSegmentator分解CT图像为解剖区域嵌入,通过Transformer建模跨器官关系
  • 在Merlin和CT-RATE上零样本分类性能超越基线,训练时间少于1小时
  • 适合需要高精度解剖对齐的医学影像研究者,尤其关注上下文关系建模

CT视觉语言基础模型在下游任务中表现良好,但通常使用全体积表示,削弱了细粒度解剖信号。细粒度预训练虽能对齐解剖级视觉特征与特定文本,却丢失了全体积模型提供的全局上下文。现有方法需从头训练,计算成本高。我们提出解剖上下文自适应(ACA),一种轻量级框架,在不更新冻结的CT基础模型的前提下,实现解剖级视觉-语言对齐并增强全局上下文。ACA利用TotalSegmentator将CT体积分解为解剖级嵌入,通过变压器捕获跨解剖关系,并与放射科报告中提取的器官级及扫描级文本对齐。在Merlin和CT-RATE数据集上的评估显示,ACA在零样本发现分类任务中持续优于冻结的基础模型和现有细粒度方法,且一旦嵌入缓存,训练耗时不足一小时。此外,ACA的跨解剖变压器学习到的注意力权重揭示了合理的跨解剖上下文路由路径。结果表明,ACA是一种轻量级方法,可在保持并增强全局解剖上下文的同时,实现解剖基础的视觉-语言对齐。

原文摘要 · Abstract (English)

CT vision-language foundation models have demonstrated promising performance across downstream tasks, but are typically trained with whole-volume representations that dilute fine-grained anatomical signals. Fine-grained vision-language pre-training addresses this by aligning anatomy-level visual features with anatomy-specific text, but in doing so discards the global context that whole-volume models provide. Furthermore, existing fine-grained approaches train from scratch, making them computationally expensive. We introduce Anatomy Contextualized Adaptation (ACA), a lightweight framework that adapts frozen CT foundation model representations for anatomy-level vision-language alignment while enhancing global contextualization. ACA uses TotalSegmentator to decompose CT volumes into anatomy-level embeddings, which are refined via a transformer that captures cross-anatomy relationships, and aligned to both per-anatomy and scan-level text extracted from radiology reports. Evaluated on Merlin and CT-RATE, ACA consistently outperforms both the frozen foundation model baselines and existing fine-grained methods in zero-shot finding classification, while requiring less than one hour of training once embeddings are cached. The attention weights learned by ACA's inter-anatomy transformer additionally indicate plausible cross-anatomy context routing. Altogether, these results support ACA as a lightweight approach for adapting CT foundation models to anatomically grounded vision-language alignment while preserving and enhancing global anatomical context.

医学影像视觉语言解剖对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。