arXiv:2606.00602cs.CV2026-06

通过解剖结构与语义自适应预训练,提升胸部CT的医学影像表征能力。

ASAP: Advancing Medical Volumetric Representation Learning with Anatomy-aware Semantically-adaptive Pre-training

论文配图:ASAP: Advancing Medical Volumetric Representation Learning with Anatomy-aware Semantically-adaptive Pre-training
图 1 · 摘自论文原文
  • 引入解剖先验与动态语义对齐,让图像与报告精准匹配局部病变区域。
  • 在15个数据集22项任务上达到顶尖性能,小样本和分布外场景表现突出。
  • 为医学体积视觉语言预训练提供标准化评测基准,适合临床研究者使用。

从医学体数据中学习可迁移且可解释的表征仍具挑战,原因在于复杂的解剖结构以及放射科报告提供的弱而异构的监督信号。本文提出解剖感知语义自适应预训练(ASAP),一种针对大规模胸部CT扫描及其对应放射科报告的细粒度医学体数据表示学习的系统性视觉-语言预训练框架。ASAP包含三个核心组件:(1) 解剖感知知识注入模块,利用现成分割工具引入器官级结构先验,促进解剖一致的表征;(2) 语义自适应选择性对齐机制,动态关联句子级发现与局部体数据区域;(3) 语义自适应融合模块,在双模态掩码建模范式下实现解剖引导视觉特征与语义锚定文本线索的有效交互。除了方法贡献,我们构建了涵盖15个数据集和22项下游任务的全面基准,覆盖异常分类、分割、疾病预后预测、报告生成、词汇分类、跨模态检索和视觉问答。该基准提供标准化评估协议,可系统评估不同临床场景与数据条件下的表征质量。大量实验表明,ASAP在各项任务与数据集上持续达到最先进水平,尤其在监督有限和分布外情况下优势显著,验证其在学习可迁移且临床有意义的体数据表征方面的有效性。

原文摘要 · Abstract (English)

Learning transferable and interpretable representations from medical volumetric scans remains challenging due to complex anatomical structures and weak, heterogeneous supervision provided by radiology reports. In this paper, we propose Anatomy-aware Semantically-Adaptive Pre-training (ASAP), a principled vision-language pre-training framework for fine-grained medical volumetric representation learning from large-scale chest CT scans and their corresponding radiology reports. ASAP integrates three key components: (1) an anatomy-aware knowledge injection module that incorporates organ-level structural priors via off-the-shelf segmentation tool to encourage anatomically coherent representations; (2) a semantically-adaptive selective alignment mechanism that dynamically associates sentence-level findings with localized volumetric regions; and (3) a semantically-adaptive fusion module for effective interaction between anatomically informed visual features and grounded textual cues under dual-modal masked modeling paradigm. Beyond methodological contributions, we establish a comprehensive benchmark for medical volumetric vision-language pre-training on chest CT, covering 15 datasets and 22 downstream tasks spanning abnormality classification, segmentation, disease prognosis prediction, report generation, vocabulary classification, cross-modal retrieval and visual question answering. This benchmark provides standardized evaluation protocols to systematically assess representation quality under diverse clinical settings and data regimes. Extensive experiments demonstrate that ASAP consistently achieves state-of-the-art performance across tasks and datasets, with particularly pronounced gains under limited supervision and distribution shift, validating its effectiveness in learning transferable and clinically meaningful volumetric representations.

医学影像视觉语言预训练解剖结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。