arXiv:2509.08311cs.CV2025-09中稿 · MICCAI 2025被引 4

SimCroP通过跨粒度对齐提升胸部CT影像的病理结构识别能力。

SimCroP: Radiograph Representation Learning with Similarity-driven Cross-granularity Pre-training

  • 基于相似性驱动的跨粒度对齐,精准匹配报告句子与影像区域。
  • 在五个公开数据集上实现图像分类与分割性能超越现有方法。
  • 适合需要细粒度医学影像理解的研究者和临床辅助系统开发者。

医学视觉-语言预训练在大量配对的放射影像与报告中学习代表性特征方面展现出巨大潜力。然而,在计算机断层扫描(CT)中,病变分布具有空间稀疏性,且报告中各病理描述之间的复杂隐含关系及其对应的影像子区域带来了额外挑战。本文提出一种针对胸部CT的相似性驱动跨粒度预训练框架(SimCroP),结合相似性驱动对齐与跨粒度融合机制以提升影像解析能力。首先,采用多模态掩码建模优化编码器,以理解影像中的精细低层语义;其次,设计相似性驱动对齐模块,使编码器能自适应选择并匹配报告中每句话对应的影像块;最后,跨粒度融合模块整合实例级与词-块级的多模态信息,帮助模型更好捕捉稀疏影像中的关键病灶结构,从而提升多尺度下游任务表现。SimCroP在大规模配对的CT-报告数据集上进行预训练,并在五个公开数据集的图像分类与分割任务上验证。实验结果表明,该方法在多项指标上优于前沿的自监督学习与医学视觉-语言预训练方法。代码与模型已开源于 https://github.com/ToniChopp/SimCroP。

原文摘要 · Abstract (English)

Medical vision-language pre-training shows great potential in learning representative features from massive paired radiographs and reports. However, in computed tomography (CT) scans, the distribution of lesions which contain intricate structures is characterized by spatial sparsity. Besides, the complex and implicit relationships between different pathological descriptions in each sentence of the report and their corresponding sub-regions in radiographs pose additional challenges. In this paper, we propose a Similarity-Driven Cross-Granularity Pre-training (SimCroP) framework on chest CTs, which combines similarity-driven alignment and cross-granularity fusion to improve radiograph interpretation. We first leverage multi-modal masked modeling to optimize the encoder for understanding precise low-level semantics from radiographs. Then, similarity-driven alignment is designed to pre-train the encoder to adaptively select and align the correct patches corresponding to each sentence in reports. The cross-granularity fusion module integrates multimodal information across instance level and word-patch level, which helps the model better capture key pathology structures in sparse radiographs, resulting in improved performance for multi-scale downstream tasks. SimCroP is pre-trained on a large-scale paired CT-reports dataset and validated on image classification and segmentation tasks across five public datasets. Experimental results demonstrate that SimCroP outperforms both cutting-edge medical self-supervised learning methods and medical vision-language pre-training methods. Codes and models are available at https://github.com/ToniChopp/SimCroP.

医学影像视觉语言预训练跨粒度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。