arXiv:2605.13544cs.CV2026-05被引 1

提出跨解剖结构对比学习,提升3D医学图像理解的鲁棒性。

CA-GCL: Cross-Anatomy Global-Local Contrastive Learning for Robust 3D Medical Image Understanding

论文配图:CA-GCL: Cross-Anatomy Global-Local Contrastive Learning for Robust 3D Medical Image Understanding
图 1 · 摘自论文原文
  • 通过全局-局部对比学习分离不同解剖结构的文本表征。
  • 在标准和非标准提示下均保持稳定,准确率更高且波动更小。
  • 适合临床部署,对描述不完整或变化敏感的场景有强适应力。

细粒度视觉-语言预训练(FVLP)在3D医学图像理解中展现出巨大潜力,能够对齐解剖级视觉表示与对应文本描述。然而,现有FVLP方法常出现文本嵌入空间的表征坍塌问题,导致不同解剖结构的文本嵌入高度聚集、难以区分。这种分布退化使模型对提示词变化极度敏感,阻碍其可靠临床应用。为此,我们提出一种新型跨解剖结构全局-局部对比学习框架(CA-GCL)。CA-GCL引入全局对比目标,强制解剖类别在潜在空间中分离,有效抑制局部对齐引发的聚集倾向。同时,基于置换不变性和部分完整性设计临床感知文本增强策略,提升对描述不完整性的鲁棒性。在CT-RATE和Rad-ChestCT数据集上的大量实验表明,CA-GCL在零样本异常检测性能上与现有VLP方法相当,但在提示词变化下的鲁棒性显著提升:在标准模板下获得更高平均AUC且方差更低;在非标准模板下表现稳定,而基线模型则大幅下降。结果验证了CA-GCL在鲁棒3D医学图像理解中的有效性。

原文摘要 · Abstract (English)

Fine-grained Vision-Language Pre-training (FVLP) demonstrates significant potential in 3D medical image understanding by aligning anatomy-level visual representations with corresponding textual descriptions. However, existing FVLP paradigms often suffer from severe representation collapse in the textual embedding space, where text embeddings of distinct anatomical structures become highly clustered and indistinguishable. This distributional degeneracy renders the model hypersensitive to prompt variations, hindering reliable clinical deployment. To address these challenges, we propose a novel Cross-Anatomy Global-Local Contrastive Learning framework (CA-GCL). CA-GCL introduces a global contrastive objective that enforces separation between anatomical categories in the latent space, effectively counteracting the aggregation tendency induced by local alignment. Furthermore, we incorporate a clinical-aware text augmentation strategy based on permutation invariance and partial completeness to enhance robustness against descriptive incompleteness. Extensive evaluations on the CT-RATE and Rad-ChestCT datasets show that CA-GCL achieves comparable zero-shot abnormality detection performance to existing VLP paradigms, while demonstrating substantially better robustness to prompt variations: on canonical templates it obtains higher mean AUC with lower variance, and on non-canonical templates it remains stable whereas baselines degrade markedly. These results validate CA-GCL as an effective framework for robust 3D medical image understanding.

3D医学图像对比学习视觉语言模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。