通过建模正常解剖结构提升医学图像语义密度,改善图文对齐效果。
Boosting Vision Semantic Density with Anatomy Normality Modeling for Medical Vision-language Pre-training
- 用疾病级对比学习增强视觉语义,区分正常与异常样本。
- 基于VQ-VAE建模解剖正常分布,放大异常信号差异。
- 在54种疾病上达84.9%平均AUC,零样本性能领先。
视觉语言预训练(VLP)在构建多功能通用医疗诊断能力方面具有巨大潜力。然而,低信噪比的医学图像与高信噪比的报告之间存在语义密度差距,导致视觉对齐偏差。本文提出提升视觉语义密度以改善对齐效果:一方面,通过疾病级视觉对比学习增强模型对各解剖结构正常与异常样本的区分能力;另一方面,引入解剖正常性建模方法,利用VQ-VAE在潜在空间重建正常视觉嵌入,通过异常样本的分布偏移放大异常信号,增强模型对异常特征的感知与判别能力。强化后的视觉表示有效捕捉诊断相关语义,促进与诊断报告更高效准确的对齐。我们在两个胸部CT数据集CT-RATE和Rad-ChestCT,以及一个腹部CT数据集MedVL-CT69K上进行了大量实验,全面评估了胸腹CT场景下的多种诊断任务表现,实现了最先进的零样本性能。特别地,该方法在15个器官的54种疾病上平均AUC达到84.9%,显著优于现有方法。此外,我们还展示了预训练模型优异的迁移学习能力。代码已开源。
原文摘要 · Abstract (English)
Vision-language pre-training (VLP) has great potential for developing multifunctional and general medical diagnostic capabilities. However, aligning medical images with a low signal-to-noise ratio (SNR) to reports with a high SNR presents a semantic density gap, leading to visual alignment bias. In this paper, we propose boosting vision semantic density to improve alignment effectiveness. On one hand, we enhance visual semantics through disease-level vision contrastive learning, which strengthens the model's ability to differentiate between normal and abnormal samples for each anatomical structure. On the other hand, we introduce an anatomical normality modeling method to model the distribution of normal samples for each anatomy, leveraging VQ-VAE for reconstructing normal vision embeddings in the latent space. This process amplifies abnormal signals by leveraging distribution shifts in abnormal samples, enhancing the model's perception and discrimination of abnormal attributes. The enhanced visual representation effectively captures the diagnostic-relevant semantics, facilitating more efficient and accurate alignment with the diagnostic report. We conduct extensive experiments on two chest CT datasets, CT-RATE and Rad-ChestCT, and an abdominal CT dataset, MedVL-CT69K, and comprehensively evaluate the diagnosis performance across multiple tasks in the chest and abdominal CT scenarios, achieving state-of-the-art zero-shot performance. Notably, our method achieved an average AUC of 84.9% across 54 diseases in 15 organs, significantly surpassing existing methods. Additionally, we demonstrate the superior transfer learning capabilities of our pre-trained model. Code is available at https://github.com/alibaba-damo-academy/ViSD-Boost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。