arXiv:2509.16673cs.CV2025-09

用疾病感知的混合方法提升医学视觉语言预训练效果

MedCutMix: A Data-Centric Approach to Improve Radiology Vision-Language Pre-training with Disease Awareness

  • 在病历文本中进行诊断句裁剪混合,增强多模态学习
  • 在4个放射科诊断数据集上表现超越现有方法
  • 适合关注医学图像与报告对齐的研究者

视觉-语言预训练(VLP)因其能减少人工标注需求并提升下游任务语义理解能力而受到越来越多关注。然而,其依赖图像-文本配对数据集,面临隐私问题及标注成本高昂的挑战。数据增强成为可行策略,但现有方法难以捕捉医学数据中细微复杂的变异,多样性有限。为此,我们提出MedCutMix,一种新型多模态疾病中心的数据增强方法。MedCutMix在医学报告中执行诊断句级裁剪混合,并建立诊断句与医学图像之间的交叉注意力,以指导影像模态内的注意流形混合。该方法在四个下游放射科诊断数据集上均优于先前方法,显著提升了放射科VLP的性能与泛化能力。

原文摘要 · Abstract (English)

Vision-Language Pre-training (VLP) is drawing increasing interest for its ability to minimize manual annotation requirements while enhancing semantic understanding in downstream tasks. However, its reliance on image-text datasets poses challenges due to privacy concerns and the high cost of obtaining paired annotations. Data augmentation emerges as a viable strategy to address this issue, yet existing methods often fall short of capturing the subtle and complex variations in medical data due to limited diversity. To this end, we propose MedCutMix, a novel multi-modal disease-centric data augmentation method. MedCutMix performs diagnostic sentence CutMix within medical reports and establishes the cross-attention between the diagnostic sentence and medical image to guide attentive manifold mix within the imaging modality. Our approach surpasses previous methods across four downstream radiology diagnosis datasets, highlighting its effectiveness in enhancing performance and generalizability in radiology VLP.

视觉语言预训练医学图像数据增强疾病感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。