提出渐进式局部对齐网络,提升医学图像与文本的精准匹配能力。
Progressive Local Alignment for Medical Multimodal Pre-training
- 基于对比学习设计软区域对齐机制,支持不规则结构灵活识别。
- 通过渐进式训练迭代优化词-像素对应关系,显著提升对齐精度。
- 适用于多模态医学诊断任务,尤其适合零样本分类场景。
医学图像与文本的局部对齐对精准诊断至关重要,但因缺乏自然的局部配对及刚性区域识别方法的局限性而难以实现。传统方法依赖硬边界,引入不确定性,而医学影像需柔性软区域识别以应对不规则结构。为此,本文提出渐进式局部对齐网络(PLAN),设计基于对比学习的局部对齐方法,建立有意义的词-像素关联,并引入渐进学习策略,迭代优化这些关系,提升对齐精度与鲁棒性。结合该方法,PLAN有效增强软区域识别能力并抑制噪声干扰。在多个医学数据集上的实验表明,PLAN在短语定位、图像-文本检索、目标检测和零样本分类任务上均超越现有最优方法,为医学图像-文本对齐设立了新基准。
原文摘要 · Abstract (English)
Local alignment between medical images and text is essential for accurate diagnosis, though it remains challenging due to the absence of natural local pairings and the limitations of rigid region recognition methods. Traditional approaches rely on hard boundaries, which introduce uncertainty, whereas medical imaging demands flexible soft region recognition to handle irregular structures. To overcome these challenges, we propose the Progressive Local Alignment Network (PLAN), which designs a novel contrastive learning-based approach for local alignment to establish meaningful word-pixel relationships and introduces a progressive learning strategy to iteratively refine these relationships, enhancing alignment precision and robustness. By combining these techniques, PLAN effectively improves soft region recognition while suppressing noise interference. Extensive experiments on multiple medical datasets demonstrate that PLAN surpasses state-of-the-art methods in phrase grounding, image-text retrieval, object detection, and zero-shot classification, setting a new benchmark for medical image-text alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。