arXiv:2511.15943cs.CV2025-11中稿 · ICLR被引 6

让医学图像理解更准:多粒度语言学习提升多标签对齐能力

Boosting Medical Visual Understanding From Multi-Granular Language Learning

  • 设计多粒度语言学习框架,融合不同层级文本描述增强对齐
  • 在多个医学数据集上超越现有方法,提升多标签识别准确率
  • 适合医疗视觉理解、多标签分类任务的研究者使用

近期图像-文本预训练进展显著提升了视觉理解能力,其中对比语言-图像预训练(CLIP)在多模态学习中起到关键作用。然而,其单标签、单一粒度对齐的局限性限制了在医学影像等复杂领域中的应用,因为医学图像常对应多种高层标签(如疾病类别)及不同标注粒度(如诊断描述、临床解释)。为此,我们提出多粒度语言学习(MGLL)框架,通过结构化多标签监督、跨粒度文本融合以及点对点约束的软标签监督,实现多标签与跨粒度对齐。MGLL采用平滑KL散度确保跨粒度一致性,同时保持计算高效,可作为即插即用模块集成至视觉-语言模型。在自建的大规模多粒度数据集上预训练,并在多个下游数据集上评估,MGLL性能优于当前主流方法。代码已开源。

原文摘要 · Abstract (English)

Recent advances in image-text pretraining have significantly enhanced visual understanding by aligning visual and textual representations. Contrastive Language-Image Pretraining (CLIP) has played a pivotal role in multimodal learning. However, its focus on single-label, single-granularity alignment limits its effectiveness in complex domains such as medical imaging, where images often correspond to multiple high-level labels (e.g., disease categories) across different annotation granularities (e.g., diagnostic description, clinical explanation). To address this, we propose Multi-Granular Language Learning (MGLL), a contrastive learning framework designed to improve both multi-label and cross-granularity alignment. MGLL leverages structured multi-label supervision, integrates textual descriptions across granularities, and introduces soft-label supervision with point-wise constraints to enhance alignment. MGLL employs smooth Kullback-Leibler (KL) divergence to ensure cross-granularity consistency while maintaining computational efficiency as a plug-and-play module for vision-language models. Pretrained on our constructed large-scale multi-granular datasets and evaluated across multiple datasets, MGLL outperforms other state-of-the-art methods in downstream tasks. The code is available at https://github.com/HUANGLIZI/MGLL.

医学视觉多标签学习多粒度对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。