arXiv:2508.10528cs.CVcs.AI2025-08被引 9

构建大规模医学图文对齐数据集,提升病灶精确定位能力

Med-GLIP: Advancing Medical Language-Image Pre-training with Large-scale Grounded Dataset

  • 基于530万条区域标注,覆盖7种影像模态的细粒度医学图像对齐数据集
  • 在多个基准上超越现有模型,实现从器官到病灶的多粒度识别
  • 适用于医学视觉问答与报告生成,助力智能诊断系统开发

医学图像定位旨在将自然语言描述与医学图像中的特定区域对齐,是智能诊断、视觉问答(VQA)和自动报告生成(MRG)的基础任务。然而,现有研究受限于模态覆盖不足、标注粗粒度以及缺乏统一通用的对齐框架。为此,我们构建了大规模医学图像对齐数据集Med-GLIP-5M,包含超过530万条跨七种影像模态的区域级标注,涵盖多种解剖结构与病理发现。该数据集支持分割与定位任务,具有从器官边界到细粒度病灶的层级标签。基于此,我们提出Med-GLIP,一种模态感知的对齐框架,在Med-GLIP-5M上训练。无需显式设计专家模块,其通过多样化数据隐式学习层次化语义理解,能够精准区分如肺部与肺炎病灶等多粒度结构。大量实验表明,Med-GLIP在多个对齐基准上持续优于现有先进模型。进一步将空间输出融入下游任务,如医学VQA与报告生成,均带来显著性能提升。数据集即将公开。

原文摘要 · Abstract (English)

Medical image grounding aims to align natural language phrases with specific regions in medical images, serving as a foundational task for intelligent diagnosis, visual question answering (VQA), and automated report generation (MRG). However, existing research is constrained by limited modality coverage, coarse-grained annotations, and the absence of a unified, generalizable grounding framework. To address these challenges, we construct a large-scale medical grounding dataset Med-GLIP-5M comprising over 5.3 million region-level annotations across seven imaging modalities, covering diverse anatomical structures and pathological findings. The dataset supports both segmentation and grounding tasks with hierarchical region labels, ranging from organ-level boundaries to fine-grained lesions. Based on this foundation, we propose Med-GLIP, a modality-aware grounding framework trained on Med-GLIP-5M. Rather than relying on explicitly designed expert modules, Med-GLIP implicitly acquires hierarchical semantic understanding from diverse training data -- enabling it to recognize multi-granularity structures, such as distinguishing lungs from pneumonia lesions. Extensive experiments demonstrate that Med-GLIP consistently outperforms state-of-the-art baselines across multiple grounding benchmarks. Furthermore, integrating its spatial outputs into downstream tasks, including medical VQA and report generation, leads to substantial performance gains. Our dataset will be released soon.

医学图像图文对齐多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。