用高质量医学定位数据提升视觉语言模型的临床描述准确性
MedGround: Bridging the Evidence Gap in Medical Vision-Language Models with Verified Grounding Data
- 通过分割标注自动生成精准的医学定位数据,利用专家掩码定位解剖结构
- 训练后模型在35K样本上实现更优的定位准确率与多对象语义区分能力
- 适合医疗AI研究者,尤其关注医学图文对齐与可解释性的团队
视觉语言模型虽能生成逼真的临床叙述,却常缺乏视觉依据。我们提出MedGround,一个自动化管道,将分割资源转化为高质量的医学指代定位数据。基于专家掩码作为空间锚点,精确提取解剖位置、形状和空间特征,并引导模型生成符合形态与位置的自然查询。通过多阶段验证系统(格式检查、几何/医学先验规则、图像判别)过滤模糊或视觉不支持样本。最终构建MedGround-35K,一个新型多模态医学数据集。实验证明,使用该数据集训练的模型在指代定位任务中表现更优,增强多目标语义消歧能力,并在未见场景中具有良好泛化性。本工作展示了可扩展的数据驱动方法,使医学语言可追溯至可验证的视觉证据。数据与代码将在论文录用后公开。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) can generate convincing clinical narratives, yet frequently struggle to visually ground their statements. We posit this limitation arises from the scarcity of high-quality, large-scale clinical referring-localization pairs. To address this, we introduce MedGround, an automated pipeline that transforms segmentation resources into high-quality medical referring grounding data. Leveraging expert masks as spatial anchors, MedGround precisely derives localization targets, extracts shape and spatial cues, and guides VLMs to synthesize natural, clinically grounded queries that reflect morphology and location. To ensure data rigor, a multi-stage verification system integrates strict formatting checks, geometry- and medical-prior rules, and image-based visual judging to filter out ambiguous or visually unsupported samples. Finally, we present MedGround-35K, a novel multimodal medical dataset. Extensive experiments demonstrate that VLMs trained with MedGround-35K consistently achieve improved referring grounding performance, enhance multi-object semantic disambiguation, and exhibit strong generalization to unseen grounding settings. This work highlights MedGround as a scalable, data-driven approach to anchor medical language to verifiable visual evidence. Dataset and code will be released publicly upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。