让医学视觉语言模型同时看懂图像和理解文字,实现精准定位与描述。
MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models

- 用统一的词元空间编码图像区域,让文字和图像标记混排处理。
- 在184万数据上训练,多任务表现超越现有模型,定位准确率提升显著。
- 适合医疗图像分析、AI辅助诊断等需要图文联动的场景。
医学视觉语言模型擅长描述图像内容,但精确的视觉感知、分割与定位仍具挑战。现有方法要么将区域表述为坐标串,要么依赖外部模块,导致感知与理解分离,造成区域与语言对齐的表征鸿沟。我们提出MedUP,一种原生融合感知与理解的医学视觉语言模型。核心是UniMedTok,一个将掩码编码为大语言模型词汇表中离散词元的区域分词器,使模型能无缝混合掩码词元与文本。我们构建了包含184万样本的UniMed-Train数据集,涵盖文本引导分割、区域定位理解、医学VQA及基于思维链的分割任务,并引入UniMed-Bench进行统一评估。大量实验表明,MedUP在所有任务上均优于原生、代理型及双解码器医学视觉语言模型,且在分割任务上保持与专用分割器相当的竞争力,验证了统一感知与理解建模的强大潜力。
原文摘要 · Abstract (English)
Medical Vision-Language Models (Med-VLMs) excel at verbalizing visual content, yet precise visual perception, segmentation, and grounding remain challenging. Existing approaches either verbalize regions as coordinate strings or rely on external modules that decouple perception from understanding, creating representation gaps for region-language alignment. We present MedUP, a Med-VLM that natively unifies perception and understanding within a shared token space. At its core lies UniMedTok, a region tokenizer that encodes masks as discrete tokens in the LLM vocabulary, enabling the model to seamlessly interleave mask tokens with text. We curate UniMed-Train, a 1.84M-instance corpus spanning text-guided segmentation, region-grounded understanding, medical VQA and CoT-based segmentation, and introduce UniMed-Bench for unified evaluation. Extensive experiments show that MedUP outperforms native, agentic, and dual-decoder Med-VLMs across all tasks while remaining competitive with specialist segmentors, demonstrating the strong potential of unified understanding and perception modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。