arXiv:2510.10011cs.CV2025-10CVPR被引 32

MIMO让医学视觉模型能理解图像细节并定位答案位置。

MIMO: A medical vision language model with visual referring multimodal input and pixel grounding multimodal output

  • 输入融合图像线索与文本指令,理解复杂医学图像
  • 输出将术语与图像像素精准对应,实现定位式回答
  • 基于89.5万样本的MIMOSeg数据集,支持多模态任务

当前医学视觉语言模型广泛用于医学视觉问答任务,但存在两个问题:输入仅依赖文本指令,缺乏对图像中视觉线索的直接理解;输出仅提供文本答案,无法关联图像中的关键区域。为此,我们提出统一的医学视觉语言模型MIMO,具备视觉指代多模态输入和像素定位多模态输出能力。MIMO不仅能结合视觉线索与文本指令理解复杂医学图像与语义,还能将文本输出中的医学术语在图像中进行像素级定位。为克服医学领域数据稀缺问题,我们构建了MIMOSeg,一个包含89.5万样本的综合性医学多模态数据集,涵盖四种不同视角,覆盖基础指令遵循与复杂问答任务,支持多模态输入与输出。我们在多个下游医学多模态任务上进行实验,结果表明MIMO能独特地整合视觉指代与像素定位能力,这是以往模型不具备的。

原文摘要 · Abstract (English)

Currently, medical vision language models are widely used in medical vision question answering tasks. However, existing models are confronted with two issues: for input, the model only relies on text instructions and lacks direct understanding of visual clues in the image; for output, the model only gives text answers and lacks connection with key areas in the image. To address these issues, we propose a unified medical vision language model MIMO, with visual referring Multimodal Input and pixel grounding Multimodal Output. MIMO can not only combine visual clues and textual instructions to understand complex medical images and semantics, but can also ground medical terminologies in textual output within the image. To overcome the scarcity of relevant data in the medical field, we propose MIMOSeg, a comprehensive medical multimodal dataset including 895K samples. MIMOSeg is constructed from four different perspectives, covering basic instruction following and complex question answering with multimodal input and multimodal output. We conduct experiments on several downstream medical multimodal tasks. Extensive experimental results verify that MIMO can uniquely combine visual referring and pixel grounding capabilities, which are not available in previous models.

医学视觉多模态像素定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。