arXiv:2505.11852cs.CV2025-05NeurIPS被引 10

首个面向医学影像序列定位的基准,解决多时序图像精准对齐难题。

MedSG-Bench: A Benchmark for Medical Image Sequences Grounding

  • 构建两类序列图像定位任务:变化区域检测与一致性语义识别。
  • 覆盖76个数据集、10种模态,含9630个问答对,验证模型严重不足。
  • 提供18.8万条指令微调数据与专用模型,助力临床时序分析研究。

视觉定位在多模态大语言模型(MLLMs)中对精准感知与推理至关重要,尤其在医学影像领域。现有医疗视觉定位基准多聚焦单图场景,而真实临床应用常涉及序列图像,需跨模态精确定位病灶及追踪疾病进展(如治疗前后对比),要求细粒度跨图像语义对齐与上下文感知推理。为弥补现有基准对图像序列的忽视,我们提出首个专用于医学影像序列定位的基准——MedSG-Bench。该基准包含8个基于VQA的任务,分为两大范式:1)图像差异定位,聚焦于跨图像变化区域检测;2)图像一致性定位,强调序列图像间一致或共享语义的识别。MedSG-Bench涵盖76个公开数据集、10种医学成像模态,覆盖广泛解剖结构与疾病类型,共包含9,630个问答对。我们对通用型MLLM(如Qwen2.5-VL)与医学专用型MLLM(如HuatuoGPT-vision)进行评测,发现即便先进模型在医学序列定位任务中仍存在显著局限。为推动该领域发展,我们构建了针对序列视觉定位的大规模指令微调数据集MedSG-188K,并进一步开发了MedSeq-Grounder模型,以促进未来对医学序列图像的细粒度理解研究。相关资源已开放于https://huggingface.co/MedSG-Bench。

原文摘要 · Abstract (English)

Visual grounding is essential for precise perception and reasoning in multimodal large language models (MLLMs), especially in medical imaging domains. While existing medical visual grounding benchmarks primarily focus on single-image scenarios, real-world clinical applications often involve sequential images, where accurate lesion localization across different modalities and temporal tracking of disease progression (e.g., pre- vs. post-treatment comparison) require fine-grained cross-image semantic alignment and context-aware reasoning. To remedy the underrepresentation of image sequences in existing medical visual grounding benchmarks, we propose MedSG-Bench, the first benchmark tailored for Medical Image Sequences Grounding. It comprises eight VQA-style tasks, formulated into two paradigms of the grounding tasks, including 1) Image Difference Grounding, which focuses on detecting change regions across images, and 2) Image Consistency Grounding, which emphasizes detection of consistent or shared semantics across sequential images. MedSG-Bench covers 76 public datasets, 10 medical imaging modalities, and a wide spectrum of anatomical structures and diseases, totaling 9,630 question-answer pairs. We benchmark both general-purpose MLLMs (e.g., Qwen2.5-VL) and medical-domain specialized MLLMs (e.g., HuatuoGPT-vision), observing that even the advanced models exhibit substantial limitations in medical sequential grounding tasks. To advance this field, we construct MedSG-188K, a large-scale instruction-tuning dataset tailored for sequential visual grounding, and further develop MedSeq-Grounder, an MLLM designed to facilitate future research on fine-grained understanding across medical sequential images. The benchmark, dataset, and model are available at https://huggingface.co/MedSG-Bench

医学影像视觉定位序列分析多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。