arXiv:2603.14579cs.CVcs.LG2026-03中稿 · MICCAI 2026

用语义采样提升医疗影像空间定位能力,显著提高模型准确率

Medical Image Spatial Grounding with Semantic Sampling

论文配图:Medical Image Spatial Grounding with Semantic Sampling
图 1 · 摘自论文原文
  • 提出语义采样方法,在推理时优化视觉语言模型的空间定位
  • 在新基准MIS-Ground上,使Qwen3-VL-32B准确率提升13.06%
  • 适用于需要精准解剖结构定位的医疗AI研究与应用

视觉语言模型(VLMs)在图像和视频的视觉定位中展现出巨大潜力。在医学影像领域,VLMs连接了目标检测与分割,实现理解与生成。然而,三维医学影像中解剖结构的空间定位面临诸多独特挑战。本研究分析了图像模态、切片方向与坐标系统对视觉组件的影响,以及解剖、方向与关系术语对语言组件的影响。结果表明,标签、边界框与掩码叠加等提示方式对模型空间定位能力影响各异。为此,我们构建了MIS-Ground基准,全面测试VLM在医疗影像空间定位中的脆弱性,并公开于github.com/asy51/mis-ground。此外,提出MIS-SemSam——一种低成本、推理时、模型无关的优化方法,通过语义采样提升空间定位能力。实验显示,该方法使Qwen3-VL-32B在MIS-Ground上的准确率提升13.06%。

原文摘要 · Abstract (English)

Vision language models (VLMs) have shown significant promise in visual grounding for images as well as videos. In medical imaging research, VLMs represent a bridge between object detection and segmentation, and report understanding and generation. However, spatial grounding of anatomical structures in the three-dimensional space of medical images poses many unique challenges. In this study, we examine image modalities, slice directions, and coordinate systems as differentiating factors for vision components of VLMs, and the use of anatomical, directional, and relational terminology as factors for the language components. We then demonstrate that visual and textual prompting systems such as labels, bounding boxes, and mask overlays have varying effects on the spatial grounding ability of VLMs. To enable measurement and reproducibility, we introduce MIS-Ground, a benchmark that comprehensively tests a VLM for vulnerabilities against specific modes of Medical Image Spatial Grounding. We release MIS-Ground to the public at github.com/asy51/mis-ground. In addition, we present MIS-SemSam, a low-cost, inference-time, and model-agnostic optimization of VLMs that improves their spatial grounding ability with the use of Semantic Sampling. We find that MIS-SemSam improves the accuracy of Qwen3-VL-32B on MIS-Ground by 13.06%.

医疗影像空间定位视觉语言模型语义采样

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。