arXiv:2603.06663cs.CVcs.AI2026-03AAAI被引 3

用场景图增强视觉提示,让多模态模型更好理解物体空间关系。

Graph-of-Mark: Promote Spatial Reasoning in Multimodal Language Models with Graph-Based Visual Prompting

  • 在图像上叠加场景图作为视觉提示,显式建模物体间关系。
  • 零样本下视觉问答与定位准确率最高提升11个百分点。
  • 适合需要空间推理的多模态应用,如智能导航、机器人感知。

训练无关的视觉提示技术(如Set-of-Mark)通过将输入图像划分为物体区域并添加带编号的标记框来增强多模态语言模型(MLM)的定位能力。然而,这些方法将标记物体视为孤立实体,忽视了它们之间的关系。为此,我们提出首个像素级视觉提示方法Graph-of-Mark(GoM),在输入图像上叠加场景图以支持空间推理。我们在3个开源MLM和4个数据集上评估GoM,进行了对绘制组件的广泛消融实验,并研究了文本提示中附加图描述的影响。结果表明,GoM在零样本条件下持续提升了MLM对物体位置和相对方向的理解能力,使视觉问答和定位任务的基础准确率最高提升11个百分点。

原文摘要 · Abstract (English)

Recent advances in training-free visual prompting, such as Set-of-Mark, have emerged as a promising direction for enhancing the grounding capabilities of multimodal language models (MLMs). These techniques operate by partitioning the input image into object regions and annotating them with marks, predominantly boxes with numeric identifiers, before feeding the augmented image to the MLM. However, these approaches treat marked objects as isolated entities, failing to capture the relationships between them. On these premises, we propose Graph-of-Mark (GoM), the first pixel-level visual prompting technique that overlays scene graphs onto the input image for spatial reasoning tasks. We evaluate GoM across 3 open-source MLMs and 4 different datasets, conducting extensive ablations on drawn components and investigating the impact of auxiliary graph descriptions in the text prompt. Our results demonstrate that GoM consistently improves the zero-shot capability of MLMs in interpreting object positions and relative directions, improving base accuracy in visual question answering and localization up to 11 percentage points.

空间推理视觉提示多模态场景图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。