arXiv:2608.21170cs.CVcs.AI2026-08中稿 · EMNLP

通过视觉提示提升模型空间推理能力,效果显著。

Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds

  • 引入轻量级视觉结构提示,增强空间信息可读性。
  • 模型准确率最高提升34.0个百分点,且与训练方法互补。
  • 适合研究视觉提示与多模态推理关系的学者参考。

视觉语言模型(VLMs)在多模态推理方面进展迅速,但近期研究表明其失败常源于视觉定位与下游推理之间的交互。当前尚不明确的是,当底层推理任务不变时,视觉呈现方式如何影响模型性能与错误模式。我们在SPaRC基准上研究此问题,引入轻量级输入侧提示,在保留视觉模态的同时使空间结构更易获取。在多个VLM中,这些提示使任务准确率相比原始视觉设置最高提升34.0个百分点,并进一步与基于GRPO的训练结合,相较原始视觉输入近乎零增益,额外提升达4.6个百分点。对端到端任务求解与目标检测的分析表明,性能提升主要源于定位相关错误的减少,而规则推理仍具挑战。我们发现,视觉呈现是决定VLM基准衡量的是具身感知、下游推理,还是二者混合的关键因素。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have advanced rapidly in multimodal reasoning, yet recent work shows that their failures often reflect an interaction between visual grounding and downstream reasoning. What remains less clear is how the visual presentation of a task shapes model performance and failure modes when the underlying reasoning problem is unchanged. We study this question in SPaRC, a benchmark for grid-based visual spatial planning, by introducing lightweight input-side scaffolds that preserve the visual modality while making spatial structure more accessible. Across multiple VLMs, these scaffolds improve task accuracy over the original visual setting by up to 34.0 percentage points and further complement GRPO-based training, yielding up to 4.6 additional accuracy points compared with near-zero gains on the original visual input. Analyses on both end-to-end task solving and object detection show that these gains are closely tied to reductions in grounding-related errors, while rule reasoning remains comparatively challenging. We find that visual presentation is a central factor that determines whether VLM benchmarks measure grounded perception, downstream reasoning, or a mixture of both.

视觉提示空间推理多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。