arXiv:2605.31387cs.CLcs.RO2026-05

让视觉语言模型通过多轮对话协作建模空间结构,但效果仍有限。

Multi-Turn Multi-Agent Dialogue for Collaborative Reconstruction Improves VLM Performance on Spatial Reasoning, But Only Barely

论文配图:Multi-Turn Multi-Agent Dialogue for Collaborative Reconstruction Improves VLM Performance on Spatial Reasoning, But Only Barely
图 1 · 摘自论文原文
  • 设计多轮对话框架,让VLM通过语言协作重建目标空间结构。
  • 使用详细文本描述可提升重建成功率,分解图像表示也有帮助。
  • 揭示当前VLM在空间定位与指令生成上的关键瓶颈,适合研究机器人协作的学者。

机器人在复杂环境中依赖视觉理解物体与空间布局,在人机协作任务中需通过语言传达理解。视觉语言模型(VLMs)支持视觉解读、问答和指令执行,但在需要空间推理的协作对话任务中能力仍待探索。本文通过协作搭建结构的任务,结合视觉解读、语义定位、语言引导交互与动作生成,构建一个框架,使VLM通过对话从视觉与文本输入中重建目标结构。评估了开源与闭源VLM在不同交互设置、输入模态与图像表征下的表现。结果表明,现有VLM对视觉表征的空间推理依然困难。详细文本描述能提高跨模态条件下的重建成功率,而分解图像表征则有助于性能提升。这些发现揭示了当前VLM在视觉空间定位与具身指令生成方面的局限性。

原文摘要 · Abstract (English)

Robots operating in diverse environments rely on visual input to interpret objects and spatial layouts. In human-collaborative tasks, they are expected to communicate this understanding through language. Vision-language models (VLMs) support robotic tasks involving visual interpretation, question answering, and instruction following, but their capabilities in collaborative dialogue tasks requiring spatial reasoning remain underexplored. We study this gap through a collaborative structure-building task that combines visual interpretation, grounding, language-guided interaction, and action generation. We develop a framework in which VLMs use dialogue to reconstruct a target structure from visual and textual inputs. We evaluate open-weight and closed VLMs across interaction settings, input modalities, and image representations. Results show that spatial reasoning over visual representations remains difficult for the evaluated VLMs. Detailed text representations of the target yield higher reconstruction success across modality conditions, while decomposed image representations improve performance. These findings reveal limits in visual spatial grounding and grounded instruction generation for collaborative VLM agents.

空间推理视觉语言模型协作对话机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。