arXiv:2506.03642cs.CVcs.AI2025-06NeurIPS被引 29

用结构化提示+仿真数据提升视觉模型的3D空间理解能力

Spatial Understanding from Videos: Structured Prompts Meet Simulation Data

  • 通过分步推理提示分解复杂场景与问题
  • 在多个基准上显著提升预训练模型的空间推理性能
  • 适合研究机器人导航与具身智能的学者参考

视觉-空间理解能力是推断视觉输入中物体关系与布局的基础,对机器人导航和具身交互等下游任务至关重要。然而现有方法面临空间不确定性与数据稀缺问题,限制了预训练视觉-语言模型(VLM)的三维空间推理能力。为此,我们提出一种无需修改模型架构的统一框架,结合SpatialMind——一种将复杂场景与问题分解为可解释推理步骤的结构化提示策略,以及ScanForgeQA——一个通过自动化流程从多样化3D仿真场景构建的大规模问答数据集,用于微调。大量实验表明,该提示与微调策略在多个基准上均具有效果,且联合使用时表现更优,为未来视觉-空间理解研究提供了新思路。

原文摘要 · Abstract (English)

Visual-spatial understanding, the ability to infer object relationships and layouts from visual input, is fundamental to downstream tasks such as robotic navigation and embodied interaction. However, existing methods face spatial uncertainty and data scarcity, limiting the 3D spatial reasoning capability of pre-trained vision-language models (VLMs). To address these challenges, we present a unified framework for enhancing 3D spatial reasoning in pre-trained VLMs without modifying their architecture. This framework combines SpatialMind, a structured prompting strategy that decomposes complex scenes and questions into interpretable reasoning steps, with ScanForgeQA, a scalable question-answering dataset built from diverse 3D simulation scenes through an automated construction process designed for fine-tuning. Extensive experiments across multiple benchmarks demonstrate the individual and combined effectiveness of our prompting and fine-tuning strategies, and yield insights that may inspire future research on visual-spatial understanding.

空间理解视觉语言模型仿真数据结构化提示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。