arXiv:2511.16160cs.CV2025-11被引 5

用视频重建带度量的空间布局,提升模型空间推理能力。

Video2Layout: Recall and Reconstruct Metric-Grounded Cognitive Map for Spatial Reasoning

  • 用连续边界坐标代替离散网格,实现精确空间计算。
  • 在主流基准上比网格模型平均提升3.24%准确率。
  • 适合需要精细空间理解的多模态模型研究者。

空间智能是多模态大语言模型(MLLMs)理解物理世界的关键。现有基于网格的认知地图方法依赖离散表示,限制了细粒度空间推理能力。为此,我们提出Video2Layout框架,从视频中重建度量基准的空间布局。该框架采用连续物体边界坐标,支持定量空间计算,有效降低自然语言描述中的歧义。方法分为两个阶段:首先在AI2THOR模拟器上构建高质量数据集,通过监督微调使模型学习从视觉输入映射到精确边界坐标;随后通过强化学习微调提升模型在真实场景的泛化能力。基于此框架,我们分析影响认知地图精度的因素,并量化其与任务表现的关系。在主流空间推理基准上,我们的模型V2LO-7B相比网格模型平均提升3.24%,验证了方法优势。

原文摘要 · Abstract (English)

Spatial intelligence is a critical frontier for Multimodal Large Language Models (MLLMs), empowering them to comprehend the physical world. Drawing inspiration from human perception mechanisms, prior studies attempt to construct a spatial understanding via grid-based cognitive maps. However, current grid-based map methods rely on discretized representations, which limit the model's ability in fine-grained spatial reasoning. To overcome this limitation, we propose Video2Layout, a framework for reconstructing metric-grounded spatial layouts from video. The framework uses continuous object boundary coordinates to enable quantitative spatial computation, which effectively reduces ambiguity in natural language descriptions of spatial relationships. Specifically, our method comprises two stages. First, in supervised fine-tuning stage, we construct a high-quality dataset from the AI2THOR simulator, which enables the model to learn the mapping from visual inputs to precise boundary coordinates. Subsequently, a reinforcement fine-tuning stage enhances the model's real-world generalization capabilities. Based on the above framework, we investigate factors that affect cognitive map accuracy and quantify its relationship with task performance. Evaluated on mainstream spatial reasoning benchmarks, our model, V2LO-7B, achieves an average improvement of 3.24\% over the model trained on grid maps, validating the superiority of our method.

空间推理视频理解多模态布局重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。