arXiv:2512.01821cs.CV2025-12被引 4

让大模型通过视觉想象理解三维空间结构

Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling

  • 用隐式世界建模模拟人类空间想象,结合视觉生成反馈
  • 在多个基准上显著提升空间推理能力,相对坐标编码更优
  • 适合需要三维理解的多模态模型研究者

空间推理是多模态大语言模型(MLLMs)中关键但尚未充分发展的能力。现有方法主要依赖文本描述调优,存在视觉无能问题——仅通过文字符号学习空间概念,缺乏与视觉表现的关联。为此,本文提出MILO隐式空间世界建模范式,模拟人类空间想象。MILO通过视觉生成器提供几何感知反馈,将MLLM的符号推理隐式地锚定在感知经验中。为支持该范式,我们提出新型相对位置编码RePE,捕捉相机姿态的相对变换,性能优于绝对坐标系统。为训练构建了大规模几何感知生成数据集GeoGen,包含约2,241个视频和67,827个观察-动作-结果三元组。实验表明,该方法显著提升了多个基线模型在多项基准上的空间推理能力,实现对三维空间更全面的理解。

原文摘要 · Abstract (English)

Spatial reasoning, the ability to understand and interpret the 3D structure of the world, is a critical yet underdeveloped capability in Multimodal Large Language Models (MLLMs). Current methods predominantly rely on verbal descriptive tuning, which suffers from visual illiteracy, i.e., they learn spatial concepts through textual symbols alone, devoid of connection to their visual manifestations. To bridge this gap, this paper introduces MILO, an Implicit spatIaL wOrld modeling paradigm that simulates human-like spatial imagination. MILO integrates a visual generator to provide geometry-aware feedback, thereby implicitly grounding the MLLM's symbolic reasoning in perceptual experience. Complementing this paradigm, we propose RePE (Relative Positional Encoding), a novel encoding scheme that captures relative camera-pose transformations, offering superior performance over absolute coordinate systems. To support the training, we construct GeoGen, a large-scale Geometry-aware Generative dataset with approximately 2,241 videos and 67,827 observation-action-outcome triplets. Experiments demonstrate that our approach significantly enhances spatial reasoning capabilities across multiple baselines and benchmarks, offering a more holistic understanding of 3D space.

空间推理隐式建模多模态3D理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。