arXiv:2607.00881cs.CV2026-07

让AI理解空间关系更准,通过多视角地图解决复杂空间推理难题。

OmniView-Space: Reinforcing Spatial Reasoning via Multi-Perspective Spatial Mapping

论文配图:OmniView-Space: Reinforcing Spatial Reasoning via Multi-Perspective Spatial Mapping
图 1 · 摘自论文原文
  • 构建多视角空间映射,动态对齐视觉与文本空间信息。
  • 在多图空间推理任务中达到顶尖性能,超越现有方法。
  • 适合需要精准空间理解的视觉语言模型研究者使用。

空间智能仍是多模态大模型的核心挑战,因其需超越基础物体识别,构建连贯的空间场景表征。现有方法通常依赖文本推理或3D重建,但在多步推理中常因无法动态切换相机、物体或方向中心的参考系而失败。为此,我们提出OmniView-Space框架,通过多模态自我中心证据保持空间一致性。其包含三个核心组件:(1) 多视角空间映射(MPSM),将重建几何重新锚定至查询对齐的视觉认知图与文本空间图;(2) 工具引导的自我中心推理,训练一个交错策略以主动选择查询所需的自我锚点并请求对应MPSM证据;(3) 认知图蒸馏,利用MPSM生成的轨迹与自我帧奖励训练模型自主生成认知图进行推理。在单图与多图空间推理基准测试中,OmniView-Space实现最先进性能,且蒸馏后模型减少对外部几何管线依赖。

原文摘要 · Abstract (English)

Spatial intelligence remains a persistent challenge for Multimodal Large Language Models (MLLMs), as it requires coherent spatial scene representations beyond basic object recognition. Existing methods typically build such representations through textual reasoning or 3D reconstruction. However, they often falter during multi-step reasoning, particularly when required to dynamically re-anchor evidence to the specific camera-, object-, or direction-centric reference frames demanded by complex queries. To address this, we propose OmniView-Space, a framework designed to maintain spatial consistency through multimodal egocentric evidence. Our approach consists of three core components: (1) Multi-Perspective Spatial Mapping (MPSM), which re-anchors reconstructed geometry into a query-aligned visual cognitive map and a textual spatial graph; (2) Tool-Guided Egocentric Reasoning, an interleaved policy trained to actively select the ego anchor required by the query and request the corresponding MPSM evidence; and (3) Cognitive-Map Distillation, which uses MPSM-generated trajectories and ego-frame rewards to train the model to reason with self-generated cognitive maps. Experiments on single- and multi-image spatial reasoning benchmarks show that OmniView-Space achieves state-of-the-art performance. Furthermore, the distilled model maintains this performance while reducing reliance on external geometry pipelines.

空间推理多模态认知图视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。