arXiv:2606.08952cs.AI2026-06被引 1

让大模型学会从局部观察构建全局空间地图,提升空间推理能力。

AlloSpatial: Agentic Harness Framework for Spatial Reasoning in Foundation Models

论文配图:AlloSpatial: Agentic Harness Framework for Spatial Reasoning in Foundation Models
图 1 · 摘自论文原文
  • 通过认知映射沙盒将视角观察转为结构化全局空间先验
  • 在无训练条件下提升现有模型5%-18%的空间推理性能
  • 适合需要可靠空间理解的机器人与自动驾驶场景

多模态基础模型虽有显著进展,但在物理世界的空间推理上仍脆弱。核心瓶颈在于无法将局部自我中心观测转化为全局分配坐标空间表示。为此,我们提出 AlloSpatial,一种用于基础模型分配坐标空间认知的代理框架。AlloSpatial 引入 World2Mind,一个即插即用的认知映射沙盒,可将自我中心观测转化为结构化的分配坐标先验,包括分配坐标空间树和路线图,支持查询物体拓扑、几何关系、通行性及轨迹。为在重建噪声和模糊视觉证据下可靠利用这些先验,AlloSpatial 引入空间推理约束框架,实现工具使用判断、模态解耦线索收集和几何-语义仲裁。我们进一步通过基于约束门控轨迹级奖励的冷启动强化学习,将该过程内化至 Qwen3-VL。在 VSI-Bench 与 MindCube 上的实验表明,AlloSpatial 在无需训练的情况下使专有模型性能提升5%-18%,而仅靠分配坐标空间树即可在移除视觉输入时维持强空间推理能力。训练后的 AlloSpatial 代理甚至超越更大规模通用模型与竞争性空间基线,表明结构化分配坐标表示、主动工具使用与可验证推理是通往具备空间能力的基础模型的可行路径。

原文摘要 · Abstract (English)

Multimodal Foundation Models (MFMs) have made substantial progress, yet remain fragile in spatial reasoning over the physical world. A key bottleneck lies in their inability to transform local egocentric observations into a global allocentric spatial representation. To address this, we propose AlloSpatial, an agentic framework for allocentric spatial cognition in foundation models. AlloSpatial introduces World2Mind, a plug-and-play cognitive mapping sandbox that converts egocentric observations into structured allocentric priors, including Allocentric-Spatial Trees and route maps that support querying object topology, geometric relations, passability, and trajectories. To utilize these priors reliably under noisy reconstruction and ambiguous visual evidence, AlloSpatial introduces a Spatial Reasoning Harness for tool-use judgment, modality-decoupled cue collection, and geometry-semantic arbitration. We further internalize this process in Qwen3-VL through cold-start reinforcement learning with a harness-gated trajectory-level reward. Experiments on VSI-Bench and MindCube show that AlloSpatial improves proprietary models by 5%-18% in a training-free setting, while ASTs alone support strong spatial reasoning even when visual inputs are removed. The trained AlloSpatial agents further outperform larger general-purpose models and competitive spatial baselines, suggesting that structured allocentric representations, active tool use, and verifiable reasoning offer a promising route toward spatially capable foundation models.

空间推理认知映射大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。