arXiv:2509.00465cs.ROcs.AI2025-09

让机器人听懂指令并自主导航,靠的是对空间的智能感知与推理。

Embodied Spatial Intelligence: from Implicit Scene Modeling to Spatial Reasoning

  • 用隐式神经模型构建高精度场景表示,支持自监督标定与大规模重建
  • 设计新导航基准与3D语言定位方法,提升长程决策能力
  • 适合做具身智能、多模态理解与机器人导航研究者参考

本论文提出「具身空间智能」框架,解决机器人基于自然语言指令在真实世界中感知与行动的挑战。为弥合大语言模型(LLMs)与物理具身之间的鸿沟,本文在场景表征与空间推理两方面做出贡献:感知方面,采用隐式神经模型实现鲁棒、可扩展、高精度的场景表示,涵盖自监督相机标定、高保真深度场生成与大规模场景重建;推理方面,通过引入新型导航评估基准、将语言嵌入3D空间的方法,以及状态反馈机制,增强大模型的长期决策能力。该工作为机器人在复杂语言指令下稳健感知环境并智能执行任务奠定了基础。

原文摘要 · Abstract (English)

This thesis introduces "Embodied Spatial Intelligence" to address the challenge of creating robots that can perceive and act in the real world based on natural language instructions. To bridge the gap between Large Language Models (LLMs) and physical embodiment, we present contributions on two fronts: scene representation and spatial reasoning. For perception, we develop robust, scalable, and accurate scene representations using implicit neural models, with contributions in self-supervised camera calibration, high-fidelity depth field generation, and large-scale reconstruction. For spatial reasoning, we enhance the spatial capabilities of LLMs by introducing a novel navigation benchmark, a method for grounding language in 3D, and a state-feedback mechanism to improve long-horizon decision-making. This work lays a foundation for robots that can robustly perceive their surroundings and intelligently act upon complex, language-based commands.

具身智能空间推理视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。