arXiv:2508.15874cs.ROcs.AI2025-08被引 6

让机器人看懂空间关系,精准执行复杂操作任务

Spatial Policy: Guiding Visuomotor Robotic Manipulation with Spatial-Aware Modeling and Reasoning

  • 通过显式建模空间规划表,引导视觉预测与动作生成
  • 在Meta-World和iTHOR上分别提升33%和25%性能
  • 适合需要空间推理的机器人操控场景,真实世界可验证

以视觉为中心的分层具身模型展现出巨大潜力,但现有方法缺乏空间感知能力,难以在复杂环境中将视觉计划转化为可执行控制。为此,我们提出Spatial Policy(SP),一种统一的空间感知视觉-运动机器人操作框架,通过显式空间建模与推理实现端到端控制。首先设计空间条件驱动的具身视频生成模块,基于空间计划表进行空间引导的预测;其次提出基于流的行动预测模块,协同推断可执行动作;最后引入空间推理反馈策略,通过双阶段重规划优化空间计划表。大量实验表明,SP显著优于现有最先进方法,在Meta-World上提升超33%,在iTHOR上提升超25%,覆盖23个具身控制任务。此外,我们在真实机器人上验证了其可行性,证明该框架具备实际应用价值。代码与模型权重已开源。

原文摘要 · Abstract (English)

Vision-centric hierarchical embodied models have demonstrated strong potential. However, existing methods lack spatial awareness capabilities, limiting their effectiveness in bridging visual plans to actionable control in complex environments. To address this problem, we propose Spatial Policy (SP), a unified spatial-aware visuomotor robotic manipulation framework via explicit spatial modeling and reasoning. Specifically, we first design a spatial-conditioned embodied video generation module to model spatially guided predictions through the spatial plan table. Then, we propose a flow-based action prediction module to infer executable actions with coordination. Finally, we propose a spatial reasoning feedback policy to refine the spatial plan table via dual-stage replanning. Extensive experiments show that SP substantially outperforms state-of-the-art baselines, achieving over 33% improvement on Meta-World and over 25% improvement on iTHOR, demonstrating strong effectiveness across 23 embodied control tasks. We additionally evaluate SP in real-world robotic experiments to verify its practical viability. SP enhances the practicality of embodied models for robotic control applications. Code and checkpoints are maintained at https://plantpotatoonmoon.github.io/SpatialPolicy/.

机器人操控空间推理视觉运动具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。