arXiv:2603.08383cs.RO2026-03

用状态图引导规划,让机器人长时序操作更可靠。

MoMaStage: Skill-State Graph Guided Planning and Closed-Loop Execution for Long-Horizon Indoor Mobile Manipulation

  • 构建技能-状态图约束任务分解与执行路径。
  • 真实场景测试中任务成功率显著提升,错误率下降。
  • 适合需要长期自主操作的移动机器人系统。

室内移动操作(MoMA)使机器人能将自然语言指令转化为物理动作,但长时序执行仍面临误差累积和环境适应性差的问题。现有学习方法难以保持长期逻辑一致性,而依赖显式场景建模的方法则因结构僵化降低动态环境下的适应性。为此,我们提出MoMaStage,一种基于视觉-语言模型的结构化框架,无需显式场景地图。该框架将视觉-语言模型嵌入分层技能库与拓扑感知的技能-状态图中,限定任务分解与技能组合在可行转移空间内,确保生成计划在逻辑与拓扑上均有效。为增强鲁棒性,系统引入闭环执行机制,通过本体感知反馈检测偏差,并触发图约束的语义重规划,保持计划与实际动作一致。在物理丰富仿真及真实环境中的大量实验表明,MoMaStage优于当前最优基线,在长时序操作中实现更高规划成功率、更低令牌开销,并显著提升整体任务完成率。

原文摘要 · Abstract (English)

Indoor mobile manipulation (MoMA) enables robots to translate natural language instructions into physical actions, yet long-horizon execution remains challenging due to cascading errors and limited generalization across diverse environments. Learning-based approaches often fail to maintain logical consistency over extended horizons, while methods relying on explicit scene representations impose rigid structural assumptions that reduce adaptability in dynamic settings. To address these limitations, we propose MoMaStage, a structured vision-language framework for long-horizon MoMA that eliminates the need for explicit scene mapping. MoMaStage grounds a Vision-Language Model (VLM) within a Hierarchical Skill Library and a topology-aware Skill-State Graph, constraining task decomposition and skill composition within a feasible transition space. This structured grounding ensures that generated plans remain logically consistent and topologically valid with respect to the agent's evolving physical state. To enhance robustness, MoMaStage incorporates a closed-loop execution mechanism that monitors proprioceptive feedback and triggers graph-constrained semantic replanning when deviations are detected, maintaining alignment between planned skills and physical outcomes. Extensive experiments in physics-rich simulations and real-world environments demonstrate that MoMaStage outperforms state-of-the-art baselines, achieving substantially higher planning success, reducing token overhead, and significantly improving overall task success rates in long-horizon mobile manipulation. Video demonstrations are available on the project website: https://chenxuli-cxli.github.io/MoMaStage/.

移动操作长时序规划视觉语言模型闭环控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。