提出双系统框架,让机器人更智能地完成复杂长时序操作任务。
Goal2Skill: Long-Horizon Manipulation with Adaptive Planning and Reflection

- 高阶规划与低阶执行分离,实现分步决策与动态调整。
- 在RMBench任务上成功率32.4%,远超最强基线的9.8%。
- 适合需要记忆、纠错和多阶段协调的复杂机械操作场景。
近期视觉-语言-动作(VLA)系统在具身操作中表现出强大能力。然而,现有VLA策略依赖有限观测窗口和端到端动作预测,在长时序、记忆依赖、部分可观测、遮挡和多阶段依赖的任务中表现脆弱。这类任务不仅需要精准的视觉运动控制,还需持续记忆、自适应任务分解和显式失败恢复。为此,我们提出一种双系统框架,用于长时序具身操作。该框架显式分离高层语义推理与底层运动执行:高层规划器采用基于VLM的代理模块,维护结构化任务记忆,实现目标分解、结果验证和错误驱动修正;底层执行器则基于VLA的视觉运动控制器,通过几何保持的过滤观测进行扩散式动作生成。两者形成规划与执行间的闭环,支持记忆感知推理、自适应重规划和鲁棒在线恢复。在代表性RMBench任务上的实验表明,所提框架显著优于基线方法,平均成功率达32.4%,而最强基线仅为9.8%。消融实验进一步验证了结构化记忆与闭环恢复对长时序操作的关键作用。
原文摘要 · Abstract (English)
Recent vision-language-action (VLA) systems have demonstrated strong capabilities in embodied manipulation. However, most existing VLA policies rely on limited observation windows and end-to-end action prediction, which makes them brittle in long-horizon, memory-dependent tasks with partial observability, occlusions, and multi-stage dependencies. Such tasks require not only precise visuomotor control, but also persistent memory, adaptive task decomposition, and explicit recovery from execution failures. To address these limitations, we propose a dual-system framework for long-horizon embodied manipulation. Our framework explicitly separates high-level semantic reasoning from low-level motor execution. A high-level planner, implemented as a VLM-based agentic module, maintains structured task memory and performs goal decomposition, outcome verification, and error-driven correction. A low-level executor, instantiated as a VLA-based visuomotor controller, carries out each sub-task through diffusion-based action generation conditioned on geometry-preserving filtered observations. Together, the two systems form a closed loop between planning and execution, enabling memory-aware reasoning, adaptive replanning, and robust online recovery. Experiments on representative RMBench tasks show that the proposed framework substantially outperforms representative baselines, achieving a 32.4% average success rate compared with 9.8% for the strongest baseline. Ablation studies further confirm the importance of structured memory and closed-loop recovery for long-horizon manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。