arXiv:2608.05999cs.RO2026-08

让机器人分步执行复杂任务,提升长程操作成功率

Beyond Flat Policies: Hierarchical Post-Training for Embodied Agents in Robotic Manipulation

论文配图:Beyond Flat Policies: Hierarchical Post-Training for Embodied Agents in Robotic Manipulation
图 1 · 摘自论文原文
  • 将任务规划与动作执行解耦,分层优化提高可解释性
  • 在线强化学习使动作策略持续改进,任务成功率显著提升
  • 适合需要长期规划的机器人操作场景,如家庭服务机器人

视觉-语言-动作(VLA)模型在机器人操作中表现出色,但现有后训练方法多将其视为扁平策略,难以显式建模任务进展,导致长程操作鲁棒性差。尽管分层方法引入任务分解,却主要依赖离线演示的监督学习,无法通过在线交互持续优化。为此,我们提出分层机器人控制框架HiRoC,将高层任务规划与低层动作执行解耦:规划器将复杂任务分解为可执行子目标以提供语义引导,执行器则通过强化学习持续优化子目标条件下的动作生成。为促进两模块协同,我们在强化学习前对齐执行器与规划器生成的子目标,缓解了规划与执行间的分布偏移。在多个机器人操作基准上的实验表明,HiRoC持续优于强基线。全面分析验证了分层后训练的有效性及各关键组件的贡献。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models have demonstrated remarkable capabilities in robotic manipulation by leveraging pretrained vision-language models. However, existing post-training methods predominantly optimize VLA models as flat policies, making it difficult to explicitly model task progression and perform robust long-horizon manipulation. Although hierarchical approaches introduce task decomposition, they mainly rely on supervised learning from offline demonstrations and cannot effectively improve execution through online interaction. To address this limitation, we propose Hierarchical Robotic Control (HiRoC), a hierarchical post-training framework that decouples high-level task planning from low-level action execution. The planner decomposes complex tasks into executable subgoals to provide explicit semantic guidance, while the executor continuously improves subgoal-conditioned action generation through reinforcement learning. To enable effective collaboration between the two modules, we further align the executor with planner-generated subgoals before reinforcement learning, mitigating the distribution misalignment between planning and execution. Extensive experiments across diverse robotic manipulation benchmarks demonstrate that HiRoC consistently outperforms strong baselines. Comprehensive analyses further validate the effectiveness of hierarchical post-training and the contribution of each key component.

机器人操作分层控制强化学习任务规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。