arXiv:2603.05291cs.RO2026-03中稿 · ICML

让机器人高层规划与底层控制在线协同优化,提升长时序操作能力。

Online Self-Training for Co-Adaptation in Hierarchical Diffusion Policies

  • 通过环境反馈筛选成功轨迹,双向蒸馏优化规划与控制模块。
  • 在CALVIN基准上,小模型超越两倍大的离线模型性能。
  • 适合需要在线自适应的复杂机器人任务场景。

层级策略将语言驱动的长时序机器人操作分解为高层规划器与底层控制器。然而,两者有效协同需共享兼容的子目标分布。本文提出ORCHID,一种自训练框架,通过迭代精炼实现层次扩散策略的稳定在线优化。该方法利用环境反馈过滤策略样本,识别规划与控制协同成功的轨迹,并通过监督学习将其回传至两个模块。这一过程引发双向协同进化:规划器基于控制器的实际可达能力设定子目标,控制器则针对规划器生成的轨迹结构进行专精。得益于对筛选后在线样本的监督蒸馏,ORCHID避免了基于梯度的在线强化学习训练中扩散模型常见的不稳定性。在CALVIN基准上,一个轻量级初始弱模型的表现优于纯离线方法,包括一个规模为其两倍的视觉-语言-动作模型。

原文摘要 · Abstract (English)

Hierarchical policies decompose language-conditioned long-horizon robotic manipulation into a high-level planner and a low-level controller. However, effective coordination between HL and LL requires that both components operate on compatible subgoal distributions. We propose ORCHID, a self-training framework that enables stable online improvement of hierarchical diffusion policies by aligning planning and control through iterative refinement. By filtering policy samples via environment feedback, ORCHID identifies trajectories where the planner and controller are jointly successful and distills them back into both modules via supervised learning. This process induces a bidirectional co-adaptation: the planner grounds its subgoals in the actual reaching capabilities of the controller, while the controller specializes in the trajectory structures the planner produces. By relying on supervised distillation of filtered on-policy samples, ORCHID avoids the instability typical of online hierarchical gradient-based RL training with diffusion models. On the CALVIN benchmark, ORCHID allows a lightweight, initially weak model to outperform pure offline methods, including a Vision-Language-Action model twice its size.

机器人扩散模型自训练协同优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。