arXiv:2512.19269cs.ROcs.LG2025-12被引 2

用在线回溯目标提升机器人低层策略,性能翻倍

Translating Flow to Policy via Hindsight Online Imitation

  • 通过回溯执行结果自动标注高层目标,在线优化策略
  • 在仿真与真实场景中性能提升超2倍,优于现有方法
  • 可从跨体感视频训练的规划器获取策略,适合迁移学习

近期层级机器人系统利用高层规划器生成任务计划,低层策略生成具体动作。该设计使规划器可在无动作或非机器人数据源(如视频)上训练,实现可迁移的高层指导。然而,将高层计划转化为可执行动作仍具挑战,尤其受限于高质量机器人数据的稀缺。为此,我们提出通过在线交互改进低层策略。具体方法为:收集在线轨迹,从达成结果中回溯标注对应高层目标,并聚合这些回溯重标注的经验以更新目标条件化的模仿策略。所提方法HinFlow采用二维点流作为高层规划器,在多种仿真与物理世界操作任务中,性能较基线策略提升超过2倍,显著优于现有方法。此外,该框架支持从跨体感视频数据训练的规划器获取策略,展现出可扩展且可迁移的机器人学习潜力。

原文摘要 · Abstract (English)

Recent advances in hierarchical robot systems leverage a high-level planner to propose task plans and a low-level policy to generate robot actions. This design allows training the planner on action-free or even non-robot data sources (e.g., videos), providing transferable high-level guidance. Nevertheless, grounding these high-level plans into executable actions remains challenging, especially with the limited availability of high-quality robot data. To this end, we propose to improve the low-level policy through online interactions. Specifically, our approach collects online rollouts, retrospectively annotates the corresponding high-level goals from achieved outcomes, and aggregates these hindsight-relabeled experiences to update a goal-conditioned imitation policy. Our method, Hindsight Flow-conditioned Online Imitation (HinFlow), instantiates this idea with 2D point flows as the high-level planner. Across diverse manipulation tasks in both simulation and physical world, our method achieves more than $2\times$ performance improvement over the base policy, significantly outperforming the existing methods. Moreover, our framework enables policy acquisition from planners trained on cross-embodiment video data, demonstrating its potential for scalable and transferable robot learning.

机器人学习模仿学习在线优化迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。