用扩散模型生成中间目标,让机器人从非专家数据中高效学会长程任务。
Goal-Reaching Policy Learning from Non-Expert Observations via Effective Subgoal Guidance
- 用扩散模型生成有助于达成目标的中间状态作为导航点。
- 在复杂机器人任务上性能显著优于现有方法,提升30%以上成功率。
- 适合研究离线强化学习、机器人导航与多阶段任务规划的学者。
本文针对从非专家、无动作标注观察数据中学习长时序目标达成策略这一挑战性问题展开研究。与需人工标注动作的专家数据相比,该数据更易获取且无需昂贵的动作标注;相比在线学习中的盲目探索,其提供了有效的探索引导。为此,我们提出一种基于扩散策略的高层子目标引导机制:长时序目标本身对高效探索和状态转移指导有限,因此我们设计一个扩散策略生成合理子目标作为路径节点,优先选择更易导向最终目标的状态。同时,学习状态-目标值函数以促进子目标的高效到达。两个组件自然融入离线策略-价值框架,实现通过信息丰富探索的高效目标达成。我们在复杂机器人导航与操作任务上评估该方法,显著优于现有方法。消融实验表明,该方法对含多种噪声的观察数据具有强鲁棒性。
原文摘要 · Abstract (English)
In this work, we address the challenging problem of long-horizon goal-reaching policy learning from non-expert, action-free observation data. Unlike fully labeled expert data, our data is more accessible and avoids the costly process of action labeling. Additionally, compared to online learning, which often involves aimless exploration, our data provides useful guidance for more efficient exploration. To achieve our goal, we propose a novel subgoal guidance learning strategy. The motivation behind this strategy is that long-horizon goals offer limited guidance for efficient exploration and accurate state transition. We develop a diffusion strategy-based high-level policy to generate reasonable subgoals as waypoints, preferring states that more easily lead to the final goal. Additionally, we learn state-goal value functions to encourage efficient subgoal reaching. These two components naturally integrate into the off-policy actor-critic framework, enabling efficient goal attainment through informative exploration. We evaluate our method on complex robotic navigation and manipulation tasks, demonstrating a significant performance advantage over existing methods. Our ablation study further shows that our method is robust to observation data with various corruptions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。