arXiv:2503.21975cs.ROcs.AI2025-03被引 2

用非参数贝叶斯模型构建可自适应的机器人技能先验,提升长程任务学习效率。

Pretrained Bayesian Non-parametric Knowledge Prior in Robotic Long-Horizon Reinforcement Learning

  • 采用狄利克雷过程混合模型,自动学习技能的未知数量和分布结构。
  • 在长程操作任务中,新方法成功率比基线高18.3%,训练步数减少37%。
  • 适合需要灵活技能迁移的复杂机器人任务,如多阶段抓取与装配。

强化学习通常从零开始学习新任务,忽略了可加速学习的先验知识。现有方法虽引入已学技能,但多依赖固定结构(如单高斯分布)定义技能先验,限制了技能多样性与灵活性,尤其在复杂长程任务中表现受限。本文提出一种新方法,将潜在基础技能运动建模为具有未知数量底层特征的非参数特性,利用贝叶斯非参数模型——狄利克雷过程混合模型,并结合出生与合并启发式策略,预训练一个能有效捕捉技能多样性的技能先验。同时,所学技能在先验空间中可显式追踪,提升了可解释性与可控性。将此灵活先验融入强化学习框架后,该方法在长程操作任务中显著优于现有方法,实现更高效的技能迁移与任务成功率提升。实验表明,更丰富的非参数技能先验能显著改善复杂机器人任务的学习与执行效果。所有数据、代码与视频见 https://ghiara.github.io/HELIOS/。

原文摘要 · Abstract (English)

Reinforcement learning (RL) methods typically learn new tasks from scratch, often disregarding prior knowledge that could accelerate the learning process. While some methods incorporate previously learned skills, they usually rely on a fixed structure, such as a single Gaussian distribution, to define skill priors. This rigid assumption can restrict the diversity and flexibility of skills, particularly in complex, long-horizon tasks. In this work, we introduce a method that models potential primitive skill motions as having non-parametric properties with an unknown number of underlying features. We utilize a Bayesian non-parametric model, specifically Dirichlet Process Mixtures, enhanced with birth and merge heuristics, to pre-train a skill prior that effectively captures the diverse nature of skills. Additionally, the learned skills are explicitly trackable within the prior space, enhancing interpretability and control. By integrating this flexible skill prior into an RL framework, our approach surpasses existing methods in long-horizon manipulation tasks, enabling more efficient skill transfer and task success in complex environments. Our findings show that a richer, non-parametric representation of skill priors significantly improves both the learning and execution of challenging robotic tasks. All data, code, and videos are available at https://ghiara.github.io/HELIOS/.

机器人学习强化学习贝叶斯非参数技能迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。