arXiv:2502.01800cs.ROcs.AI2025-02ICML被引 5

用流模型自动学习随机化分布,提升机器人技能训练的鲁棒性。

Flow-based Domain Randomization for Learning and Sequencing Robotic Skills

  • 基于归一化流构建可学习的环境随机化分布。
  • 在六个仿真和一个真实机器人任务中表现优于传统方法。
  • 可用于检测分布外情况,支持不确定感知的多步规划。

强化学习中的领域随机化是一种提升控制策略在仿真中训练后鲁棒性的成熟技术。通过在训练中随机化环境属性,学习到的策略能对随机化维度上的不确定性具有鲁棒性。以往环境分布通常手工设定,本文提出通过熵正则化奖励最大化,利用基于归一化流的神经采样分布自动发现最优分布。该架构更具灵活性,比现有学习简单参数化分布的方法表现更优,在六个仿真和一个真实机器人任务中得到验证。最后,结合特权价值函数,探索了所学采样分布用于分布外检测的能力,支持不确定感知的多步操作规划。

原文摘要 · Abstract (English)

Domain randomization in reinforcement learning is an established technique for increasing the robustness of control policies trained in simulation. By randomizing environment properties during training, the learned policy can become robust to uncertainties along the randomized dimensions. While the environment distribution is typically specified by hand, in this paper we investigate automatically discovering a sampling distribution via entropy-regularized reward maximization of a normalizing-flow-based neural sampling distribution. We show that this architecture is more flexible and provides greater robustness than existing approaches that learn simpler, parameterized sampling distributions, as demonstrated in six simulated and one real-world robotics domain. Lastly, we explore how these learned sampling distributions, combined with a privileged value function, can be used for out-of-distribution detection in an uncertainty-aware multi-step manipulation planner.

机器人学习强化学习域随机化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。