arXiv:2501.14856cs.ROcs.AI2025-01ICLR被引 3

用能量模型生成奖励,让机器人仅凭观察就能学会复杂动作。

Noise-conditioned Energy-based Annealed Rewards (NEAR): A Generative Framework for Imitation Learning from Observation

  • 用噪声扰动专家轨迹,学习数据的能量函数作为奖励
  • 在多人类动作任务上表现媲美对抗式方法
  • 避免对抗训练难题,适合物理模拟场景

本文提出一种基于能量模型的模仿学习框架——噪声条件能量衰减奖励(NEAR),可仅通过状态轨迹学习复杂的、依赖物理规律的机器人运动策略。该算法构建专家运动数据分布的多个扰动版本,利用去噪得分匹配学习平滑且明确的数据分布能量函数,并将其作为强化学习的奖励函数。同时提出渐进式切换策略,确保策略生成样本流形上的奖励始终有效。我们在复杂人形任务如行走和武术动作上进行了评估,与仅依赖状态的对抗式模仿学习方法(如AMP)对比,该框架避开了对抗训练的优化难题,在多个量化指标上达到相近性能。

原文摘要 · Abstract (English)

This paper introduces a new imitation learning framework based on energy-based generative models capable of learning complex, physics-dependent, robot motion policies through state-only expert motion trajectories. Our algorithm, called Noise-conditioned Energy-based Annealed Rewards (NEAR), constructs several perturbed versions of the expert's motion data distribution and learns smooth, and well-defined representations of the data distribution's energy function using denoising score matching. We propose to use these learnt energy functions as reward functions to learn imitation policies via reinforcement learning. We also present a strategy to gradually switch between the learnt energy functions, ensuring that the learnt rewards are always well-defined in the manifold of policy-generated samples. We evaluate our algorithm on complex humanoid tasks such as locomotion and martial arts and compare it with state-only adversarial imitation learning algorithms like Adversarial Motion Priors (AMP). Our framework sidesteps the optimisation challenges of adversarial imitation learning techniques and produces results comparable to AMP in several quantitative metrics across multiple imitation settings.

模仿学习能量模型机器人控制强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。