arXiv:2410.03246cs.ROcs.AI2024-10

用专家动作隐变量提升机器人行走的自然性和泛化能力

Latent Action Priors for Locomotion with Deep Reinforcement Learning

  • 从少量专家演示中学习动作隐变量作为先验知识
  • 使强化学习策略探索更高效,迁移任务性能显著提升
  • 适合需要自然运动控制的机器人研究者

深度强化学习让机器人通过与环境交互学习复杂行为,但算法无约束导致结果脆弱且不自然,尤其在直接关节扭矩控制中难以融入归纳偏置。本文提出一种针对行走控制的归纳偏置:从少量专家演示中学习的动作隐变量。该先验使策略能直接利用专家动作中的知识,促进更高效的探索。实验表明,智能体性能可超越演示者的奖励水平,且在迁移任务中表现显著提升。结合风格奖励进行模仿时,能更接近专家行为。视频与代码见 https://sites.google.com/view/latent-action-priors。

原文摘要 · Abstract (English)

Deep Reinforcement Learning (DRL) enables robots to learn complex behaviors through interaction with the environment. However, due to the unrestricted nature of the learning algorithms, the resulting solutions are often brittle and appear unnatural. This is especially true for learning direct joint-level torque control, as inductive biases are difficult to integrate into the learning process. We propose an inductive bias for learning locomotion that is especially useful for torque control: latent actions learned from a small dataset of expert demonstrations. This prior allows the policy to directly leverage knowledge contained in the expert's actions and facilitates more efficient exploration. We observe that the agent is not restricted to the reward levels of the demonstration, and performance in transfer tasks is improved significantly. Latent action priors combined with style rewards for imitation lead to a closer replication of the expert's behavior. Videos and code are available at https://sites.google.com/view/latent-action-priors.

强化学习机器人控制动作先验模仿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。