arXiv:2503.10484cs.RO2025-03被引 3

用想象的过渡动作提升机器人行走效率与鲁棒性平衡

Learning Robotic Policy with Imagined Transition: Mitigating the Trade-off between Robustness and Optimality

  • 用理想环境下的最优策略和动态模型生成想象过渡
  • 训练更快,理想条件下跟踪误差降低,外部扰动下更稳定
  • 适合需要兼顾性能与抗干扰的机器人控制场景

现有四足机器人行走学习方法通常依赖大量领域随机化来缓解仿真到现实的差距并增强鲁棒性,通过在多种环境参数和传感器噪声下训练策略以应对不确定性。然而,理想条件下的最优表现与极端情况下的稳定性存在冲突,导致策略偏向保守,牺牲峰值性能。本文提出一种两阶段框架,通过融合想象过渡来缓解这一权衡。该框架将想象过渡(由理想设定下的最优策略与动态模型生成)作为示范输入,增强传统强化学习方法。实验表明,该方法显著减轻了领域随机化带来的负面影响,实现训练加速、分布内跟踪误差减小以及分布外鲁棒性提升。

原文摘要 · Abstract (English)

Existing quadrupedal locomotion learning paradigms usually rely on extensive domain randomization to alleviate the sim2real gap and enhance robustness. It trains policies with a wide range of environment parameters and sensor noises to perform reliably under uncertainty. However, since optimal performance under ideal conditions often conflicts with the need to handle worst-case scenarios, there is a trade-off between optimality and robustness. This trade-off forces the learned policy to prioritize stability in diverse and challenging conditions over efficiency and accuracy in ideal ones, leading to overly conservative behaviors that sacrifice peak performance. In this paper, we propose a two-stage framework that mitigates this trade-off by integrating policy learning with imagined transitions. This framework enhances the conventional reinforcement learning (RL) approach by incorporating imagined transitions as demonstrative inputs. These imagined transitions are derived from an optimal policy and a dynamics model operating within an idealized setting. Our findings indicate that this approach significantly mitigates the domain randomization-induced negative impact of existing RL algorithms. It leads to accelerated training, reduced tracking errors within the distribution, and enhanced robustness outside the distribution.

机器人控制强化学习鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。