让机器人在仿真中学会感知环境差异,提升真实世界表现。
Can Context Bridge the Reality Gap? Sim-to-Real Transfer of Context-Aware Policies
- 用上下文感知机制让策略根据环境参数动态调整
- 在控制基准和真实机械臂任务中均优于传统方法
- 适合关注仿真到现实迁移的机器人学习研究者
仿真到现实的迁移仍是强化学习在机器人领域的一大挑战,因仿真与现实间存在环境动态差异,导致训练好的策略难以泛化。领域随机化(DR)通过在训练中引入广泛随机化的动态参数缓解此问题,但常伴随性能下降。现有方法通常训练对这些变化无感的策略,而本文探索是否可通过让策略依赖动态参数估计(即上下文)来改善迁移效果。为此,我们在基于DR的强化学习框架中集成上下文估计模块,并系统比较了多种主流监督策略。在标准控制基准和使用Franka Emika Panda机械臂的真实推物任务中评估结果表明,上下文感知策略在所有设置下均优于无上下文基线,但最优监督策略随任务而异。
原文摘要 · Abstract (English)
Sim-to-real transfer remains a major challenge in reinforcement learning (RL) for robotics, as policies trained in simulation often fail to generalize to the real world due to discrepancies in environment dynamics. Domain Randomization (DR) mitigates this issue by exposing the policy to a wide range of randomized dynamics during training, yet leading to a reduction in performance. While standard approaches typically train policies agnostic to these variations, we investigate whether sim-to-real transfer can be improved by conditioning the policy on an estimate of the dynamics parameters -- referred to as context. To this end, we integrate a context estimation module into a DR-based RL framework and systematically compare SOTA supervision strategies. We evaluate the resulting context-aware policies in both a canonical control benchmark and a real-world pushing task using a Franka Emika Panda robot. Results show that context-aware policies outperform the context-agnostic baseline across all settings, although the best supervision strategy depends on the task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。