用仿真+真机双源数据,高效优化机器人策略。
Simulation-Aided Policy Tuning for Black-Box Robot Learning
- 利用仿真与真实机器人数据联合建模策略改进
- 仅需少量真实交互即可实现高概率性能提升
- 适合数据稀缺的机器人快速适应新任务
机器人如何在少量数据下学习并适应新任务?系统化探索和仿真对高效机器人学习至关重要。本文提出一种面向黑箱策略搜索的新算法,专注于数据高效的策略优化。该算法直接在真实机器人上运行,将仿真视为加速学习的额外信息源。核心是构建一个概率模型,不仅通过机器人实验,还结合仿真数据来学习策略参数与学习目标之间的依赖关系。这显著减少了与真实机器人的交互次数。基于该模型,每次策略更新都能以高概率保证性能提升,从而实现快速、目标导向的学习。我们在模拟微调任务上评估了该算法,并展示了双信息源优化方法的数据效率。在真实机器人实验中,借助一个不完美的仿真器,成功实现了机械臂的快速且成功的任务学习。
原文摘要 · Abstract (English)
How can robots learn and adapt to new tasks and situations with little data? Systematic exploration and simulation are crucial tools for efficient robot learning. We present a novel black-box policy search algorithm focused on data-efficient policy improvements. The algorithm learns directly on the robot and treats simulation as an additional information source to speed up the learning process. At the core of the algorithm, a probabilistic model learns the dependence of the policy parameters and the robot learning objective not only by performing experiments on the robot, but also by leveraging data from a simulator. This substantially reduces interaction time with the robot. Using this model, we can guarantee improvements with high probability for each policy update, thereby facilitating fast, goal-oriented learning. We evaluate our algorithm on simulated fine-tuning tasks and demonstrate the data-efficiency of the proposed dual-information source optimization algorithm. In a real robot learning experiment, we show fast and successful task learning on a robot manipulator with the aid of an imperfect simulator.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。