用大规模强化学习微调,让机器人在新环境新任务中表现大幅提升。
FLaRe: Achieving Masterful and Adaptive Robot Policies with Large-Scale Reinforcement Learning Fine-Tuning
- 基于预训练表示和梯度稳定技术,通过强化学习微调提升机器人泛化能力。
- 在未见环境中平均成功率79.5%,比现有最优方法提升23.6%(仿真)和30.7%(实机)。
- 仅用稀疏奖励即可快速适应新机器人形态和新技能,一天内完成微调。
近年来,机器人领域已尝试通过大规模多任务行为克隆构建通用机器人策略。然而,直接部署这些策略表现不佳,难以应对未见状态与任务。本文提出FLaRe——一种大规模强化学习微调框架,整合鲁棒预训练表示、大规模训练与梯度稳定技术。该方法将预训练策略对齐至任务完成目标,在先前演示及全新任务与机器人形态上均达到当前最优(SoTA)性能。具体而言,在一系列长时程移动操作任务中,FLaRe在未见环境中的平均成功率达79.5%,相比之前最优方法在仿真中提升23.6%,在真实机器人上提升30.7%。仅使用稀疏奖励,即可实现对预训练数据外新能力的泛化,且几乎无需人工干预。此外,我们展示了在不到一天时间内即可快速适配新机器人形态与行为的能力。视频详见项目网站:https://robot-flare.github.io/
原文摘要 · Abstract (English)
In recent years, the Robotics field has initiated several efforts toward building generalist robot policies through large-scale multi-task Behavior Cloning. However, direct deployments of these policies have led to unsatisfactory performance, where the policy struggles with unseen states and tasks. How can we break through the performance plateau of these models and elevate their capabilities to new heights? In this paper, we propose FLaRe, a large-scale Reinforcement Learning fine-tuning framework that integrates robust pre-trained representations, large-scale training, and gradient stabilization techniques. Our method aligns pre-trained policies towards task completion, achieving state-of-the-art (SoTA) performance both on previously demonstrated and on entirely novel tasks and embodiments. Specifically, on a set of long-horizon mobile manipulation tasks, FLaRe achieves an average success rate of 79.5% in unseen environments, with absolute improvements of +23.6% in simulation and +30.7% on real robots over prior SoTA methods. By utilizing only sparse rewards, our approach can enable generalizing to new capabilities beyond the pretraining data with minimal human effort. Moreover, we demonstrate rapid adaptation to new embodiments and behaviors with less than a day of fine-tuning. Videos can be found on the project website at https://robot-flare.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。