用人类引导的残差强化学习,让机器人模型1.5小时就达95%成功率
HiL-ResRL: A Model-Agnostic Finetuning Adapter via Human-in-the-loop Residual Reinforcement Learning

- 以人类反馈为引导,训练残差策略修正模仿学习错误
- 仅1.5小时真实训练,机器人任务成功率超95%
- 无需改动模型结构,适配多种视觉语言动作模型
生成式模仿学习近年来显著推动了机器人操作领域的发展。然而,现有多数模型严重依赖行为克隆(BC),该范式存在误差累积和分布偏移问题,导致实际工业部署效果受限。为此,我们提出一种新型即插即用的微调流程,旨在促进视觉-语言-动作(VLA)模型在真实环境中的稳健部署。与当前受特定模型架构限制的强化学习微调方法不同,本框架具备模型无关性,可适配多种VLA模型。我们将VLA生成的动作视为统一接口,并在此基础上训练一个残差策略,用于纠正次优动作并缓解模仿学习中的分布偏移。同时,引入人机协同指导,确保探索安全并提升训练效率。我们在真实机器人环境中直接开展实验,结果表明:仅需1.5小时的真实世界在线强化学习训练,平均成功率即可超过95%。本工作为行为克隆模型在工业场景中的部署提供了切实可行的解决方案。
原文摘要 · Abstract (English)
Recent advancements in generative imitation learning have significantly propelled the field of robotic manipulation. However, the majority of existing models rely heavily on Behavior Cloning (BC), a paradigm that suffers from compounding errors and distributional shift. Consequently, the efficacy of these models in practical industrial deployments remains limited. To address these challenges, we introduce a novel, plug-and-play fine-tuning pipeline designed to facilitate the robust deployment of Vision-Language-Action (VLA) models in real-world environments. In contrast to contemporary reinforcement learning (RL) fine-tuning strategies, which are often constrained by specific model architectures, our proposed framework is model-agnostic and adaptable to a diverse range of VLA models. We conceptualize VLA-generated actions as a unified interface, upon which we train a residual policy. This policy is designed to rectify suboptimal actions and address the distributional shift inherent in imitation learning. Additionally, we incorporate human-in-the-loop guidance to ensure safe exploration and maximize training efficiency. We conduct experiments directly in real-world robotic settings. The results demonstrate that within only 1.5 hour of real-world online RL training, the average success rate exceeds 95% on real robots. Our work presents a practical solution for deploying behavior cloning models in industrial scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。