用物体姿态提升视觉语言动作模型的实机表现,零样本迁移成功率达76%。
Object-Centric Residual RL for Zero-Shot Sim-to-Real VLA Enhancement

- 以物体姿态为输入,构建可跨域迁移的残差强化学习框架。
- 在真实机器人上零样本提升成功率至76%,比原模型提高34个百分点。
- 无需真实世界训练,适合希望低成本增强机械臂泛化能力的研究者。
视觉-语言-动作(VLA)模型虽能泛化于多样操作任务,但基于模仿学习的策略在精细物理交互中仍易因执行误差累积而失效。纯仿真训练的强化学习能否实现零样本提升真实世界VLA鲁棒性?残差强化学习通过在冻结的VLA之上学习修正策略提供自然框架,但现有方法面临仿真到现实的根本困境:依赖状态信息的方法需有损蒸馏部署;基于图像的方法受视觉域差距影响;真实世界训练成本高且不安全。本文提出一种面向物体的残差强化学习框架,利用物体位姿精炼VLA动作,构建紧凑观测空间,确保仿真与现实间一致迁移。为对齐域间差异,我们在仿真中重播相同遥操作示范,训练真实VLA的仿真对应版本。残差策略仅在仿真中训练,注入位姿噪声和丢弃,并实现零样本迁移到真实机器人。在真实Franka Research 3(FR3)机器人上,五项操作任务的成功率从42%提升至76%,改进后的轨迹还可用于重新训练基础VLA实现自进化,无需额外遥操作。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models can generalize across diverse manipulation tasks, but their imitation-learning-based policies remain brittle in precise physical interactions due to compounding execution errors; Can a reinforcement learning policy trained purely in simulation improve the robustness of real-world VLAs zero-shot? Residual RL, which learns a corrective policy on top of a frozen VLA, offers a natural framework, but existing approaches face a fundamental sim-to-real dilemma: privileged-state methods require lossy distillation for deployment; image-based methods suffer from the visual domain gap; and real-world RL is costly and unsafe. We propose an object-centric residual RL framework that refines VLA actions using object poses, enabling a compact observation space that transfers consistently between simulation and reality. To align the two domains, we additionally replay the same teleoperation demonstrations in simulation to train a sim counterpart of the real-world VLA. The residual RL policy is trained only in simulation with pose noise injection and dropout, and transfers zero-shot to the real robot. Across five manipulation tasks on a real Franka Research 3 (FR3) robot, our method improves the success rate from 42% to 76% zero-shot, and the improved rollouts can be further reused to retrain the base VLA for self-improvement without additional teleoperation. Project page: https://www.microsoft.com/en-us/research/articles/object-centric-residual-rl/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。