arXiv:2602.12628cs.RO2026-02被引 3

用强化学习让仿真与真实数据协同训练,提升机器人任务成功率和泛化能力。

Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models

  • 先用真实与仿真数据微调,再在仿真中通过强化学习优化策略
  • 相比纯真实数据训练,任务成功率提升24%(OpenVLA)和20%(π₀.₅)
  • 适合希望低成本提升机器人部署性能的研究者与工程师

仿真为视觉-语言-动作(VLA)模型训练提供了可扩展且低成本的途径,减少对昂贵真实机器人演示的依赖。然而,现有仿真-真实协同训练方法多采用监督微调(SFT),将仿真视为静态演示源,未充分利用大规模闭环交互。为此,本文提出基于强化学习的仿真-真实协同训练(RL-Co)框架,通过交互式仿真保留真实能力。方法分为两阶段:首先在真实与仿真数据混合集上进行SFT暖启动;随后在仿真中以强化学习微调策略,并加入真实数据的辅助监督损失,防止灾难性遗忘。我们在四个真实桌面操作任务上,使用OpenVLA和π₀.₅两种代表性VLA架构进行评估,结果表明,相比仅用真实数据微调和SFT协同训练,本方法在真实世界中成功率分别提升24%(OpenVLA)和20%(π₀.₅)。此外,该方法还显著增强对未见任务变体的泛化能力,并大幅提升真实数据效率,为利用仿真提升真实机器人部署提供了一条实用且可扩展的路径。

原文摘要 · Abstract (English)

Simulation offers a scalable and low-cost way to enrich vision-language-action (VLA) training, reducing reliance on expensive real-robot demonstrations. However, most sim-real co-training methods rely on supervised fine-tuning (SFT), which treats simulation as a static source of demonstrations and does not exploit large-scale closed-loop interaction. Consequently, real-world gains and generalization are often limited. In this paper, we propose an RL-based sim-real Co-training (RL-Co) framework that leverages interactive simulation while preserving real-world capabilities. Our method follows a generic two-stage design: we first warm-start the policy with SFT on a mixture of real and simulated demonstrations, then fine-tune it with reinforcement learning in simulation while adding an auxiliary supervised loss on real-world data to anchor the policy and mitigate catastrophic forgetting. We evaluate our framework on four real-world tabletop manipulation tasks using two representative VLA architectures, OpenVLA and $π_{0.5}$, and observe consistent improvements over real-only fine-tuning and SFT-based co-training, including +24% real-world success on OpenVLA and +20% on $π_{0.5}$. Beyond higher success rates, RL co-training yields stronger generalization to unseen task variations and substantially improved real-world data efficiency, providing a practical and scalable pathway for leveraging simulation to enhance real-robot deployment.

强化学习机器人仿真训练VLA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。