arXiv:2503.14526cs.CVcs.GR2025-03被引 32

用仿真生成真实感机器人视频,低成本提升机器人学习效果

ReBot: Scaling Robot Learning with Real-to-Sim-to-Real Robotic Video Synthesis

  • 真实轨迹进仿真,自动扩增物体多样性
  • 合成视频保持物理真实与时间连贯性,性能提升显著
  • 全自动流程适配新任务,适合部署到新机器人

视觉-语言-动作(VLA)模型通过直接在真实机器人数据集(如Open X-Embodiment)上训练策略展现出巨大潜力。然而,真实数据采集成本高昂,制约了数据规模扩展,进而限制了VLA的泛化能力。本文提出ReBot,一种全新的真实→仿真→真实机器人视频生成方法,用于扩展真实数据集并适应VLA模型至目标领域,解决机器人操作中的最后一公里部署难题。具体而言,ReBot在仿真中重放真实机器人轨迹以多样化操控物体(真实→仿真),并将仿真运动与修复的真实背景结合,合成物理逼真且时间一致的机器人视频(仿真→真实)。该方法具备三大优势:1)利用真实数据最小化仿真到现实的差距;2)发挥仿真的可扩展性;3)可全自动将预训练VLA模型泛化至新领域。大量仿真与真实环境实验表明,ReBot显著提升了VLA的性能与鲁棒性。例如,在SimplerEnv中使用WidowX机器人时,ReBot使Octo的域内性能提升7.2%,OpenVLA提升21.8%;域外泛化能力分别提升19.9%和9.4%。在Franka机器人的真实测试中,Octo成功率提升17%,OpenVLA提升20%。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models present a promising paradigm by training policies directly on real robot datasets like Open X-Embodiment. However, the high cost of real-world data collection hinders further data scaling, thereby restricting the generalizability of VLAs. In this paper, we introduce ReBot, a novel real-to-sim-to-real approach for scaling real robot datasets and adapting VLA models to target domains, which is the last-mile deployment challenge in robot manipulation. Specifically, ReBot replays real-world robot trajectories in simulation to diversify manipulated objects (real-to-sim), and integrates the simulated movements with inpainted real-world background to synthesize physically realistic and temporally consistent robot videos (sim-to-real). Our approach has several advantages: 1) it enjoys the benefit of real data to minimize the sim-to-real gap; 2) it leverages the scalability of simulation; and 3) it can generalize a pretrained VLA to a target domain with fully automated data pipelines. Extensive experiments in both simulation and real-world environments show that ReBot significantly enhances the performance and robustness of VLAs. For example, in SimplerEnv with the WidowX robot, ReBot improved the in-domain performance of Octo by 7.2% and OpenVLA by 21.8%, and out-of-domain generalization by 19.9% and 9.4%, respectively. For real-world evaluation with a Franka robot, ReBot increased the success rates of Octo by 17% and OpenVLA by 20%. More information can be found at: https://yuffish.github.io/rebot/

机器人学习视频生成仿真VLA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。