arXiv:2606.31101cs.RO2026-06中稿 · CVPR

用合成数据训练的视觉动作模型,零样本部署到真实机械臂上成功率达35%。

Efficient Sim-to-Real Transfer of World-Action Models from Synthetic Priors

论文配图:Efficient Sim-to-Real Transfer of World-Action Models from Synthetic Priors
图 1 · 摘自论文原文
  • 基于视频扩散模型构建世界-动作模型,利用合成数据训练。
  • 仅用约800条模拟演示,零样本在真实机械臂上实现35%成功率。
  • 首次实现世界-动作模型从仿真到真实的零样本迁移,适合机器人学习研究者。

弥合仿真到现实的差距是部署学习到的操控策略的核心挑战。模拟到现实的学习具有吸引力,因为它可以用可扩展的合成数据替代昂贵的真实机器人演示,但此前尚未证明世界-动作模型能从仿真迁移到真实机器人操控。我们研究了是否可以从合成先验中训练世界-动作模型,并在真实世界中零样本部署。为此,我们基于适应于视觉-运动控制的Cosmos Policy(一种视频扩散模型),构建了经过广泛领域随机化的仿真环境,并使用AnyTask运动规划管道生成示范。我们在物体抓取、抽屉开启和拾取放置任务上进行评估,每个任务使用约800条合成示范,且无需真实示范。当零样本部署到Franka机器人时,该策略平均成功率达到35%。据我们所知,这是首个成功实现世界-动作模型在机器人操控中从仿真到真实环境的零样本迁移。

原文摘要 · Abstract (English)

Bridging the sim-to-real gap is a core challenge in deploying learned manipulation policies. Sim-to-real learning is attractive because it can replace expensive real robot demonstrations with scalable synthetic data, yet world-action models have not previously been shown to transfer from simulation to real robotic manipulation. We study whether a world-action model can be trained from synthetic priors and deployed zero-shot in the real world. To this end, we build upon Cosmos Policy, a video diffusion model adapted for visuomotor control. We construct simulation environments with extensive domain randomization and generate demonstrations using the AnyTask motion planning pipeline. We evaluate our approach across object lifting, drawer opening, and pick-and-place tasks using ${\sim}800$ synthetic demonstrations per task and no real demonstrations. When deployed zero-shot on a Franka Robot, our policy attains a 35\% average success rate. To our knowledge, this represents the first successful sim-to-real transfer of a world-action model for robotic manipulation.

Sim-to-Real视觉动作模型零样本机器人学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。