arXiv:2603.16861cs.RO2026-03被引 6

用180万条仿真数据训练机器人,零样本直接在真实世界完成抓取和操作。

MolmoB0T: Large-Scale Simulation Enables Zero-Shot Manipulation

  • 通过大规模仿真生成多样化任务数据,训练视觉语言模型实现零样本迁移。
  • 在真实桌面场景中成功率达79.2%,远超对比模型的39.2%。
  • 开源全流程工具链,适合边缘部署与强化学习微调,适合工业应用。

当前主流观点认为,仅靠仿真不足以实现有效的真实世界迁移,通常需真实数据或任务微调来弥合差距。我们挑战这一假设:当仿真数据足够大规模且多样化时,零样本迁移到真实世界不仅可行,而且高效,适用于静态与移动操作。我们提出MolmoBot-Engine,一个全开源的跨机器人、任务与环境的程序化数据生成管道,构建于MolmoSpaces中。基于此,我们发布了包含180万条专家轨迹的MolmoBot-Data数据集,涵盖刚性物体操作与拾放任务。训练了三类策略:基于Molmo2的多帧视觉语言模型MolmoBot,采用流匹配动作头;复现$π_0$架构的MolmoBot-Pi0;以及轻量级适配边缘部署与强化学习微调的MolmoBot-SPOC。我们在Franka FR3与Rainbow Robotics RB-Y1两款平台上评估,涵盖桌面操作、开门、抽屉与柜体交互及移动拾放任务。无需任何真实世界微调,策略即可在未见物体与环境中成功执行。在桌面拾放任务中,MolmoBot在4个真实场景下成功率高达79.2%,显著优于$π_{0.5}$的39.2%。结果表明,结合程序化环境生成与多样刚性资产,可训练出广泛泛化的稳健操作策略。

原文摘要 · Abstract (English)

A prevailing view in robot learning is that simulation alone is not enough; effective sim-to-real transfer is widely believed to require at least some real-world data collection or task-specific fine-tuning to bridge the gap between simulated and physical environments. We challenge that assumption. With sufficiently large-scale and diverse simulated synthetic training data, we show that zero-shot transfer to the real world is not only possible, but effective for both static and mobile manipulation. We introduce MolmoBot-Engine, a fully open-source pipeline for procedural data generation across robots, tasks, and diverse simulated environments in MolmoSpaces. With it, we release MolmoBot-Data, a dataset of 1.8 million expert trajectories for articulated object manipulation and pick-and-place tasks. We train three policy classes: MolmoBot, a Molmo2-based multi-frame vision-language model with a flow-matching action head; MolmoBot-Pi0, which replicates the $π_0$ architecture to enable direct comparison; and MolmoBot-SPOC, a lightweight policy suitable for edge deployment and amenable to RL fine-tuning. We evaluate on two robotic platforms: the Franka FR3 for tabletop manipulation tasks and the Rainbow Robotics RB-Y1 mobile manipulator for door opening, drawer manipulation, cabinet interaction, and mobile pick-and-place. Without any real-world fine-tuning, our policies achieve zero-shot transfer to unseen objects and environments. On tabletop pick-and-place, MolmoBot achieves a success rate of 79.2% in real world evaluations across 4 settings, outperforming $π_{0.5}$ at 39.2%. Our results demonstrate that procedural environment generation combined with diverse articulated assets can produce robust manipulation policies that generalize broadly to the real world. Technical website: https://allenai.github.io/MolmoBot

机器人仿真迁移零样本视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。