arXiv:2503.08593cs.RO2025-03被引 6

用程序化生成模拟训练大模型,让机器人学会在真实世界推物

Proc4Gem: Foundation models for physical agency through procedural generation

  • 用程序化生成带物理接触的多样化仿真环境
  • 仅用仿真数据微调的Gemini模型可直接控制机器人推物
  • 适合对具身智能和仿真训练感兴趣的开发者

在机器人学习中,通常要么忽略环境语义,只关注整体控制;要么忽略接触动力学,只依赖视觉和语言。本文提出利用生成建模、逼真渲染和程序化生成技术,同时处理这两方面需求。通过在语义多样且接触丰富的仿真环境中生成精确物理的轨迹,我们将行为提炼为大型多模态模型,实现直接跨域到真实世界——即Proc4Gem系统。实验表明,仅在仿真数据上微调的Gemini模型,可通过自然语言指令控制四足机器人,在未见过的真实环境中用身体推动物体至未知目标。结果展示了用仿真赋予基础模型物理行为能力的巨大潜力。视频见官网:https://sites.google.com/view/proc4gem

原文摘要 · Abstract (English)

In robot learning, it is common to either ignore the environment semantics, focusing on tasks like whole-body control which only require reasoning about robot-environment contacts, or conversely to ignore contact dynamics, focusing on grounding high-level movement in vision and language. In this work, we show that advances in generative modeling, photorealistic rendering, and procedural generation allow us to tackle tasks requiring both. By generating contact-rich trajectories with accurate physics in semantically-diverse simulations, we can distill behaviors into large multimodal models that directly transfer to the real world: a system we call Proc4Gem. Specifically, we show that a foundation model, Gemini, fine-tuned on only simulation data, can be instructed in language to control a quadruped robot to push an object with its body to unseen targets in unseen real-world environments. Our real-world results demonstrate the promise of using simulation to imbue foundation models with physical agency. Videos can be found at our website: https://sites.google.com/view/proc4gem

具身智能仿真训练大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。