用混合仿真生成真实机器人训练数据,降低仿真到现实的差距。
ComSim: Building Scalable Real-World Robot Data Generation via Compositional Simulation
- 结合经典仿真与神经仿真,生成带动作的真实视频对。
- 仅需少量真实数据即可扩展出大规模多样场景数据集。
- 适合需要大量真实感训练数据的机器人研究者使用。
近期基于大语言模型和世界模型的基础模型显著提升了机器人能力,使其能自主完成复杂任务。然而,获取大规模高质量机器人训练数据仍具挑战,因通常需大量人工参与,且难以覆盖多样化真实环境。为此,我们提出一种新型混合方法——组合仿真(Compositional Simulation),融合经典仿真与神经仿真,生成准确的动作-视频配对,同时保持真实世界一致性。该方法采用闭环真实-仿真-真实数据增强流程,利用少量真实数据生成覆盖更广真实场景的大规模多样化训练数据集。我们训练神经仿真器将经典仿真视频转换为真实世界表征,提升在真实环境中训练策略模型的准确性。通过大量实验验证,该方法显著缩小了仿真到现实的域差距,提升了真实策略模型训练的成功率。本方法为生成鲁棒训练数据及弥合仿真与真实机器人之间的差距提供了可扩展解决方案。
原文摘要 · Abstract (English)
Recent advancements in foundational models, such as large language models and world models, have greatly enhanced the capabilities of robotics, enabling robots to autonomously perform complex tasks. However, acquiring large-scale, high-quality training data for robotics remains a challenge, as it often requires substantial manual effort and is limited in its coverage of diverse real-world environments. To address this, we propose a novel hybrid approach called Compositional Simulation, which combines classical simulation and neural simulation to generate accurate action-video pairs while maintaining real-world consistency. Our approach utilizes a closed-loop real-sim-real data augmentation pipeline, leveraging a small amount of real-world data to generate diverse, large-scale training datasets that cover a broader spectrum of real-world scenarios. We train a neural simulator to transform classical simulation videos into real-world representations, improving the accuracy of policy models trained in real-world environments. Through extensive experiments, we demonstrate that our method significantly reduces the sim2real domain gap, resulting in higher success rates in real-world policy model training. Our approach offers a scalable solution for generating robust training data and bridging the gap between simulated and real-world robotics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。