用视频自动生成可编辑的仿真场景,提升机器人策略泛化能力
SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation

- 从视频自动构建可编辑的数字孪生场景
- 7个任务5种模型下仿真预测真实性能相关性达0.911
- 生成场景变体使真实任务成功率提升最高40%
在真实世界中训练和评估机器人策略成本高且难以扩展。我们提出SimFoundry,一个模块化、自动化的系统,能够从视频零样本生成可用于仿真的数字孪生。SimFoundry支持物体、场景和任务的编辑,可自动生成多样化的数字表亲:保留功能特性的现实场景变体。在包含多步操作、刚性物体交互和双臂协作的挑战性真实任务中,基于SimFoundry数据训练的策略可实现零样本迁移。其生成的数字表亲(原场景、物体和任务的变体)有助于策略泛化至新的真实条件。在7个操作任务和5种策略架构上,仿真评估与真实性能的平均皮尔逊相关系数为0.911,最大排名偏差为0.018。在零样本评估中,使用物体、场景和任务变体训练的策略,在真实世界中的任务成功率分别平均提升17%、21%和40%。
原文摘要 · Abstract (English)
Training and evaluating robot policies in the real world is costly and difficult to scale. We introduce SimFoundry, a modular and automated system for zero-shot real-to-sim scene construction from a video. SimFoundry generates sim-ready digital twins and supports object, scene, and task editing, enabling the automated generation of diverse digital cousins: affordance-preserving variations of reconstructed real-world scenes. Policies trained on SimFoundry data transfer zero-shot to challenging real tasks involving multi-step manipulation, articulated object interaction, and bimanual interaction, and its digital cousins (variations of the original scene, objects, and tasks) facilitate generalization to new real-world conditions. Across 7 manipulation tasks and 5 policy architectures, SimFoundry simulation evaluations strongly predict real-world performance, with mean Pearson correlation 0.911 and mean maximum ranking violation 0.018. When evaluating sim-trained policies zero-shot in the real world, policies trained with object, scene, and task cousins in simulation show average task success rate improvements of 17%, 21%, and 40%, respectively. Additional details at https://research.nvidia.com/labs/gear/simfoundry/ .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。