用真实场景生成高保真仿真环境,提升机器人学习泛化能力
From Seeing to Simulating: Generative High-Fidelity Simulation with Digital Cousins for Generalizable Robot Learning and Evaluation

- 从真实全景图生成高保真仿真场景,支持语义与几何编辑
- 生成多样孪生场景,使机器人在未知环境中表现显著提升
- 适用于复杂布局长时序导航,适合需要泛化能力的机器人训练
在真实环境中学习鲁棒机器人策略需要多样化数据增强,但实物采集和环境重配置成本高昂。为此,我们提出一种生成式框架,实现从真实全景图到高保真仿真场景的映射,并通过语义与几何编辑合成多样化孪生场景。结合高质量物理引擎和真实资产,生成场景支持交互操作任务。同时,采用多房间拼接技术构建一致性大规模环境,用于复杂布局下的长时序导航。实验表明,该平台具有强模拟到现实的相关性,且大规模数据生成显著提升对未见场景和物体变化的泛化能力,验证了数字孪生在可泛化机器人学习与评估中的有效性。
原文摘要 · Abstract (English)
Learning robust robot policies in real-world environments requires diverse data augmentation, yet scaling real-world data collection is costly due to the need for acquiring physical assets and reconfiguring environments. Therefore, augmenting real-world scenes into simulation has become a practical augmentation for efficient learning and evaluation. We present a generative framework that establishes a generative real-to-sim mapping from real-world panoramas to high-fidelity simulation scenes, and further synthesize diverse cousin scenes via semantic and geometric editing. Combined with high-quality physics engines and realistic assets, the generated scenes support interactive manipulation tasks. Additionally, we incorporate multi-room stitching to construct consistent large-scale environments for long-horizon navigation across complex layouts. Experiments demonstrate a strong sim-to-real correlation validating our platform's fidelity, and show that extensively scaling up data generation leads to significantly better generalization to unseen scene and object variations, demonstrating the effectiveness of Digital Cousins for generalizable robot learning and evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。