arXiv:2607.06699cs.RO2026-07被引 2

用一张照片生成可交互的仿真环境,提升机器人学习泛化能力

RoboSnap: One-Shot Real-to-Sim Scene Generation for Generalizable Robot Learning and Evaluation

论文配图:RoboSnap: One-Shot Real-to-Sim Scene Generation for Generalizable Robot Learning and Evaluation
图 1 · 摘自论文原文
  • 分层设计:关键交互区优化物理稳定性,背景用3D高斯溅射保真视觉
  • 在真实场景上实现可靠轨迹重放,支持任务定制数据生成与评估
  • 适合需要高效仿真环境的机器人研究者,尤其关注真实到仿真迁移

将真实世界场景重建为可交互的仿真环境,有助于提升机器人学习的泛化能力并实现策略评估的可重复性。然而,构建既物理稳定又视觉逼真的场景仍耗时且成本高昂。本文提出RoboSnap,一种从单张RGB图像生成仿真就绪场景的实时到仿真框架。核心思想是分层设计:将物理交互关键区域与周围视觉上下文分离处理——碰撞感知的前景资产经过优化以确保机器人交互的稳定性,而背景则通过3D高斯溅射技术在新视角下保持视觉忠实度。在DROID场景和真实机器人任务上的实验表明,RoboSnap可在重建场景中实现可靠的轨迹重放,支持面向特定任务的合成数据生成用于策略训练,并在策略评估中展现出有意义的仿真-真实相关性。为进一步支持实时到仿真研究,我们构建了配套数据集DROID-Sim,基于DROID中的564个真实场景。大量实验表明,实时到仿真方法的价值不仅在于高保真视觉重建,更在于将真实环境转化为可复用的机器人学习与评估基础设施。

原文摘要 · Abstract (English)

Recovering real-world scenes as interactive simulation environments can enable generalizable robot learning and reproducible policy evaluation. However, constructing scenes that are both physically stable and visually faithful remains slow and expensive. In this work, we present RoboSnap, a real-to-sim framework that turns a single RGB image into a simulation-ready scene. The key idea is a layered design that separates the physics-critical interaction area from the surrounding visual context: collision-aware foreground assets are refined for stable robot interaction, while a 3D Gaussian splatting visual layer preserves faithful background appearance under novel views. Experiments on DROID scenes and real-robot tasks show that RoboSnap achieves reliable trajectory replay in the recovered scenes, supports task-specific synthetic data generation for policy training, and yields meaningful sim-real correlation for policy evaluation. To further support real-to-sim research, we introduce DROID-Sim, a real-to-sim companion dataset constructed from 564 real-world scenes in DROID. Extensive experiments suggest that the value of real-to-sim methods lies not only in high-fidelity visual reconstruction, but in turning real environments into reusable infrastructure for robot learning and evaluation.

真实到仿真机器人学习3D重建仿真生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。