arXiv:2501.06693cs.CVcs.RO2025-01CVPR被引 46

用单目视频生成可交互的逼真城市模拟环境,提升机器人导航性能。

Vid2Sim: Realistic and Interactive Simulation from Video for Urban Navigation

  • 基于单目视频重建三维场景,构建可物理交互的仿真环境。
  • 在数字孪生与真实世界中,导航成功率分别提升31.2%和68.3%。
  • 适合需要高保真仿真训练的自动驾驶与城市机器人研究者。

仿真中的现实差距长期制约机器人学习在真实世界的部署。以往方法主要依赖领域随机化与系统识别,但受限于仿真引擎的固有约束。本文提出Vid2Sim框架,通过高效低成本的从现实到仿真(real2sim)管道,实现神经三维场景重建与仿真。输入单目视频,即可生成具有视觉真实感且可物理交互的三维仿真环境,支持复杂城市环境中视觉导航智能体的强化学习。大量实验表明,与以往仿真方法相比,该方法在数字孪生与真实世界中导航成功率分别提升31.2%和68.3%。

原文摘要 · Abstract (English)

Sim-to-real gap has long posed a significant challenge for robot learning in simulation, preventing the deployment of learned models in the real world. Previous work has primarily focused on domain randomization and system identification to mitigate this gap. However, these methods are often limited by the inherent constraints of the simulation and graphics engines. In this work, we propose Vid2Sim, a novel framework that effectively bridges the sim2real gap through a scalable and cost-efficient real2sim pipeline for neural 3D scene reconstruction and simulation. Given a monocular video as input, Vid2Sim can generate photorealistic and physically interactable 3D simulation environments to enable the reinforcement learning of visual navigation agents in complex urban environments. Extensive experiments demonstrate that Vid2Sim significantly improves the performance of urban navigation in the digital twins and real world by 31.2% and 68.3% in success rate compared with agents trained with prior simulation methods.

城市导航仿真生成视觉导航三维重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。