构建高保真城市环境仿真,提升具身智能评估可靠性
Wanderland: Geometrically Grounded Simulation for Open-World Embodied AI
- 基于多传感器采集与几何重建,实现真实场景到仿真的精准映射
- 实验证明仅用图像的方案在新视角合成中表现差,几何精度直接影响导航性能
- 提供可复现的测试基准,适合导航、三维重建与视图合成模型评测
具身智能(如视觉导航)的可复现闭环评估仍是重大瓶颈。高保真仿真结合逼真传感器渲染与复杂开放世界中的几何交互,是潜在解决方案。尽管近期视频-3DGS方法简化了开放世界场景捕获,但其仍因视觉与几何的仿真-现实差距而不适合作为基准。为此,我们提出Wanderland,一个从真实到仿真的框架,具备多传感器采集、可靠重建、精确几何与鲁棒视图合成能力。通过该流程,我们构建了多样化的室内外城市场景数据集,并系统证明:仅依赖图像的流程扩展性差,几何质量影响新视角合成效果,这些因素均损害导航策略学习与评估可靠性。除了作为具身导航的可信测试平台外,其丰富的原始传感器数据还支持三维重建与新视角合成模型的评测。本工作为开放世界具身智能的可复现研究建立了新基础。项目网站:https://ai4ce.github.io/wanderland/
原文摘要 · Abstract (English)
Reproducible closed-loop evaluation remains a major bottleneck in Embodied AI such as visual navigation. A promising path forward is high-fidelity simulation that combines photorealistic sensor rendering with geometrically grounded interaction in complex, open-world urban environments. Although recent video-3DGS methods ease open-world scene capturing, they are still unsuitable for benchmarking due to large visual and geometric sim-to-real gaps. To address these challenges, we introduce Wanderland, a real-to-sim framework that features multi-sensor capture, reliable reconstruction, accurate geometry, and robust view synthesis. Using this pipeline, we curate a diverse dataset of indoor-outdoor urban scenes and systematically demonstrate how image-only pipelines scale poorly, how geometry quality impacts novel view synthesis, and how all of these adversely affect navigation policy learning and evaluation reliability. Beyond serving as a trusted testbed for embodied navigation, Wanderland's rich raw sensor data further allows benchmarking of 3D reconstruction and novel view synthesis models. Our work establishes a new foundation for reproducible research in open-world embodied AI. Project website is at https://ai4ce.github.io/wanderland/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。