用图像生成可交互的3D环境,让机器人导航训练更高效真实。
Image2Sim: Scaling Embodied Navigation via Generative Neural Simulator

- 从带位姿的RGB-D图像生成3D特征高斯场景,一步完成
- 构建近2万张可交互场景,合成超1000万条导航样本
- 训练模型零样本迁移至真实世界,适合做机器人导航研究
具身导航旨在构建能理解多模态目标、在三维空间中推理并可靠到达目的地的智能体。然而,进展受限于缺乏可扩展、高保真且物理真实的交互环境。真实扫描数据集虽视觉逼真,但规模有限;合成模拟器易扩展,却存在显著的仿真到现实差距。我们提出Image2Sim,一个实时神经模拟框架,可从带位姿的RGB-D图像序列构建高质量交互环境。核心思想是将3D空间定位与逼真观测生成解耦。场景构建采用前馈特征高斯模型,单次处理即可将带位姿的RGB-D观测转化为3D特征高斯表示。渲染方面,提出几何感知的一步像素流模型,将稀疏噪声的高斯投影转换为高质量全景RGB-D观测。Image2Sim还可作为全自动具身数据引擎,大规模生成高保真观测、可执行动作和多样化导航指令。它将大量视频与图像转换为近20,000个交互场景,并合成超过1000万条导航训练样本。在这些神经环境中训练的导航模型在主要基准上表现显著提升,并能有效零样本迁移到真实世界。结果表明,可扩展的神经模拟可作为具身导航规模化训练的实用基础。
原文摘要 · Abstract (English)
Embodied navigation aims to build agents that interpret multimodal goals, reason in 3D space, and reach target destinations reliably in the real world. However, progress remains constrained by the lack of scalable, high-fidelity, and physically grounded interactive environments. Although real-world scanned datasets offer visual realism, they are limited by scale. In contrast, synthetic simulators scale more easily but often exhibit large sim-to-real gaps. We introduce Image2Sim, a real-time neural simulation framework that constructs high-quality interactive environments from posed RGB-D image sequences. The central idea is to decouple 3D spatial anchoring from photorealistic observation synthesis. For scene construction, Image2Sim uses a feed-forward feature Gaussian model that lifts posed RGB-D observations into a 3D feature-Gaussian representation in a single pass. For rendering, we propose a Geometry-Aware One-Step Pixel Flow model that transforms sparse and noisy Gaussian projections into high-quality panoramic RGB-D observations. Image2Sim also serves as a fully automated embodied data engine that generates high-fidelity observations, executable actions, and diverse navigation instructions at scale. It converts large collections of videos and images into nearly 20K interactive scenes and synthesizes more than 10 million navigation training samples. Navigation models trained entirely in these neural environments achieve strong improvements on major benchmarks and transfer effectively to real-world zero-shot settings. These results suggest that scalable neural simulation can serve as a practical training substrate for embodied navigation at scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。