arXiv:2606.17520cs.ROcs.CV2026-06

用全景相机快速重建高保真仿真环境,缩小模拟与现实差距。

GASE: Gaussian Splatting-Based Automated System for Reconstructing Embodied-Simulation Environments

论文配图:GASE: Gaussian Splatting-Based Automated System for Reconstructing Embodied-Simulation Environments
图 1 · 摘自论文原文
  • 通过多视角视频自动扫描环境,结合相机位姿提取前景物体
  • 分割准确率提升超10%,背景修复质量达当前最优水平
  • 适合机器人学习中的场景构建,尤其擅长抓取与导航任务

在真实世界训练具身智能体需要专业人员和昂贵硬件。仿真环境提供了大规模、低成本的数据增强替代方案。因此,以最小化仿真到现实的差距,快速构建高保真仿真场景成为机器人学习的关键目标。尽管基于重建的方法视觉质量优异,但现有流程存在数据采集效率低和前景物体提取效果差的问题。为此,我们提出GASE——一种高度自动化的仿真场景构建系统。GASE利用全景相机阵列的多视角视频流实现快速环境扫描。为确保资产生成质量,其流水线引入基于相机位姿的策略,在2D域中鲁棒提取跨帧物体,随后进行高保真场景补全。前景物体与静态背景分别独立重建,并无缝导入物理仿真器用于策略训练。大量实验表明,GASE在分割准确率上比现有基于3D高斯的方法提升超过10%,同时达到最先进的补全质量。此外,真实机器人部署在操作与导航任务中,性能差距小于10%(相比纯真实数据训练的策略)。这些结果证实GASE是弥合仿真到现实差距的高效且有效方案。代码将开源。

原文摘要 · Abstract (English)

Training embodied agents in the real world requires skilled operators and expensive hardware. Simulation environments offer a compelling alternative by enabling large-scale, cost-effective data augmentation. Consequently, rapidly constructing high-fidelity simulation scenes with a minimal sim-to-real gap has become a critical objective in robot learning. While reconstruction-based methods provide superior visual quality, current workflows are hindered by inefficient data acquisition and subpar foreground object extraction. We thus propose GASE, a highly automated system for simulation scene construction. GASE leverages multi-view video streams from panoramic camera arrays to enable rapid environment scanning. To ensure high-quality asset generation, our pipeline introduces a camera-pose-based strategy that robustly extracts objects across frames in the 2D domain, followed by high-fidelity scene inpainting. Foreground objects and the static background are then reconstructed independently and seamlessly imported into physics simulators for policy training. Extensive experiments demonstrate that GASE outperforms existing 3D Gaussian-based methods in segmentation accuracy by over 10\% while achieving state-of-the-art inpainting quality. Furthermore, real-robot deployments across manipulation and navigation tasks maintains a performance gap of less than 10\% compared to policies trained purely on real-world data. These results confirm that GASE provides an efficient and highly effective solution for bridging the sim-to-real gap. Code will be released.

仿真构建高斯溅射机器人学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。