用3D高斯泼溅实现每秒超十万步的逼真机器人仿真
GaussGym: An open-source real-to-sim framework for learning locomotion from pixels
- 将3D高斯泼溅作为渲染器嵌入物理引擎,实现高速仿真
- 在消费级显卡上达到10万+步/秒,视觉保真度高
- 支持海量真实场景数据,适合快速构建训练环境
我们提出一种新颖的摄影级机器人仿真方法,将3D高斯泼溅作为即插即用渲染器集成到向量化物理模拟器(如IsaacGym)中。该方法在消费级GPU上实现超过10万步/秒的惊人速度,同时保持高视觉保真度,并在多种任务中验证其有效性。我们进一步展示了其在仿真到现实迁移中的应用潜力。相比仅依赖深度感知,丰富的视觉语义显著提升了导航与决策能力,例如避开不良区域。此外,可轻松整合来自iPhone扫描、大型场景数据集(如GrandTour、ARKit)及生成视频模型(如Veo)的数千个环境,实现快速构建逼真训练世界。本工作打通了高吞吐仿真与高保真感知的壁垒,推动可扩展、泛化性强的机器人学习。所有代码与数据将开源,相关视频、代码和数据详见https://escontrela.me/gauss_gym/
原文摘要 · Abstract (English)
We present a novel approach for photorealistic robot simulation that integrates 3D Gaussian Splatting as a drop-in renderer within vectorized physics simulators such as IsaacGym. This enables unprecedented speed -- exceeding 100,000 steps per second on consumer GPUs -- while maintaining high visual fidelity, which we showcase across diverse tasks. We additionally demonstrate its applicability in a sim-to-real robotics setting. Beyond depth-based sensing, our results highlight how rich visual semantics improve navigation and decision-making, such as avoiding undesirable regions. We further showcase the ease of incorporating thousands of environments from iPhone scans, large-scale scene datasets (e.g., GrandTour, ARKit), and outputs from generative video models like Veo, enabling rapid creation of realistic training worlds. This work bridges high-throughput simulation and high-fidelity perception, advancing scalable and generalizable robot learning. All code and data will be open-sourced for the community to build upon. Videos, code, and data available at https://escontrela.me/gauss_gym/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。