用真实场景生成逼真仿真环境,让机器人视觉导航更靠谱
VR-Robo: A Real-to-Sim-to-Real Framework for Visual Robot Navigation and Locomotion
- 用多视角图像重建3D高保真场景,生成可交互的数字孪生环境
- 在仿真中训练的视觉导航策略可直接迁移到真实世界,无需额外调参
- 适合需要快速适应新复杂环境的家居与工厂机器人应用
近期足式机器人运动的成功得益于强化学习与物理仿真器的结合。然而,这些策略在真实环境中部署时常因仿真到现实的差距而失效,因为仿真器通常无法还原真实的视觉效果和复杂的现实几何结构。此外,缺乏真实感的视觉渲染限制了策略在依赖RGB感知的高层任务(如自居视角导航)中的表现。本文提出一种真实到仿真再到真实的框架,通过基于3D高斯泼溅(3DGS)的多视角图像场景重建,生成支持自居视角视觉感知和基于网格物理交互的高保真数字孪生仿真环境。为验证其有效性,我们在仿真中训练强化学习策略完成视觉目标追踪任务。大量实验表明,该框架实现了仅依赖RGB的仿真到现实策略迁移。同时,该框架支持机器人策略在复杂新环境中的快速适应与高效探索,展现出在家庭和工厂场景中的应用潜力。
原文摘要 · Abstract (English)
Recent success in legged robot locomotion is attributed to the integration of reinforcement learning and physical simulators. However, these policies often encounter challenges when deployed in real-world environments due to sim-to-real gaps, as simulators typically fail to replicate visual realism and complex real-world geometry. Moreover, the lack of realistic visual rendering limits the ability of these policies to support high-level tasks requiring RGB-based perception like ego-centric navigation. This paper presents a Real-to-Sim-to-Real framework that generates photorealistic and physically interactive "digital twin" simulation environments for visual navigation and locomotion learning. Our approach leverages 3D Gaussian Splatting (3DGS) based scene reconstruction from multi-view images and integrates these environments into simulations that support ego-centric visual perception and mesh-based physical interactions. To demonstrate its effectiveness, we train a reinforcement learning policy within the simulator to perform a visual goal-tracking task. Extensive experiments show that our framework achieves RGB-only sim-to-real policy transfer. Additionally, our framework facilitates the rapid adaptation of robot policies with effective exploration capability in complex new environments, highlighting its potential for applications in households and factories.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。