RealMirror让机器人在仿真中训练后直接用在真实世界,无需调参。
RealMirror: A Comprehensive, Open-Source Vision-Language-Action Platform for Embodied AI
- 用生成模型和3D高斯点云重建真实环境,实现零样本仿真到现实迁移。
- 构建了包含多场景、长轨迹的开源基准,支持不同模型公平对比。
- 全链路开源,无需真机器人即可完成从数据到推理的全流程研究。
具身智能中的视觉-语言-动作(VLA)系统面临数据成本高、缺乏标准评估基准以及仿真与现实差距大的挑战。为此,我们提出RealMirror——一个综合性、开源的具身智能VLA平台。该平台构建了高效低成本的数据采集、模型训练与推理系统,支持无需真实机器人的端到端研究。为促进模型迭代与公平比较,我们设计了专用于人形机器人的VLA基准,涵盖多个任务场景、丰富轨迹数据及多种主流模型。通过融合生成模型与3D高斯溅射技术,实现逼真环境与机器人建模,成功验证零样本仿真到现实(Sim2Real)迁移:仅在仿真中训练的模型可直接在真实机器人上执行任务,无需微调。综上,通过整合核心组件,RealMirror为人类机器人VLA模型研发提供强大框架,显著加速进展。
原文摘要 · Abstract (English)
The emerging field of Vision-Language-Action (VLA) for humanoid robots faces several fundamental challenges, including the high cost of data acquisition, the lack of a standardized benchmark, and the significant gap between simulation and the real world. To overcome these obstacles, we propose RealMirror, a comprehensive, open-source embodied AI VLA platform. RealMirror builds an efficient, low-cost data collection, model training, and inference system that enables end-to-end VLA research without requiring a real robot. To facilitate model evolution and fair comparison, we also introduce a dedicated VLA benchmark for humanoid robots, featuring multiple scenarios, extensive trajectories, and various VLA models. Furthermore, by integrating generative models and 3D Gaussian Splatting to reconstruct realistic environments and robot models, we successfully demonstrate zero-shot Sim2Real transfer, where models trained exclusively on simulation data can perform tasks on a real robot seamlessly, without any fine-tuning. In conclusion, with the unification of these critical components, RealMirror provides a robust framework that significantly accelerates the development of VLA models for humanoid robots. Project page: https://terminators2025.github.io/RealMirror.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。