arXiv:2512.16881cs.ROcs.LG2025-12被引 30

用视频重建真实场景,实现机器人政策高保真仿真评估。

PolaRiS: Scalable Real-to-Sim Evaluations for Generalist Robot Policies

  • 通过神经重建将实拍视频转为可交互仿真环境。
  • 在未见过的仿真环境中实现零样本评估,相关性显著提升。
  • 适合研究通用机器人策略与仿真评估的开发者使用。

机器人学习研究的一大挑战在于准确衡量和比较机器人策略的性能。传统实地测试因随机性、可复现性差且耗时长而困难,尤其对覆盖多种场景与任务的通用策略更难评估。仿真评估虽具可扩展性,但现有仿真基准与真实世界间的视觉与物理差距导致其难以有效指导策略优化。此外,构建逼真多样的仿真环境通常需大量人工投入。为此,我们提出PolaRiS:一种基于神经重建的实时到仿真评估框架,可将短时视频扫描转化为可交互仿真环境,并设计简单数据协同训练方法,弥合残余真实-仿真差异,实现未见仿真环境中的零样本评估。通过大量仿真与真实世界配对测试,证明PolaRiS评估结果与真实世界通用策略表现的相关性远超现有仿真基准。其简洁性也支持快速生成多样化仿真环境,推动下一代机器人基础模型的分布式、普惠化评估。

原文摘要 · Abstract (English)

A significant challenge for robot learning research is our ability to accurately measure and compare the performance of robot policies. Benchmarking in robotics is historically challenging due to the stochasticity, reproducibility, and time-consuming nature of real-world rollouts. This challenge is exacerbated for recent generalist policies, which has to be evaluated across a wide variety of scenes and tasks. Evaluation in simulation offers a scalable complement to real world evaluations, but the visual and physical domain gap between existing simulation benchmarks and the real world has made them an unreliable signal for policy improvement. Furthermore, building realistic and diverse simulated environments has traditionally required significant human effort and expertise. To bridge the gap, we introduce Policy Evaluation and Environment Reconstruction in Simulation (PolaRiS), a scalable real-to-sim framework for high-fidelity simulated robot evaluation. PolaRiS utilizes neural reconstruction methods to turn short video scans of real-world scenes into interactive simulation environments. Additionally, we develop a simple simulation data co-training recipe that bridges remaining real-to-sim gaps and enables zero-shot evaluation in unseen simulation environments. Through extensive paired evaluations between simulation and the real world, we demonstrate that PolaRiS evaluations provide a much stronger correlation to real world generalist policy performance than existing simulated benchmarks. Its simplicity also enables rapid creation of diverse simulated environments. As such, this work takes a step towards distributed and democratized evaluation for the next generation of robotic foundation models.

机器人仿真评估神经重建通用策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。