用视觉语言模型实现真实与仿真间校准,自动评估机器人操作质量。
R2S-Eval: Robot Evaluation with Real-to-Sim Calibration via Vision-Language Models

- 通过真实到仿真的校准生成仿真轨迹视频,减少硬件实验次数。
- 利用视觉语言模型对比视频质量,得出与人类偏好一致的排名结果。
- 适合需要高效、稳定、高质量评估的机器人研发团队使用。
随着通用视觉-语言-动作(VLA)模型在物理机器人上的部署,机器人操作策略的评估变得愈发重要。传统真实世界评估存在劳动密集、结果不稳定、信息量不足等问题,需重复硬件测试、手动重置场景并持续监控,且仅依赖成功率指标,难以反映执行质量。相比之下,人类通过观察完整行为来评估性能。为此,我们提出R2S-Eval:结合真实到仿真校准与视觉语言模型(VLM)偏好评估的评测流程。真实到仿真组件在与真实环境校准的仿真器中高效生成轨迹视频,降低硬件试验需求;VLM评估器对视频进行质量打分并输出成对偏好,进而聚合为策略排名。我们还设计了验证协议,以确保评估结论可靠,并缓解传统评估的主要挑战。仿真与真实实验均表明,R2S-Eval可产生稳定可靠的策略结论,与人类偏好高度一致,显著减少重复硬件操作,揭示出二元成功标签无法捕捉的行为质量差异。总体而言,R2S-Eval将机器人评估从人工计数转向自动化、统计稳定且质量敏感的评测范式。项目页面:https://r2s-eval.github.io。
原文摘要 · Abstract (English)
Evaluating robot manipulation policies is becoming increasingly important as generalist models, particularly vision-language-action (VLA) models, are deployed on physical robots. However, conventional real-world evaluation remains labor-intensive, unstable, and insufficiently informative. It requires repeated hardware trials, manual scene resets, and continuous operator monitoring, may produce different policy rankings across repeated evaluations, and primarily relies on success-rate metrics that provide limited information about execution quality. In contrast, humans assess robot performance by observing and comparing complete behaviors rather than relying solely on binary success outcomes. To this end, we propose R2S-Eval, an evaluation pipeline that combines real-to-sim calibration with vision-language model (VLM) preference evaluation. The real-to-sim component efficiently generates rollout videos in a simulator calibrated to the real-world evaluation setting, thereby reducing the need for repeated hardware trials. The VLM evaluator assesses the execution quality of rollout videos and produces pairwise preferences, which are subsequently aggregated into policy rankings. We further introduce a protocol to assess whether the proposed evaluation pipeline yields validated policy conclusions while mitigating the key challenges of conventional real-world evaluation. Experiments in both simulation and real-world settings demonstrate that R2S-Eval produces reliable and stable policy conclusions, achieves agreement with human preferences, substantially reduces repeated hardware-operation effort, and reveals behavior-quality differences that are not captured by binary success labels. In general, R2S-Eval advances robot evaluation from manual success counting toward automated, statistically stable, and quality-aware evaluation of robot behavior. Project page: https://r2s-eval.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。