arXiv:2608.14284cs.ROcs.CV2026-08

用视频生成精细进程曲线,评估机器人执行过程的细节表现。

PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment

论文配图:PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment
图 1 · 摘自论文原文
  • 将机器人轨迹视频转为密集进度曲线,量化过程中的细微进展。
  • 新增3项指标,涵盖失败后恢复、失败侧进展与成功侧执行质量。
  • 提供可复现的评估工具套件,适合研究者优化和对比模型性能。

精细的机器人评估对理解具身模型至关重要,需超越简单的成败判断和规则评分。我们提出PRM-as-a-Judge 1.5,一个用于机器人过程评估的工具包,能将回放视频转化为密集进度曲线,并衍生出多项细粒度指标。该版本在1.0基础上引入三项新指标,分别刻画失败侧进展、回落后的恢复能力以及成功侧执行质量,帮助用户深入理解具身模型的能力。基于基准测试的回放视频,我们开展了全面评估,提供细粒度指标结果与关键发现。此外,我们提出RoboPulse++,用于评估过程奖励模型(PRM)的可靠性,为评估者提供更精准的测试平台。我们还发布了用户友好的评估套件,包含基准数据、指标实现与可视化工具,支持可复现的操控过程评估。我们呼吁社区重新思考机器人的评估方式,以透明、程序化和可复现的评估为基础,推动下一代具身智能的发展。

原文摘要 · Abstract (English)

Fine-grained robotic evaluation matters for understanding embodied models, going beyond binary success rates and rule-based process scores. We present PRM-as-a-Judge 1.5, a toolkit for robot process assessment that turns rollout videos into dense progress curves and derives multiple fine metrics. PRM-as-a-Judge 1.5 introduces three metrics, building on version 1.0, that characterize failure-side progress, post-drawdown recovery, and success-side execution quality, helping users understand embodied model capability. Based on the rollout videos from benchmarks, we perform a comprehensive assessment of the embodied models, providing some fine-grained metric results and key findings. We further introduce RoboPulse++ to evaluate the reliability of process reward models (PRM), providing evaluators with a more accurate testing platform. Moreover, we release a user-friendly assessment suite, including the benchmark, metric implementation, and visualization tools, to support reproducible manipulation process evaluation. We call on the community to rethink how robots are evaluated and establish transparent, procedural, and reproducible assessment as a foundation for the next generation of embodied intelligence.

机器人评估具身智能过程度量可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。