揭穿机器人抓取基准测试的四大漏洞,提醒别误信分数代表真能力。
What Are We Actually Benchmarking in Robot Manipulation?

- 针对四种失效模式设计诊断方法,检验基准测试是否可靠。
- LIBERO和CALVIN多个指标不显著,部分模型靠捷径取胜。
- 建议作者和审稿人用新工具审计结果,避免盲目宣称进展。
机器人基准测试分数通常在单一评估设置下衡量成功,却被普遍当作通用操作能力的证据。我们识别出四种削弱或无效化基准测试作为能力代理的失败模式:捷径可解性、缺乏统计显著性、渐进式过拟合和数据源依赖性。为此提出每种模式对应的诊断方法,并对LIBERO、CALVIN、SimplerEnv、RoboCasa和RoboTwin 2.0进行了审计。结果显示,LIBERO和CALVIN在多个诊断中失败;尽管近期宣传较少,RoboCasa和RoboTwin 2.0表现更稳健。在LIBERO上,一个无语言编码器的0.09B参数探针模型得分接近现有最优水平,多数宣称提升未通过统计显著性验证。在CALVIN上,将方块姿态随机化后,所有测试策略性能均下降。我们公开四个诊断方法及参考实现,供作者与审稿人使用,以判断基准分数是否真正反映进步。代码与资源见https://ripl.github.io/manipulation_benchmark_audit/。
原文摘要 · Abstract (English)
A robotics benchmark score measures success under one fixed evaluation setup, yet is routinely treated as evidence of general manipulation capability. We identify four failure modes, each of which weakens or invalidates a benchmark's role as a valid proxy for that capability: shortcut solvability, lack of statistical significance, creeping overfitting, and data-source dependence. We propose one diagnostic per failure mode. We audit LIBERO, CALVIN, SimplerEnv, RoboCasa, and RoboTwin 2.0 under these diagnostics. LIBERO and CALVIN fail multiple diagnostics. RoboCasa and RoboTwin 2.0 fail fewer, despite appearing far less often in recent progress claims. On LIBERO, a 0.09B probe with no language encoder scores at or near reported SOTA, and most reported gains are not provably statistically significant. On CALVIN, randomizing block poses within the training range drops performance for every tested policy. We release the four diagnostics with reference implementations for authors and reviewers to apply before treating a benchmark score as evidence of progress. Code and artifacts are available at https://ripl.github.io/manipulation_benchmark_audit/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。