用分布方法提升机器人视觉语言动作模型评估精度
PhAIL: A Real-Robot VLA Benchmark and Distributional Methodology

- 采用成功时间累积分布函数作为评估基础,替代传统成功率
- 在30次实验内分辨出相近模型差异,传统方法无法做到
- 适合评估机器人策略性能,尤其关注操作速度与可靠性
现实世界中对视觉-语言-动作(VLA)策略的评估仍依赖固定超时下的二元成功率,每条件仅进行≤25次滚动实验,且通常缺乏置信区间或配对统计检验;此类样本量难以可靠区分相近性能。本文提出PhAIL(物理人工智能排行榜,https://phail.ai),一个基于Franka FR3机器人的开放真实机器人基准,包含数据集、每轮实验的产物记录及端到端参考实现,引入分布式评估方法:以成功时间的累积分布函数(CDF)为评估原语,并分为两个任务。第一个是通过人类相对吞吐量(HRT,带自举置信区间)评分,以同装置下的人类遥操作为锚点;第二个是显著性检验(柯尔莫戈洛夫-斯米尔诺夫检验,按物体计算并跨物体宏平均)。在四个公开可获取的VLAs上,宏平均KS检验在每(模型,物体)单元≤30次实验下解决了两组接近对比(GR00T vs. ACT,OpenPI vs. ACT),而二元阈值指标未能分辨;最接近的一对(OpenPI vs. GR00T)仍在预算范围内未解决。表现最佳的VLA每操作比人类参考慢约7倍(RMST比值)。
原文摘要 · Abstract (English)
Real-world evaluation of vision-language-action (VLA) policies still rests on binary success rate at a fixed timeout with $N \le 25$ rollouts per condition, almost always without confidence intervals or paired statistical comparison; these cohort sizes struggle to resolve close comparisons reliably. We introduce PhAIL (Physical AI Leaderboard, https://phail.ai), an open real-robot benchmark on a Franka FR3 (dataset, per-rollout artifacts, and end-to-end reference implementation) of a distributional evaluation methodology: the time-to-success cumulative distribution function (CDF) as the evaluation primitive, with two separated jobs. The first is scoring via Human-Relative Throughput (HRT), a dimensionless scalar with bootstrap confidence intervals, anchored to same-fixture human teleoperation. The second is a significance test (Kolmogorov-Smirnov, computed per-object and macro-averaged across objects). On four publicly-available VLAs, the macro-averaged KS test resolves two close comparisons (GR00T vs. ACT, OpenPI vs. ACT) at $N \le 30$ rollouts per (model, object) cell where binary-threshold metrics do not; the closest pair (OpenPI vs. GR00T) remains unresolved within our budget. The best evaluated VLA is $\sim 7\times$ slower per operation (RMST ratio) than the human reference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。