arXiv:2603.13616cs.ROstat.AP2026-03被引 2

用更少实验次数,更准判断机器人策略优劣。

Beyond Binary Success: Sample-Efficient and Statistically Rigorous Robot Policy Comparison

  • 基于随时有效的统计推断,可动态停止实验
  • 相比传统方法减少70%测试量,比现有序列方法省50%
  • 支持细粒度任务进展评估,适合真实机器人对比

通用机器人操作策略能力日益增强,但评估仍受限于少量真实硬件回放。这种强资源约束要求更信息丰富的性能指标和可靠高效的评估流程。本文提出一种样本高效、统计严谨的机器人策略比较框架,基于安全的任意时间有效推断(SAVI),采用顺序测试,当达到预设置信水平时即可提前终止。不同于仅适用于二元成功场景的以往方法,本方法统一处理多种实用指标:从离散部分评分的任务进展到连续的回合奖励或轨迹平滑度,涵盖参数与非参数比较问题。在模拟与真实数据上的广泛验证表明,相比标准批处理方法最多减少70%评估负担,比面向二元结果的先进序列方法减少50%,且保持统计严谨性。实证显示,使用细粒度任务进展指标可更快区分竞争策略。

原文摘要 · Abstract (English)

Generalist robot manipulation policies are becoming increasingly capable, but are limited in evaluation to a small number of hardware rollouts. This strong resource constraint in real-world testing necessitates both more informative performance measures and reliable and efficient evaluation procedures to properly assess model capabilities and benchmark progress in the field. This work presents a novel framework for robot policy comparison that is sample-efficient, statistically rigorous, and applicable to a broad set of evaluation metrics used in practice. Based on safe, anytime-valid inference (SAVI), our test procedure is sequential, allowing the evaluator to stop early when sufficient statistical evidence has accumulated to reach a decision at a pre-specified level of confidence. Unlike previous work developed for binary success, our unified approach addresses a wide range of informative metrics: from discrete partial credit task progress to continuous measures of episodic reward or trajectory smoothness, spanning both parametric and nonparametric comparison problems. Through extensive validation on simulated and real-world evaluation data, we demonstrate up to 70% reduction in evaluation burden compared to standard batch methods and up to 50% reduction compared to state-of-the-art sequential procedures designed for binary outcomes, with no loss of statistical rigor. Notably, our empirical results show that competing policies can be separated more quickly when using fine-grained task progress than binary success metrics.

机器人评估统计推断样本效率策略比较

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。