arXiv:2505.08216cs.ROcs.SY2025-05被引 5

提出可保证重复性的统计查询算法,让机器人测试结果更可信。

Rethink Repeatable Measures of Robot Performance with Statistical Query

  • 对统计查询算法进行轻量级修改,确保测试结果一致。
  • 在3类典型场景中验证了方法的准确性与效率边界。
  • 适合需标准化测试的机器人研发与评估团队使用。

针对机器人性能评估中普遍存在的重复性难题,本文聚焦于标准化测试中的统计查询(Statistical Query, SQ)算法。随着系统日益复杂和随机化,传统测试难以保证不同时间、地点或人员执行时结果一致。为此,提出一种通用、轻量、自适应的算法改进方法,适用于蒙特卡洛采样、重要性采样及自适应重要性采样等各类SQ算法,可严格保证测试结果的可重复性,并提供准确性和效率的理论边界。在三类典型场景中验证:(i) 操控器的标准测试,(ii) 自动驾驶车辆运行风险评估的智能测试算法,(iii) 人形机器人行走任务中的指令跟踪性能评估。结果表明,该方法在保持高效的同时显著提升了测试的一致性。

原文摘要 · Abstract (English)

For a general standardized testing algorithm designed to evaluate a specific aspect of a robot's performance, several key expectations are commonly imposed. Beyond accuracy (i.e., closeness to a typically unknown ground-truth reference) and efficiency (i.e., feasibility within acceptable testing costs and equipment constraints), one particularly important attribute is repeatability. Repeatability refers to the ability to consistently obtain the same testing outcome when similar testing algorithms are executed on the same subject robot by different stakeholders, across different times or locations. However, achieving repeatable testing has become increasingly challenging as the components involved grow more complex, intelligent, diverse, and, most importantly, stochastic. While related efforts have addressed repeatability at ethical, hardware, and procedural levels, this study focuses specifically on repeatable testing at the algorithmic level. Specifically, we target the well-adopted class of testing algorithms in standardized evaluation: statistical query (SQ) algorithms (i.e., algorithms that estimate the expected value of a bounded function over a distribution using sampled data). We propose a lightweight, parameterized, and adaptive modification applicable to any SQ routine, whether based on Monte Carlo sampling, importance sampling, or adaptive importance sampling, that makes it provably repeatable, with guaranteed bounds on both accuracy and efficiency. We demonstrate the effectiveness of the proposed approach across three representative scenarios: (i) established and widely adopted standardized testing of manipulators, (ii) emerging intelligent testing algorithms for operational risk assessment in automated vehicles, and (iii) developing use cases involving command tracking performance evaluation of humanoid robots in locomotion tasks.

机器人测试统计查询可重复性评估标准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。