arXiv:2503.10966cs.ROstat.ML2025-03中稿 · RSS 2025被引 13

提出一种自适应停止的策略比较方法,显著减少机器人评估所需的试验次数。

Is Your Imitation Learning Policy Better than Mine? Policy Comparison with Near-Optimal Stopping

  • 采用序列化统计检验,根据中间结果动态决定是否继续试验
  • 在真实机器人实验中减少最多32%的评估试验次数
  • 特别适用于困难任务对比,可节省超160次仿真滚动

模仿学习使机器人能在复杂精细操作场景中完成长时程任务。随着新方法不断涌现,必须通过重复试验严格评估并比较其与基线性能。然而,由于人力成本高和策略推理吞吐量有限,策略比较面临小样本限制(如仅10或50次试验)。本文提出一种新型小样本统计框架,用于严谨比较两个策略。现有工作依赖批量测试,需预先固定试验次数,缺乏灵活性,且追加试验可能引发无意的p值操纵,破坏统计可靠性。相比之下,本文提出的检验方法为序列式,允许研究者根据中间结果决定是否继续试验,自适应调整样本量以匹配任务难度,节省大量时间和精力,同时保证概率正确性。大量数值模拟与真实机器人实验表明,该方法实现近最优停止,在保持统计效力与概率正确性的前提下,相较最先进基线减少最多32%的试验次数;尤其在最困难的对比任务中表现突出,在多任务场景下节省超过160次仿真滚动。

原文摘要 · Abstract (English)

Imitation learning has enabled robots to perform complex, long-horizon tasks in challenging dexterous manipulation settings. As new methods are developed, they must be rigorously evaluated and compared against corresponding baselines through repeated evaluation trials. However, policy comparison is fundamentally constrained by a small feasible sample size (e.g., 10 or 50) due to significant human effort and limited inference throughput of policies. This paper proposes a novel statistical framework for rigorously comparing two policies in the small sample size regime. Prior work in statistical policy comparison relies on batch testing, which requires a fixed, pre-determined number of trials and lacks flexibility in adapting the sample size to the observed evaluation data. Furthermore, extending the test with additional trials risks inducing inadvertent p-hacking, undermining statistical assurances. In contrast, our proposed statistical test is sequential, allowing researchers to decide whether or not to run more trials based on intermediate results. This adaptively tailors the number of trials to the difficulty of the underlying comparison, saving significant time and effort without sacrificing probabilistic correctness. Extensive numerical simulation and real-world robot manipulation experiments show that our test achieves near-optimal stopping, letting researchers stop evaluation and make a decision in a near-minimal number of trials. Specifically, it reduces the number of evaluation trials by up to 32% as compared to state-of-the-art baselines, while preserving the probabilistic correctness and statistical power of the comparison. Moreover, our method is strongest in the most challenging comparison instances (requiring the most evaluation trials); in a multi-task comparison scenario, we save the evaluator more than 160 simulation rollouts.

模仿学习策略评估统计检验机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。