用模拟数据下注提升机器人真实性能评估的精度与可信度
Betting for Sim-to-Real Performance Certificates

- 通过模拟数据下注机制动态优化真实性能评估
- 实测显示证书宽度比基线平均缩短51.6%
- 特别适合样本极少的真实场景,如机器人在线评估
在机器人系统测试中,性能证书需以指定置信度覆盖真实均值。由于真实试验成本高,样本量小,传统证书常过于宽松。本文提出一种基于模拟器的下注框架:在每次真实结果出现前,操作员根据大量模拟结果下注,依据实际结果盈亏调整对模拟器的信任。该方法实现了三个贡献:(i) 算法将大规模模拟器与有效下注关联,并用累积财富生成证书;(ii) 证明证书具有任意时间有效性,无论模拟器银行如何;(iii) 通过财富-后悔边界指导算法与模拟器配置。在合成分布和真实机器人测试中,该方法相比经典与先进基线,证书宽度平均缩小51.6%±16%,在极小样本(≤30)条件下仍实现32.26%±8%的缩减。
原文摘要 · Abstract (English)
Consider a typical test of a robot system: one observes a sequence of outcomes concerning some aspect of interest (crash or no crash, tracking error, time to completion), and reports a mean (crash risk, average error, mean time to completion) and, more importantly, an interval guaranteed to contain that mean at a prescribed confidence, referred to as a performance certificate. Given expensive real-world trials, the sample size is therefore small, and the certificate is often loose. Now consider the same procedure, except that before each real outcome is revealed, the operator ``peeks'' at a large bank of simulated results, and places a bet on where the real outcome will land. As the real outcomes settle the bets, the operator gains or loses wealth. One's ``trust'' over simulators also shifts within the portfolio. This paper develops that idea into a sim-to-real betting certificate framework with three contributions: (i) An algorithm that links a scalable bank of simulators to effective bets, and the accumulated betting wealth to the certificate. (ii) A proof that the returned certificate is anytime valid, covering the true mean with the prescribed probability, using any simulator bank. (iii) The guaranteed wealth-regret bounds yield configuration principles for the proposed algorithm and simulator bank design to deliver tight certificates. Experiments across synthetic distributions and real-world robot tests, covering both replayed standardized testing outcomes and online runtime evaluation, show the proposed method narrows the certificate by $51.6\%\pm16\%$ against classic and state-of-the-art baselines, and by $32.26\%\pm8\%$ in the extremely limited-sample regime ($\leq30$ samples).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。