arXiv:2510.04354cs.ROcs.AI2025-10被引 17

用少量真实测试校准仿真,高效可靠评估机器人策略性能。

Reliable and Scalable Robot Policy Evaluation with Imperfect Simulators

  • 结合仿真与少量真实数据,用配对样本修正仿真偏差。
  • 仅需20%-25%的硬件实验量,即可获得与全量测试相当的性能置信区间。
  • 适合大规模机器人策略评估,尤其适用于仿真不完美场景。

模仿学习、基础模型和大规模数据集的快速发展,推动了可泛化至多种任务与环境的机器人操作策略。然而,这些策略的严格评估仍具挑战:实践中常仅依赖少量硬件试验,缺乏统计保障。本文提出SureSim框架,通过将大规模仿真与小规模真实测试结合,实现对策略真实性能的可靠推断。核心思路是将真实与仿真评估的融合建模为预测驱动的推断问题,利用少量配对数据修正仿真偏差,并采用非渐近均值估计算法,提供政策平均性能的置信区间。基于物理仿真,在物体与初始状态联合分布上评估扩散策略和多任务微调的π₀策略,结果表明,该方法可在保持性能边界相似的前提下,节省20%-25%的硬件评估工作量。

原文摘要 · Abstract (English)

Rapid progress in imitation learning, foundation models, and large-scale datasets has led to robot manipulation policies that generalize to a wide-range of tasks and environments. However, rigorous evaluation of these policies remains a challenge. Typically in practice, robot policies are often evaluated on a small number of hardware trials without any statistical assurances. We present SureSim, a framework to augment large-scale simulation with relatively small-scale real-world testing to provide reliable inferences on the real-world performance of a policy. Our key idea is to formalize the problem of combining real and simulation evaluations as a prediction-powered inference problem, in which a small number of paired real and simulation evaluations are used to rectify bias in large-scale simulation. We then leverage non-asymptotic mean estimation algorithms to provide confidence intervals on mean policy performance. Using physics-based simulation, we evaluate both diffusion policy and multi-task fine-tuned \(π_0\) on a joint distribution of objects and initial conditions, and find that our approach saves over \(20-25\%\) of hardware evaluation effort to achieve similar bounds on policy performance.

机器人评估仿真校准置信区间策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。