用少量真实数据校准仿真,精准预测机器人在不同场景下的实际表现。
SCAPE: Scenario-Conditioned Simulation-Augmented Policy Evaluation

- 基于有限真实与仿真配对样本,结合大规模仿真推演,预测具体场景下的真实性能。
- 在自动驾驶和四足机器人任务中,场景级预测误差降低14.5%~34.7%。
- 适合需安全部署的机器人系统,尤其关注特定场景适配性与测试效率。
可靠性能评估是机器人学习策略在真实环境中部署的核心瓶颈。真实测试虽准确但成本高且难以扩展,而仿真测试虽易扩展却存在模拟到现实的偏差。现有仿真增强方法结合少量真实轨迹与大量仿真代理,但仅关注初始条件与部署设置下的平均性能,忽视场景特异性,难以指导何时何地安全部署。本文提出SCAPE——一种场景条件化的仿真增强策略评估框架,利用少量配对的仿真与真实样本及大规模仿真推演,预测特定场景下的真实性能。SCAPE在训练前校正仿真标签的模拟到现实偏差,并通过可认证预测校准不确定性。我们在自动驾驶与四足速度追踪任务上验证了SCAPE。在模拟到模拟研究中,相比场景条件神经网络与整体统计基线,其场景级预测误差平均降低4.9%/34.7%(驾驶)与14.5%/27.7%(四足)。我们进一步在物理机器人Unitree Go2上部署速度追踪策略。SCAPE还提升了测试样本效率,生成更窄的校准预测区间,对外分布场景泛化能力更强,并支持细粒度部署策略。
原文摘要 · Abstract (English)
Reliable performance evaluation is a central bottleneck for deploying robot-learning policies in real-world conditions. Real-world testing is faithful but costly and difficult to scale, whereas simulation-based testing scales easily but is inevitably biased by the sim-to-real gap. Existing simulation-augmented methods combine limited real-world rollouts with abundant simulation proxies, but focus on performance averaged over initial conditions and deployment settings. Such population-level averages obscure scenario-specific variation and provide limited guidance about when and where a policy can be safely deployed. We propose SCAPE, a scenario-conditioned simulation-augmented policy evaluation framework that predicts scenario-conditioned real-world policy performance using limited paired sim-and-real samples and large-scale simulation rollouts. SCAPE corrects sim-to-real bias in simulation labels before training the prediction model and calibrates prediction uncertainty through conformal prediction. We validate SCAPE on autonomous driving and quadruped velocity tracking. In sim-to-sim studies, SCAPE reduces scenario-level prediction error by 4.9%/34.7% (driving) and 14.5%/27.7% (quadruped) relative to scene-conditioned neural and aggregate statistical baselines on average. We further evaluate a velocity-tracking policy deployed on a physical Unitree Go2. SCAPE also improves testing sample efficiency, produces narrower calibrated prediction intervals, generalizes better to out-of-distribution scenarios, and enables fine-grained deployment strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。