用扩散模型在线校准仿真物理参数,让机器人在真实环境更稳健。
Neural Fidelity Calibration for Informative Sim-to-Real Adaptation
- 用条件得分网络动态校准仿真中的物理系数和环境不确定性。
- 在异常场景下微调策略,提升真实世界适应能力。
- 适合需要高鲁棒性导航的复杂物理场景机器人应用。
深度强化学习可将敏捷运动与导航技能从模拟器无缝迁移到现实世界。然而,通过领域随机化或对抗方法弥合模拟到现实的差距通常需要专家级物理知识以保证策略鲁棒性。即使如此,最先进的模拟器仍可能无法捕捉所有真实细节,且重建环境可能因感知不确定性引入误差。为此,我们提出神经保真度校准(NFC),一种利用条件得分型扩散模型在机器人执行过程中在线校准模拟器物理系数和残差保真度域的新框架。残差保真度反映模拟模型相对于真实动态的偏移,并捕捉感知环境的不确定性,使我们能够基于推断分布采样真实环境用于策略微调。该框架在三个方面具备信息性和自适应性:(a) 仅在异常场景下微调预训练策略;(b) 通过预训练 NFC 的提议先验构建序列式在线 NFC,降低扩散模型训练负担;(c) 当 NFC 不确定性较高可能损害策略改进时,采用乐观探索实现幻觉策略优化。我们的框架在具有高维参数空间的多种机器人上,相比现有最优方法实现了更高的模拟器校准精度。我们研究了残差保真度对策略改进的关键贡献,在仿真与真实世界实验中均验证其有效性。特别地,该方法在雪地断轴等挑战性真实条件下表现出强大机器人导航能力。
原文摘要 · Abstract (English)
Deep reinforcement learning can seamlessly transfer agile locomotion and navigation skills from the simulator to real world. However, bridging the sim-to-real gap with domain randomization or adversarial methods often demands expert physics knowledge to ensure policy robustness. Even so, cutting-edge simulators may fall short of capturing every real-world detail, and the reconstructed environment may introduce errors due to various perception uncertainties. To address these challenges, we propose Neural Fidelity Calibration (NFC), a novel framework that employs conditional score-based diffusion models to calibrate simulator physical coefficients and residual fidelity domains online during robot execution. Specifically, the residual fidelity reflects the simulation model shift relative to the real-world dynamics and captures the uncertainty of the perceived environment, enabling us to sample realistic environments under the inferred distribution for policy fine-tuning. Our framework is informative and adaptive in three key ways: (a) we fine-tune the pretrained policy only under anomalous scenarios, (b) we build sequential NFC online with the pretrained NFC's proposal prior, reducing the diffusion model's training burden, and (c) when NFC uncertainty is high and may degrade policy improvement, we leverage optimistic exploration to enable hallucinated policy optimization. Our framework achieves superior simulator calibration precision compared to state-of-the-art methods across diverse robots with high-dimensional parametric spaces. We study the critical contribution of residual fidelity to policy improvement in simulation and real-world experiments. Notably, our approach demonstrates robust robot navigation under challenging real-world conditions, such as a broken wheel axle on snowy surfaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。