用物理规律正则化神经控制器,让仿真训练的机器人在真实世界稳定运行。
Quantifying and Visualizing Sim-to-Real Gaps: Physics-Guided Regularization for Reproducibility
- 将PID增益作为未建模动态的代理,通过实验测量真实机器人的有效比例增益
- 训练时惩罚神经控制器输入输出敏感度与实测增益的偏差,提升仿真实现一致性
- 适用于低成本、高齿轮比机器人,无需复杂标定,适合工程落地
使用领域随机化进行机器人控制的仿真到现实迁移通常依赖低减速比、可反向驱动的执行器,但当仿真与现实差距增大时,这些方法会失效。受传统PID控制器启发,我们将其增益重新解释为复杂未建模系统动态的代理。我们提出一种基于物理的增益正则化方案,通过简单的现实实验测量机器人有效比例增益,并在训练中惩罚神经控制器局部输入输出敏感度与该值的偏差。为避免朴素领域随机化的过度保守倾向,还对控制器施加当前系统参数的条件约束。在配备110:1减速箱的现成两轮平衡机器人上,我们的增益正则化、参数条件化的RNN实现的角向调节时间与仿真高度一致;而纯领域随机化策略则持续振荡,存在显著的仿真-现实差距。结果表明,该框架轻量、可复现,适用于低成本机器人硬件。
原文摘要 · Abstract (English)
Simulation-to-real transfer using domain randomization for robot control often relies on low-gear-ratio, backdrivable actuators, but these approaches break down when the sim-to-real gap widens. Inspired by the traditional PID controller, we reinterpret its gains as surrogates for complex, unmodeled plant dynamics. We then introduce a physics-guided gain regularization scheme that measures a robot's effective proportional gains via simple real-world experiments. Then, we penalize any deviation of a neural controller's local input-output sensitivities from these values during training. To avoid the overly conservative bias of naive domain randomization, we also condition the controller on the current plant parameters. On an off-the-shelf two-wheeled balancing robot with a 110:1 gearbox, our gain-regularized, parameter-conditioned RNN achieves angular settling times in hardware that closely match simulation. At the same time, a purely domain-randomized policy exhibits persistent oscillations and a substantial sim-to-real gap. These results demonstrate a lightweight, reproducible framework for closing sim-to-real gaps on affordable robotic hardware.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。