用平滑神经代理提升腿式机器人的学习型模型预测控制可靠性
Learning Legged MPC with Smooth Neural Surrogates
- 设计可调平滑度的神经网络代理,支持接触事件下的轨迹优化
- 在复杂运动任务中,成功率从0/5提升至5/5,累计成本降低2-50倍
- 适合需要高鲁棒性腿式机器人控制的研究者与工程师
深度学习与模型预测控制(MPC)在腿式机器人中可互补协同,但将学习模型融入在线规划仍具挑战。当动力学由神经网络学习时,会面临三大问题:(1) 接触事件间的刚性突变可能源自数据;(2) 可能引入非物理的局部不光滑;(3) 训练数据集因状态快速变化导致模型误差呈现非高斯分布。针对(1)(2),提出平滑神经代理(smooth neural surrogate),通过可调平滑度的神经网络提供用于轨迹优化的精确预测与导数;针对(3),采用重尾似然函数训练模型,更匹配实际动力学误差分布。该方法显著提升学习型腿式MPC的可靠性、可扩展性与泛化能力。在零样本、难度递增的运动任务中,对简单行为累计成本平均降低10%-50%;在标准神经动力学常失效的复杂场景下,成功率从0/5升至5/5,累计成本降低约2-50倍,体现数量级的鲁棒性提升,而非小幅性能改进。
原文摘要 · Abstract (English)
Deep learning and model predictive control (MPC) can play complementary roles in legged robotics. However, integrating learned models with online planning remains challenging. When dynamics are learned with neural networks, three key difficulties arise: (1) stiff transitions from contact events may be inherited from the data; (2) additional non-physical local nonsmoothness can occur; and (3) training datasets can induce non-Gaussian model errors due to rapid state changes. We address (1) and (2) by introducing the smooth neural surrogate, a neural network with tunable smoothness designed to provide informative predictions and derivatives for trajectory optimization through contact. To address (3), we train these models using a heavy-tailed likelihood that better matches the empirical error distributions observed in legged-robot dynamics. Together, these design choices substantially improve the reliability, scalability, and generalizability of learned legged MPC. Across zero-shot locomotion tasks of increasing difficulty, smooth neural surrogates with robust learning yield consistent reductions in cumulative cost on simple, well-conditioned behaviors (typically 10-50%), while providing substantially larger gains in regimes where standard neural dynamics often fail outright. In these regimes, smoothing enables reliable execution (from 0/5 to 5/5 success) and produces about 2-50x lower cumulative cost, reflecting orders-of-magnitude absolute improvements in robustness rather than incremental performance gains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。