arXiv:2608.19443cs.ROcs.SY2026-08

用反馈策略提升采样效率,让机器人控制更稳定可靠。

Hybrid Feedback Sampling for Sample-Efficient Model Predictive Control

论文配图:Hybrid Feedback Sampling for Sample-Efficient Model Predictive Control
图 1 · 摘自论文原文
  • 结合反馈策略优化采样分布,实现局部与全局搜索平衡
  • 在高维不稳定的动态系统中,收敛速度比标准MPPI快,性能优于纯反馈策略
  • 适合复杂接触任务,已在人形机器人实机验证

由于可并行且灵活,基于采样的模型预测控制(MPC)已广泛应用于真实机器人系统。然而,对于高维且开环不稳定的动力系统,为改进控制序列所需的样本数量会随控制时域呈指数增长,导致采样效率低且数值不稳定。本文研究了基于采样的MPC中射击法的不稳定性,表明最优采样提议分布可通过使用优化后的反馈策略实现。该算法称为反馈采样MPC(FS-MPC)。FS-MPC采用混合采样设计,在系统稳定性与计算预算约束下平衡局部与全局搜索。理论分析显示,该方法收敛速度优于标准MPPI,最优性优于标准反馈采样。实验上,在多种富含接触的控制任务(如人形机器人步态-操作与灵巧操作)中,FS-MPC成功解决了传统采样方法难以处理的动力学不稳任务,且显著优于纯反馈策略。最后,我们在真实人形机器人步态与操作任务中验证了该方法的有效性。

原文摘要 · Abstract (English)

Thanks to its parallelizability and flexibility, sampling-based Model Predictive Control (MPC) has become widely popular for controlling real-world robotic systems. However, for high-dimensional and open-loop unstable dynamical systems, the required number of samples to improve the control sequence will grow exponentially with the horizon, leading to poor sample efficiency and numerical instability. This paper investigates the instability of shooting methods in sampling-based MPC and shows that the optimal sampling proposal distribution can be realized by sampling with an optimized feedback policy. We refer to this algorithm as Feedback Sampling MPC (FS-MPC). FS-MPC involves a hybrid sampling design which balances local and global search based on the system stability and the available computation budget. Our theoretical analysis shows that our hybrid sampling approach achieves faster convergence than standard MPPI and better optimality than standard feedback sampling. Empirically, in diverse contact-rich control tasks like humanoid loco-manipulation and dexterous manipulation, we show that FS-MPC successfully tackles dynamically unstable tasks where standard sample-based approaches struggle, and strictly outperforms feedback policies alone. Finally, we validate our method on humanoid robot locomotion and manipulation tasks in the real world.

机器人控制MPC采样优化反馈策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。