用目标条件值学习让机器人实时多任务控制,省时又灵活。
Goal-Conditioned Terminal Value Estimation for Real-time and Multi-task Model Predictive Control
- 通过目标条件值学习减少每次决策的计算量。
- 在斜坡上实现机器人轨迹跟踪,响应时间低于控制周期。
- 适合需要快速切换任务的动态控制场景。
虽然模型预测控制(MPC)能在每个时间步求解最优控制问题以实现非线性反馈控制,但其计算开销大,难以在控制周期内完成策略优化。为解决此问题,本文提出一种基于目标条件终端值学习的MPC框架,可在降低计算成本的同时实现多任务策略优化。通过引入分层控制结构,上层轨迹规划器输出目标条件轨迹,使机器人能够生成多样化动作。在双足倒立摆机器人模型上的实验表明,该方法结合目标条件值学习与上层轨迹规划器,可实现真正意义上的实时控制,在斜坡地形上成功追踪目标轨迹。
原文摘要 · Abstract (English)
While MPC enables nonlinear feedback control by solving an optimal control problem at each timestep, the computational burden tends to be significantly large, making it difficult to optimize a policy within the control period. To address this issue, one possible approach is to utilize terminal value learning to reduce computational costs. However, the learned value cannot be used for other tasks in situations where the task dynamically changes in the original MPC setup. In this study, we develop an MPC framework with goal-conditioned terminal value learning to achieve multitask policy optimization while reducing computational time. Furthermore, by using a hierarchical control structure that allows the upper-level trajectory planner to output appropriate goal-conditioned trajectories, we demonstrate that a robot model is able to generate diverse motions. We evaluate the proposed method on a bipedal inverted pendulum robot model and confirm that combining goal-conditioned terminal value learning with an upper-level trajectory planner enables real-time control; thus, the robot successfully tracks a target trajectory on sloped terrain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。